Hybrid-Search-RAG-Pipeline Doesn't Trust Its Own LLM

A production-style RAG stack that retrieves with dense plus BM25, fuses with RRF, reranks with a cross-encoder, and then independently audits every [n] citation the model produces.

11 min read • View on GitHub • More from melisasvr

An editorial illustration of a mechanical workshop. On the left, an ornate brass machine produces a strip of paper tape punched with bracketed number markers. On the right, a simpler machine made of calipers and a magnifying lens mounted on a rail inspects the tape as it emerges. One punched marker on the tape is cracked and the caliper is touching it. A tray of approved tape sits at the far right, and a wire basket holds a single flagged strip with a broken marker.
The pipeline's central move: a generative machine produces citations, and a separate deterministic machine inspects them before anyone believes them.
Key Takeaways

Most RAG tutorials end the same way. Retrieve some chunks, stuff them into a prompt, tell the model to cite its sources with bracketed markers, and ship it. The citations look reassuring on the screen. Almost nobody checks them.

The model is free to invent a citation number, attach a real number to the wrong passage, or write a sentence that has nothing to do with the chunk it points at. The prompt said to cite. The model complied. Compliance is not the same as correctness.

melisasvr/Hybrid-Search-RAG-Pipeline takes the next step. After the Groq-hosted model produces a cited answer, a separate module walks through every marker and decides, without calling an LLM again, whether the citation holds up. The module's docstring calls it a "second line of defense against hallucination."

Two Layers of Trust

The verifier is a two-stage check, and each stage catches a different kind of lie.

Layer 1 is structural. Every marker is checked against the number of passages that were actually retrieved. A citation pointing at [7] when only five chunks exist becomes an invalid_citation with doc_id="unknown". The model invented it, and the system says so out loud. This layer is cheap, deterministic, and it kills the most blatant failure mode outright.

Layer 2 is lexical. For each sentence that carries a marker, the verifier computes a TF-IDF cosine similarity between the sentence (with the marker stripped) and the chunk it cites. It takes the maximum across all sentences pointing at the same chunk, then buckets the result into verified, weak, or unverified using two thresholds.

The sentence splitter is the detail that shows this was debugged against real model output. A citation written after a period ("monthly. [1] Sick leave") belongs to the sentence it closes, so the boundary pattern deliberately swallows trailing markers:

_SENTENCE_BOUNDARY_RE = re.compile(r"[.!?](?:\s*\[\d+\])*")

Step through the verifier yourself: paste a sentence, watch the structural check flag invented markers, then watch the lexical check score each surviving claim against the chunk it cites.

Where the Check Actually Fails

The author flagged the limits of this approach in the code itself, which is rarer than it should be. A lexical overlap score is not an entailment judgment. Two failure modes are baked into the design.

A correct paraphrase with different vocabulary scores low and gets marked unverified. A sentence that copies the keywords from a chunk while inverting its meaning scores high and gets marked verified. The check produces both false negatives and false positives, and it produces them in predictable directions.

The sentence splitter has a second bug that the comments acknowledge. The boundary pattern matches the period in a decimal, so "The policy grants 3.5 days [1]" is split mid-number into "The policy grants 3" and "5 days [1]." For HR and policy content, where numeric answers are the whole point, that distorts the similarity score for exactly the sentences that matter most. And because the vectorizer is refit from scratch for every sentence and chunk pair, its IDF values are computed over two documents. The "TF-IDF" is effectively a stopword-filtered cosine similarity.

Check says verifiedCheck says unverified
Correct, keyword-heavy sentenceCorrect paraphrase in different words (false negative)
Keyword-copied inversion (false positive)Genuinely unsupported claim caught correctly

None of this makes the verifier useless. It makes it a screen, not a judge. The author's own recommendation is to add an LLM-as-judge pass or a trained entailment model on top. What the module gives you today is a fast, offline, deterministic first filter that catches the cheapest hallucinations before they reach a user.

Retrieval, Ranked Two Ways

The verifier is the punchline, but it sits on top of a retrieval stack that is worth understanding. Documents are embedded with MiniLM and stored in ChromaDB as dense vectors. The same text is also tokenized and indexed with BM25 as a sparse keyword index. Dense retrieval finds meaning; BM25 finds exact terms. Neither alone is enough for an internal knowledge base.

The interesting decision is how the two lists get combined. Cosine similarity and BM25 scores live on different, corpus-dependent scales, so a weighted sum requires retuning every time the corpus changes. The project fuses by rank position instead, using Reciprocal Rank Fusion:

fused[chunk_id]["fused_score"] = fused[chunk_id].get("fused_score", 0.0) + 1.0 / (k + rank)

Each chunk's fused score is the sum of 1/(k+rank) across both retrievers, with k defaulting to 60. The raw dense and BM25 scores and ranks are kept alongside the fused score, so the UI can explain why a chunk surfaced. Then the top fused candidates are hydrated from the metadata store and passed through a cross-encoder reranker, which scores each query and chunk pair directly and picks the final set. Retrieval is a funnel: cheap and wide, then narrow and expensive.

A close-up editorial illustration of two hands at a workbench. The left hand holds a strip of paper printed with a sentence and a citation marker. The right hand holds a magnifying lens over a decimal number, and the lens is cutting the paper strip in half exactly along the decimal point, splitting the sentence into two fragments. In the background, faint on a slate, a small weight chart is being rebuilt from scratch.
The sentence splitter matches the period inside a decimal, so a numeric claim gets cut in half before it is ever scored. The background chart is the vectorizer being refit for every single pair.

The Dependency Bet

The stack runs embeddings and reranking through fastembed on ONNX Runtime rather than the more common sentence-transformers on PyTorch. The reason is written into the docstrings: sentence-transformers pulls in torch and transformers, and torch was colliding with other ML tooling on the author's machine.

That is a real engineering constraint shaping a real architecture. The tradeoff is smaller installs and fewer dependency conflicts in exchange for tying yourself to fastembed's model catalog and ONNX export quality. For a reference implementation meant to be cloned and run in minutes, without Docker and without a multi-gigabyte PyTorch download, it is the right side of the bet.

The embedder also enforces a symmetry rule: ingestion and query both L2-normalize their vectors through the same class, guarding against the classic production bug of embedding documents with one model and queries with another. That guard is soft rather than structural, since the two paths are separate methods that could drift, but naming the failure mode in the docstring is itself a form of documentation.

Consistency Is a Choice

The ingestion path persists metadata first, then vectors, then the BM25 index. The comment explains why: a crash mid-index should never leave the vector or keyword stores pointing at chunk IDs that have no metadata. A dangling index entry is an orphan that retrieval cannot resolve. A metadata-only chunk is harmless, because retrieval starts from the indexes.

# Order matters: persist metadata first so a crash mid-index never
# leaves the vector/BM25 stores pointing at chunk_ids with no metadata.
chunk_store.add_chunks(chunks)
vector_store.add_chunks(chunks, embeddings)
bm25_index.add_chunks([c.chunk_id for c in chunks], [c.text for c in chunks])

The retriever has the matching defensive read path, skipping stale index entries with if chunk is None: continue. The delete path does not follow the same ordering, so deletion is not equally protected. The point is not that the code is perfect. The point is that the author reasoned about failure modes explicitly and wrote the reasoning down where the next engineer would find it.

How It Compares

Hybrid RAG is a crowded name. Several projects share the exact phrase, and they emphasize different things: breadth of integrations, offline evaluation rigor, typed contracts, latency, or teaching. Here is how the field looks through that lens.

ProjectLanguageFusionRerankCitation checkEval rigor
melisasvr/Hybrid-Search-RAG-PipelinePythonRRF (k=60)Cross-encoderTwo-layer deterministic verifierOne end-to-end smoke test
reaatech/hybrid-ragTypeScriptConfigurableYesNot documented in listingBenchmarks, ablations, cost tooling
anbsamsam17/hybrid-rag-pipelinePythonHand-written RRFCross-encoderCitation-checked generationrecall@k, nDCG, MRR, bootstrap CIs
metehanulusoy/rag-hybrid-searchPythonRRFOptional cross-encoderCitation verification plus abstentionTyped contracts
kolhesamiksha/Hybrid-Search-RAGPythonConfigurableOptionalCited answersLatency-oriented
jannctu/hybrid-search-ragPythonRRFYesCitationsExplicitly none

A note on lineage: these projects were surfaced by search, and nothing in the available sources establishes that they share code, ancestry, or influence with one another. Shared technique (BM25, RRF, reranking, citations) is the shape of the field, not evidence of descent. Treat the table as a map of emphases, not a family tree.

What It Teaches

The transferable lesson is not that you should copy this pipeline. It is that verification is a stage, not an instruction. Asking a model to cite its sources is a request. Checking the citations it returns is a system. The difference between those two things is the entire gap between a demo and something you would let answer questions about an employee handbook.

The second lesson is quieter. The docstrings read like a senior engineer's lab notebook. This is why we picked fastembed. This is why we rebuild the BM25 index instead of updating it. This is where the check fails. A codebase that names its own limits is easier to trust than one that pretends to have none, and that is a habit worth stealing regardless of which retrieval stack you run.