Skip to content
SHASHWAT // SYSTEM ARCHIVE
∞
SYSTEM.ARTICLE

Building Production RAG: Hybrid FAISS + BM25 vs. Pure Vector Search

avatarShashwat Sharma
8 min read

Pure vector search fails on the queries that matter most in production: exact product codes, rare names, error messages copy-pasted from a log file. VectorLoom is my offline RAG system that fixes this by combining FAISS with BM25 and a cross-encoder reranker, running the whole pipeline under 200ms.


What It Does and Who It's For

VectorLoom is an offline retrieval-augmented generation system. Offline means the index lives on disk, next to the documents, with no external vector database and no network call for retrieval. You point it at a folder of documents and it builds a searchable index locally.

The retrieval layer combines two methods that fail in different ways: dense vector search through FAISS, and sparse lexical search through BM25. A cross-encoder reranker sits on top and re-scores the merged candidate list before anything reaches the language model.

It's built for people who need RAG on private or air-gapped data — internal documentation, codebases, contracts — where sending chunks to a hosted vector database isn't an option, and where the queries aren't clean natural-language questions. Real users paste error messages. They search for part numbers. They type half a variable name and expect the right file to come back.

The Problem That Made Me Build It

Most RAG tutorials demo semantic search on clean, conversational questions: "What are the health benefits of green tea?" Vector search is genuinely good at this. The embedding for that question sits close to the embedding for a paragraph about antioxidants, even if the paragraph never uses the word "benefits."

Production queries don't look like that.

Someone searches for ERR_CONN_RESET_4471. That's not a concept — it's a token. A dense embedding model trained on natural language will place it somewhere in vector space based on vague subword similarity, not because it understands the error code. Meanwhile the document containing that exact string sits three positions below five documents that are semantically "about networking" but don't contain the code at all.

I hit this directly while testing VectorLoom against a support-ticket-style dataset. Queries with product SKUs, exact function names, or quoted error strings had a measurably worse retrieval rate than open-ended questions on the same corpus. Pure vector search was quietly failing on exactly the queries where users expect the most precision, because those are the queries where they already know what they're looking for.

Architecture / How It Works

The pipeline has four stages: chunking and indexing (offline), then retrieval, fusion, and reranking (at query time).

Indexing. Documents are chunked and two indexes are built from the same chunks: a FAISS index over dense embeddings, and a BM25 index over the raw tokenized text. Both live on disk alongside the source documents.

Retrieval. A query hits both indexes independently.

def retrieve(query, k=50):
    dense_hits = faiss_index.search(embed(query), k=k)
    sparse_hits = bm25_index.search(tokenize(query), k=k)
    return dense_hits, sparse_hits

Fusion. The two ranked lists are merged with Reciprocal Rank Fusion (RRF) rather than a raw score blend, because FAISS cosine scores and BM25 scores live on incompatible scales.

def reciprocal_rank_fusion(dense_hits, sparse_hits, k=60):
    scores = {}
    for rank, doc_id in enumerate(dense_hits):
        scores[doc_id] = scores.get(doc_id, 0) + 1 / (k + rank + 1)
    for rank, doc_id in enumerate(sparse_hits):
        scores[doc_id] = scores.get(doc_id, 0) + 1 / (k + rank + 1)
    return sorted(scores.items(), key=lambda x: x[1], reverse=True)

Reranking. The top candidates from the fused list (typically 20-30) go through a cross-encoder that scores each query-document pair jointly, rather than comparing pre-computed embeddings. This step is slower per document but far more accurate, so it only runs on the shortlist, not the full corpus.

Only the reranked top-k chunks are passed to the language model as context.

The Hardest Technical Decision

The hardest decision wasn't picking FAISS or BM25 — using both was obvious once the failure modes were clear. The hard part was how to merge their outputs.

My first attempt normalized both score sets to 0-1 and took a weighted average. This looked reasonable in isolation and fell apart in practice. BM25 scores are unbounded and depend on corpus statistics (term frequency, document length), so a "good" BM25 score on one corpus might be a mediocre score on another. Cosine similarity from FAISS is bounded but compresses into a narrow band for semantically similar documents, so small differences in that band get either flattened or wildly exaggerated after normalization. The blended ranking was inconsistent between corpora and hard to reason about.

I switched to Reciprocal Rank Fusion, which ignores the raw scores entirely and only uses rank position. A document's contribution is 1 / (k + rank), summed across both retrievers. This throws away information — a BM25 score of 40 and a BM25 score of 4 are treated identically if they land at the same rank — but it also removes the scale-mismatch problem completely, since rank position means the same thing regardless of which retriever produced it.

The trade-off is real: RRF can't tell you "this document is a much better match," only "this document ranked higher." For my use case that was the right trade, because the downstream reranker recovers fine-grained scoring anyway. RRF's job is just to make sure the right documents survive into the shortlist that the reranker sees.

What I Measured

I evaluated on a held-out query set split into two buckets: natural-language questions and exact-match queries (error strings, identifiers, quoted phrases), against the same document corpus.

ApproachRecall@10 (natural)Recall@10 (exact-match)p50 Latencyp95 Latency
FAISS only88.4%61.7%41ms68ms
BM25 only71.2%94.1%19ms33ms
Hybrid + RRF89.1%92.8%79ms143ms
Hybrid + RRF + rerank91.6%93.4%156ms187ms

Two things stand out. First, FAISS-only recall on exact-match queries drops over 30 points compared to natural-language queries on the identical corpus — this is the failure mode that started the whole project. Second, the reranker adds real latency (roughly 77ms at p50) for a recall gain of 2.5 points on natural-language queries and half a point on exact-match, since RRF alone already gets most of the way there for exact matches.

The full hybrid-plus-rerank pipeline stayed under my 200ms target at p95, which was the constraint that mattered more than squeezing out another point of recall.

What I'd Do Differently

I'd measure the exact-match failure mode before writing any fusion code, not after. I built the hybrid pipeline on general intuition about vector search limitations, then went looking for numbers to justify it. Running the FAISS-only baseline against a query set with real identifiers and error strings first would have told me exactly how bad the gap was and let me size the BM25 weighting decision with data instead of guesswork.

I'd also profile the reranker earlier. It was the single biggest latency contributor and I added it last, which meant the latency budget for indexing and fusion was already spent by the time I knew how much headroom the reranker actually needed.

Conclusion

Pure vector search optimizes for meaning and loses precision on tokens: codes, names, exact phrases. Pure lexical search optimizes for tokens and loses precision on meaning. Neither one is wrong — they're solving different problems, and production queries are a mix of both.

VectorLoom's answer is to run both retrievers, fuse by rank instead of raw score, and spend the latency budget on a reranker only for the shortlist that survives fusion. The result recovers most of BM25's exact-match strength and most of FAISS's semantic strength, at a latency cost that's still small enough for interactive use.

If you're building RAG on anything other than clean natural-language questions, test your exact-match recall before you ship. It's the failure mode nobody demos.