When BM25 beat the embeddings
In short: on legal text, a keyword ranking function from the 1990s wins a whole class of queries that dense embeddings lose: citations, section numbers, defined terms. I run both and fuse the…
- published
- read time
- 5 min
- words
- 905
- lang
- en
- filed under
- Engineering
In short: on legal text, a keyword ranking function from the 1990s wins a whole class of queries that dense embeddings lose: citations, section numbers, defined terms. I run both and fuse the rankings, and I decide by query type, not by an average.
The query that started it
The first version of search I built over legal text was embeddings only. Dense embeddings over chunked legislation and decisions, a vector store, the top results into the model. Plain-language questions work well in a setup like that. Now type a neutral citation, the year, court code and number a lawyer uses to name a case. A common failure looks like a handful of decisions about similar things, and not the case you named.
That's not a bug in the embedding model. It's what embeddings are for. They map text to meaning, and a citation string has almost no meaning. "Section 7" and "section 15" sit very close together in that space, because they look alike and are about the same kind of thing. To a lawyer they're different rules.
BM25 doesn't care about meaning. It scores a passage by how many of the query's exact terms it contains, weighted by how rare each term is across the corpus and adjusted for passage length. A rare token like a case number or a defined term is exactly what it's good at.
Which queries go which way
Reading the misses from a retrieval eval, the split is easy to see once you sort queries by kind. This is the general pattern, not a measured benchmark from my system.
| Query type | Shape of the query | Usually wins | Why |
|---|---|---|---|
| Citation lookup | A neutral citation or a statute and section number | BM25 | Rare exact tokens, no meaning to embed |
| Defined term | A term the act defines, in quotes | BM25 | The exact string is the signal |
| Party name | A company or person named in a case | BM25 | Names are rare tokens |
| Plain-language question | "Can my employer cut my hours without notice?" | Embeddings | None of the statute's words appear in it |
| Concept search | "duty to accommodate, undue hardship" | Close, hybrid best | Shared terms and shared meaning |
A user doesn't know which kind of query they're typing, and a lawyer types all five before lunch. So picking one retriever means losing a row of that table every time.
Fusing two rankings
BM25 scores and cosine similarities live on different scales, so adding them is meaningless unless you normalise, and normalising is fiddly. Reciprocal rank fusion skips the problem. It throws the scores away and only uses each passage's position in each list.
def rrf(rankings, k=60, limit=10):
"""rankings: lists of passage IDs, best first, one list per retriever."""
scores = {}
for ranking in rankings:
for rank, pid in enumerate(ranking, start=1):
scores[pid] = scores.get(pid, 0.0) + 1.0 / (k + rank)
return sorted(scores, key=scores.get, reverse=True)[:limit]
dense = vector_search(query, limit=50)
sparse = bm25_search(query, limit=50)
top = rrf([dense, sparse])
A passage that ranks well in either list gets a good fused score, and one that ranks well in both gets the best. The constant k damps the advantage of being first; 60 is the usual default and I haven't found a reason to move it. Take more candidates from each retriever than you need at the end, so a passage at rank 30 in one list still has a chance.
Backfilling the keyword side
The awkward part was not the fusion. It was that I already had an index full of embedded chunks and no keyword index beside it. Adding BM25 meant a backfill: walk every existing chunk, tokenise it the same way queries will be tokenised, and write it to the keyword index under the same passage ID as the vector.
Three things I'd tell myself before starting:
- Same IDs on both sides. Fusion joins on passage ID. If one side uses chunk IDs and the other uses document IDs, nothing lines up and the fused list quietly degrades to whichever side has more hits.
- Tokenise for law, not for English. A default tokenizer that splits "s. 15(1)" into loose pieces, or drops numbers as noise, throws away exactly the tokens BM25 was brought in for. Keep section numbers and citation parts as tokens.
- Make it resumable. A backfill over a large corpus will stop halfway at least once. Write in batches, record the last ID done, and let a rerun skip what's already there.
Decide by query type, not by the average
The easy mistake is to run the eval, see that embeddings have higher overall recall than BM25, and stop there. An average like that is dominated by plain-language questions, because that's most of what people write when they build a question set. Citation queries are a minority of the rows, so embeddings can lose most of them and still win the average. A lawyer who types a citation and doesn't get the case won't type a second one.
So here is what I'd do in your place. Tag every question in your retrieval eval with its kind. Run BM25 alone, embeddings alone, and the fused list, and print recall at your k for each kind separately. If one retriever wins some rows and loses others, you don't need to pick. Fuse them, and keep the per-kind numbers in front of you every time you change something.
related