The unreasonable effectiveness of BM25 for agentic search — Jo Kristian Bergum, Hornet.dev
Summary
A 30-year-old lexical scoring function is making an unexpected comeback in AI agents because LLMs now possess enough general knowledge to formulate better queries—BM25 effectiveness isn't about the algorithm improving, but about having smarter users asking better questions.
Key Takeaways
- BM25 scoring hasn't changed in 30 years, but its effectiveness increased dramatically because modern LLMs with parametric knowledge can formulate superior queries—the 'user' got more powerful, not the algorithm.
- Current LLM context windows (~350,000 tokens) are the limiting factor for retrieval-augmented systems; retrieval is essential because you can't fit all knowledge at once, similar to 1.4MB floppy disk constraints.
- Evaluate agentic search using multi-turn benchmarks like BrowseQA Plus (830 riddle-like questions, 100K document corpus) rather than single-shot queries, since agents can formulate far more queries than humans.
- Agentic search requires three components: capable models with tool use, a harness for exposing retrieval functions (via tool calling or code mode), and a retrieval engine optimized for billion-scale document sets.
- BM25 shifted from being a baseline comparison metric in IR research to a preferred production solution for agent loops because end-to-end accuracy depends heavily on retrieval quality feeding the context window.
Related topics
Transcript Excerpt
[music] So great being here. Uh I'm Joe Bergam. I'm the CEO of Hornet Dev and I'm here today to talk about the unreasonable effectiveness of BM25 for Agentic Search. So how many of you heard about BM25 before? Is it new or is it Well, quite a few. So that's great. Uh I'm also watching the World Cup. Norway is playing there uh against the Ivory Coast second half. Norway is leading so that's good. So yeah and at Hornet we are building retrieval infrastructure for agents and uh I've been working on search and retrieval problems for a long time. As you can tell I'm gray hair. Uh been working in this space for more than 20 years. Um and in the talk today I'll talk about why this kind of 30-year-old lexical scoring function is uh making a strong comeback. First I will talk a little bit about wha…
More from ai.engineer
- Grep or Embeddings? Agentic Search Over Company Documents — George He, LlamaIndex
- Agents That Write Their Own Tools at Runtime — Sandhya Subramani, AWS
- Dashboards Are Dead — Sarah Simionescu, Composio
- The State of AI in Software Development: Data from 400+ Orgs — Justin Reock, DX
- The Death of the Code Review: What the Data Actually Says — Laurie Voss, Arize AI