Grep or Embeddings? Agentic Search Over Company Documents — George He, LlamaIndex
Summary
Grep-based local file search outperforms vector embeddings for smaller datasets (100-1000 files), but enterprises with millions of mixed-format documents must choose between token-expensive local traversal and pre-indexed retrieval—the key is matching your data scale to the right orchestration primitive.
Key Takeaways
- Scale determines architecture: 100-1000 files work with local grep/file search without token waste; millions of files require pre-indexing or hybrid approaches to avoid burning context.
- Vector databases create maintenance debt at enterprise scale—Anthropic's Claude Code avoided pre-indexing entirely because syncing a vectorized index with constantly-updating documents is harder than the performance gain.
- Code search succeeds with grep because repositories are small, semi-structured, syntactically validated, and locally available—but company data lakes with mixed PDFs, images, and schematics require different retrieval primitives entirely.
- Use MCP servers, CloudMD, or instruction files to maintain traversal hierarchies—don't rely on folder structure alone for enterprise data retrieval since corporate databases lack code's inherent organization.
- Context engineering is really data architecture—the real challenge is keeping the right retrieval primitives in place based on your specific data scale and curation effort capacity.
Related topics
Transcript Excerpt
[music] We're going to get going. Nice to meet you everyone. Uh my name is George. Uh today uh we'll be going over a talk on search document search and uh how we get better results here at Llama Index. Um my name is George. I'm the head of engineering here at Llama Index. We are a series A startup. We focus quite a bit on document parsing, orchestration and uh knowledge management. Um today the topic of the talk will center around how we've helped enterprise customers scale up search GP and generally manage context using aentic orchestration and just general infrastructure tooling. Uh the plan today is going to be going through a discussion or a debate around uh how to best orchestrate and perform document search uh how to get the best results and we'll iterate through their uh technical c…
More from ai.engineer
- Agents That Write Their Own Tools at Runtime — Sandhya Subramani, AWS
- Dashboards Are Dead — Sarah Simionescu, Composio
- The State of AI in Software Development: Data from 400+ Orgs — Justin Reock, DX
- The Death of the Code Review: What the Data Actually Says — Laurie Voss, Arize AI
- Software Engineering Is Becoming Factory Engineering — Zach Lloyd, Warp