HeadlinesBriefing HeadlinesBriefing.com

Benchmarking retrieval for agents on messy company knowledge

Hacker News •
×

How teams do retrieval is changing fast. Cursor recently stopped searching your code via embeddings in favour of relying only on grep and indexed search. Others are swapping their traditional rerankers for new models like Jev. One of the most important kinds of knowledge that agents rely on is company knowledge: documentation, tickets, chat messages, internal wikis, and code. Yet most teams cannot tell which retrieval works best on it.

Kapa is a platform for indexing your company knowledge and letting your agents search them for context. The Company Knowledge Bench is what we built for ourselves to improve our own system: 1,000 eval cases annotated from real production data. Fixed retrieval (traditional RAG), agentic grep and Kapa: 7 retrievers on 1,000 eval cases. The takeaway: a frontier model with nothing but grep matches a tuned modern retrieval pipeline at 0.61, but takes five times as long. An optimized agentic retriever (Kapa Deep) does better still: 0.65 in about five seconds.

Public benchmarks do not work for us because none of them cover all the use cases and types of queries that we see. Use cases include developers asking about a product over docs, API specs, code and GitHub issues; sales and employees asking about products over Slack, Confluence, Notion; support teams drafting replies from old tickets. Queries vary from long messages to precise searches like "webhook retry backoff config". The Company Knowledge Bench scores retrieval on completeness, minimality, and source quality.

Source: Hacker News · Summarized by HeadlinesBriefing