Benchmarking retrieval for agents on messy real-world company knowledge - kapa.ai - Instant AI answers to technical questions
Pangram verdict · v3.3
We believe that this text is a mix of AI and human-written content.
AI likelihood · overall
MixedArticle text · 1,545 words · 2 segments analyzed
How teams do retrieval is changing fast, in every domain. Cursor recently stopped searching your code via embeddings in favour of relying only on grep and indexed search. Others are swapping their traditional rerankers for new models like Jev.One of the most important kinds of knowledge that agents rely on is company knowledge: documentation, tickets, chat messages, internal wikis, and code. Yet most teams cannot tell which retrieval works best on it, because they have no way to measure how their retrieval performs on real data.Kapa is a platform for indexing your company knowledge and letting your agents search them for context. You connect your sources, Kapa turns them into one searchable knowledge base, and any agent can query it for the information it needs. Retrieval is the core of what we build, and we change how we do it constantly.The Company Knowledge Bench is what we built for ourselves to improve our own system: 1,000 eval cases annotated from real production data. In this post we explain how it works, and score a few common retrieval implementations on it.Company Knowledge Bench results: score against time per queryFixed retrieval, agentic grep and Kapa: 7 retrievers on 1,000 eval cases. Higher and further left is better.Fixed retrieval (traditional RAG)Agentic retriever with grepKapaBest score for the timeRetrieval score0.40.50.60.70 s5 s10 s15 sTime per query, secondsHybrid search 0.41Hybrid search + rerank 0.50Query decomposition 0.56Kapa Default 0.61Kapa Deep 0.65Agent + grep (Luna) 0.54Agent + grep (Sol) 0.61The takeaway: a frontier model with nothing but grep matches a tuned modern retrieval pipeline at 0.61, but takes five times as long. An optimized agentic retriever (Kapa Deep) does better still: 0.65 in about five seconds.A note on bias: retrieval is core to our product, and all seven retrievers were built by us and use documents ingested with our pipeline.Why public benchmarks do not work for usPublic benchmarks do not work for us because none of them cover all the use cases and types of queries that we see.Use casesTeams index their sources in Kapa and connect an agent to the retrieval, and that agent can serve quite different use cases, each with its own documents and its own people asking questions:Developers ask about a product over its docs, API specs, code and GitHub issues, often through Claude Code.Sales and other employees ask about products, processes and customers over Slack, Confluence, Notion and Google Drive.Support teams draft replies to new tickets from old tickets, help center articles and internal handbooks.QueriesAs these use cases show, many different kinds of agents send queries to Kapa’s retrieval, and the queries look different depending on who wrote them:People put everything into one message: several questions, a pasted error, a reference to something said earlier.Agents like Claude Code write their own, breaking a complex question into short, precise searches, like webhook retry backoff configPublic retrieval benchmarks are mostly too narrow, built around one domain like law or medicine, or too artificial, built from synthetic documents and questions. None of them cover this range, so we built our own.How we score retrieval for agentsBefore we can measure good retrieval, we have to define what it is.A few terms first. A query is what is sent to the retriever. The corpus is everything a team has indexed in Kapa. A chunk is a short piece of it, such as a section of a page. A retriever takes a query and returns the most relevant chunks from the corpus for it.Our definition:The retriever’s goal is to collect a minimal set of chunks that completely answers the query, using the highest-quality sources available.The Company Knowledge Bench rests on three properties:Completeness. The agent you hand the chunks to needs nothing else to answer the query. If the set is incomplete, the agent gives an incomplete or incorrect answer.Minimality. Removing any chunk would make the set incomplete. Extra chunks the query does not need cost money and make it harder for the model to reason.Source preference. A corpus often holds several sets of chunks that could answer the same query, and they are almost never of equal quality. Our benchmark only accepts the preferred ones. Deciding which source is preferred is hard and has a lot of grey zones, but in general it comes down to two factors:Source authority. A dedicated reference page outranks a tutorial, an issue thread, or a blog post that repeats the same fact.Currency. A current source outranks an outdated one. A Slack thread from last week beats a Confluence page last edited three years ago.Beyond these, the benchmark enforces rules for specific situations. For example:The query is ambiguous. The retriever has to return chunks for every reasonable reading of it. Picking one reading is the agent’s job.A problem has several valid solutions. The retriever has to return chunks for each documented solution, so the agent can explain the options and their trade-offs.The corpus holds nothing that answers the query. The retriever has to return the strongest evidence there is: a statement that the feature is not supported, a complete list the feature is missing from, or a documented workaround. If none of those exists, the right result is nothing at all.All of it is written down precisely in our benchmark labelling handbook.What an eval case looks likeOur benchmark is a set of eval cases. Each one consists of a real production query, a snapshot of the corpus as it was when the query was asked, and a retrieval criterion: a boolean expression that specifies which chunks from the corpus are valid to retrieve for the query.Here is a simplified eval case, scored against the chunks one retriever returned:QueryRetrieval criterionChunks the retriever returnedScorePass. The criterion requires chunk_1 and chunk_2, and both were returned.Precision: 2 of 4, or 50%. chunk_3 and chunk_4 are not in the criterion.The forum post is not in the criterion because of source preference. It answers the same part of the query as the pricing page, which plans include SSO, and the pricing page has more source authority. The changelog is simply not relevant. Both count against precision, but for different reasons: one is irrelevant, the other is relevant but not preferred.In total, we created 1,000 eval cases for the Company Knowledge Bench, and they are what the results later in this post are based on.How we created the benchmarkLabelling a thousand eval cases by hand is not feasible. Sampling them properly makes it even harder: the queries come from every type of source we ingest, internal as well as public corpora, and the industries we serve. So to achieve this scale, the labelling has to be done by agents.Of course, letting agents create the benchmark only works if they do it correctly. A wrong criterion means a wrong score for every retriever we test against it. So before the agents can build our benchmark, we need a second benchmark that measures how well they build it: a benchmark for benchmark creation.That second benchmark is a set of 170 eval cases that humans labelled completely by hand, following the same benchmark labelling handbook. Each has a real query, a corpus, and a retrieval criterion a person wrote after working through the corpus themselves, so together they encode the handbook’s rules.So the process was: build agents that can do the same task the humans did for this set, run them on its queries, and compare the criteria they write with the human ones. We kept improving the agents until they reached a high enough agreement rate with the human labellers. Only then did we let them generate the 1,000 eval cases of the benchmark from production queries.The labelling is split between two kinds of agents:Candidate agents scan the full corpus for chunks that might belong in the retrieval criterion. They are built for recall and return a large set of possibly relevant chunks.Criteria agents take those candidates and turn them into the actual retrieval criterion, following the rules in the benchmark labelling handbook.The agents do not have to reproduce the human criteria exactly for the benchmark to give a good signal, but they do have to come close. Getting there took many iterations, on the agents and on the handbook itself. We made sure the agents reached sufficient agreement with the human labellers, and validated them extensively, before moving on to generate the benchmark.The result is labels that are close to what a human would write, and unlike hand labels they scale well beyond the 1,000 eval cases we generated for this benchmark.How the retrievers compareWe ran seven retrievers on the Company Knowledge Bench.
All of them search the same corpora, ingested and chunked the same way, so the only thing that differs is the retrieval strategy.Fixed pipelinesThese are the traditional RAG approaches of the last few years. They run the same steps for every query, with no model deciding what to do next.Hybrid search. The baseline most RAG pipelines start from. The query is matched against a keyword index and an embedding index at once, using gemini-embedding-001 for the embeddings, and the top 15 chunks are returned.Hybrid search with a reranker. The same search, but it retrieves the top 100 chunks and passes them through a reranker, Voyage AI’s rerank-2, which reads the query and each chunk together and scores how well the chunk answers it.