Kapa, a platform for indexing company knowledge for AI agents, has introduced "Company Knowledge Bench," a new evaluation framework designed to address the limitations of existing public benchmarks. According to the company, most current benchmarks rely on synthetic documents or narrow domains like law or medicine, failing to capture the diverse and messy nature of real-world enterprise queries.
Product LaunchesKapaCompany Knowledge Bench
Kapa Introduces Company Knowledge Bench to Evaluate Agent Retrieval on Real-World Data
The benchmark consists of 1,000 evaluation cases annotated from real production data. To ensure scalability and accuracy, Kapa utilized agents to generate the labels, validating their performance against a human-labeled set of 170 cases. The benchmark evaluates retrievers based on three key properties: completeness, minimality, and source preference, the latter of which prioritizes authoritative and current documents.
Kapa tested seven different retrieval strategies on the benchmark, ranging from traditional hybrid search to advanced agentic retrievers. The results highlighted a significant trade-off between accuracy, latency, and cost. While an optimized agentic retriever, Kapa Deep, achieved the highest score of 0.65 in approximately five seconds, simpler "agentic grep" methods using large language models were found to be significantly more expensive and slower, often returning excessive tokens.
The findings suggest that as enterprises deploy more agents, optimizing the balance between retrieval accuracy and the operational costs of latency and token usage remains a critical challenge.
Sources
- Benchmarking retrieval for agents on messy real-world company knowledge (Hacker News Frontpage, 2026-10-02)