
3/23/2026 · Exa Labs
What this post added
This post introduces WebCode, a new set of coding evaluations for web search for coding agents. It details methods for evaluating extraction faithfulness against golden references, using both LLM-judged and deterministic metrics. It also introduces a discriminative evaluation framework for RAG to assess groundedness independently of synthesis, demonstrating its importance in isolating search provider capabilities. The post details the generation of question-answer pairs from long-context documentation and the evaluation of retrieval quality on the entire web, measuring groundedness and citation precision.