Academic Publication Search
How Search Quality Shapes RL Outcomes

How Search Quality Shapes RL Outcomes

5/13/2026 · Exa Labs

What this post added

This post details an experiment comparing Exa's search backend with Google's SERP for training RL agents. It highlights how using Exa as the retrieval backend during RL training leads to higher performance with less training compute. The experiment involved training Qwen3-4B-Instruct-2507 models with LoRA adapters, using a consistent system prompt and tool configuration (5 live web results, 2000 character snippets). The analysis focuses on the impact of the retrieval backend on agent performance, token efficiency, and the trade-offs between search engine performance during training versus inference.

Read the original post ↗