
Introducing Evals
8/4/2026
Introduced Evals, a new system for testing and measuring Hex Agent performance. Evals leverage an LLM-as-judge to evaluate agent reasoning and tool usage, not just final answers. The system is integrated with the Hex CLI and supports version-controlled test cases and context previews for safe testing of changes. Evals enable scheduled checks for regressions and configuration sweeps across different LLMs.





