
Agent and Model Evaluations in Gemini Enterprise Agent Platform are now GA- Google Developers Blog
7/31/2026
Introduces the general availability of agent and model evaluations within the Gemini Enterprise Agent Platform. This feature provides a unified engine for measuring and comparing agents and models using consistent metrics across development and production. Key components include over 20 pre-built metrics (quality, safety, grounding, tool use, reference-based scoring), adaptive rubrics that tailor judging criteria, and the ability to define custom code-based or LLM-as-a-judge metrics. Experiment management allows for local or server-side runs with auditable and reproducible artifacts stored in Cloud Storage. Online monitors integrate with existing telemetry to grade live production traffic, producing score-over-time charts and drift alerts. Case generation and simulation tools (user simulator, environment simulator) are provided to bootstrap evaluation datasets and test agent behavior under various conditions, including simulated backend failures.























