BlogsModalGenerative AI Evals and Compute Scaling

Generative AI Evals and Compute Scaling

Generative AI Evals and Compute Scaling

1
posts
2025

Modal's platform now supports advanced techniques for engineering generative AI applications, including robust evaluation frameworks ('evals') and dynamic inference-time compute scaling. This enables developers to systematically improve the quality and performance of AI-driven systems, such as generating aesthetically pleasing and functional QR codes. The platform facilitates the development lifecycle from initial demos to production-ready systems by providing tools for objective measurement, alignment with human judgment, and scalable execution.

2025

How we used evals and inference-time compute scaling to generate beautiful QR codes that actually work

7/2/2025

This post details the application of 'evals' and 'inference-time compute scaling' to improve the quality of generative AI applications, specifically focusing on the creation of scannable and aesthetically pleasing QR codes (QArt codes). It outlines the process of defining system goals (scannability and aesthetic quality), operationalizing these goals into measurable metrics, developing automated evaluation procedures using libraries like QReader and aesthetic rating predictors, and iteratively aligning these automated evals with human judgment through experimentation. The post also highlights the use of inference-time compute scaling to achieve performance targets (under 20s p95) and meet service-level objectives (95% scan rate).