
How we used evals and inference-time compute scaling to generate beautiful QR codes that actually work
7/2/2025
This post details the application of 'evals' and 'inference-time compute scaling' to improve the quality of generative AI applications, specifically focusing on the creation of scannable and aesthetically pleasing QR codes (QArt codes). It outlines the process of defining system goals (scannability and aesthetic quality), operationalizing these goals into measurable metrics, developing automated evaluation procedures using libraries like QReader and aesthetic rating predictors, and iteratively aligning these automated evals with human judgment through experimentation. The post also highlights the use of inference-time compute scaling to achieve performance targets (under 20s p95) and meet service-level objectives (95% scan rate).