
7/28/2025
What this post added
Introduces Together Evaluations, a new framework for benchmarking LLM response quality. This framework enables users to define custom benchmarks and use LLMs as judges to evaluate model performance across 'Classify', 'Score', and 'Compare' modes. It details data upload formats (JSONL, CSV), system template configuration, and model selection for evaluation. The post also highlights the integration with existing serverless inference APIs and provides links to documentation, UI, and tutorial notebooks.