LLM Evaluation Framework
Together Evaluations now supports comparing top commercial APIs vs. open source models

Together Evaluations now supports comparing top commercial APIs vs. open source models

2/2/2026

What this post added

This post extends the LLM Evaluation Framework by adding support for proprietary model providers (OpenAI, Anthropic, Google) as both judge and target models. It also enables evaluation of Together fine-tuned models (LoRA serverless Inference, Dedicated Endpoints) and provides new recipes for optimizing and evaluating open-source models, including fine-tuning open models to outperform proprietary judges and using GEPA for automated prompt optimization.

Read the original post ↗