
2/2/2026
What this post added
This post extends the LLM Evaluation Framework by adding support for proprietary model providers (OpenAI, Anthropic, Google) as both judge and target models. It also enables evaluation of Together fine-tuned models (LoRA serverless Inference, Dedicated Endpoints) and provides new recipes for optimizing and evaluating open-source models, including fine-tuning open models to outperform proprietary judges and using GEPA for automated prompt optimization.