
7/20/2026
What this post added
This post details the technical collaboration with Heidi Health, focusing on achieving superior frontier model performance. Key contributions include: 1. Demonstrating Supervised Fine-Tuning (SFT) to match foundation model behavior and Reinforcement Fine-Tuning (RFT) for 'deep thinking' by incorporating preference signals. 2. Highlighting the critical role of data quality through synthetic rewrites and LLM-as-a-Judge filtering for de-noising preference data. 3. Emphasizing the necessity of scaling effective batch sizes to over 1.5 million tokens using gradient accumulation to stabilize training and improve win rates against proprietary models. The partnership resulted in a 3.5x reduction in latency.