
5/28/2026 · Wei Wei, Weiren Yu, Tianshu Bao, Lance Wang, Chris Achard
What this post added
This post details the Google Tunix Hackathon, which challenged developers to transform non-reasoning base models (Gemma-2-2B and Gemma-3-1B) into general reasoning models using Tunix and Kaggle TPUs. It highlights winning techniques such as Supervised Fine-Tuning (SFT), Rubric-Based Reinforcement Learning (GRPO) with LLM-as-judge reward systems, SimPO for strict formatting, and TF-IDF rewards. The post also showcases domain-specific reasoning training for medical, chemistry, legal, and robotics applications, and provides resources for developers to start post-training their own reasoning models.