Google Developer Platform
How the community trained Gemma to "Think" with Tunix and TPUs- Google Developers Blog

How the community trained Gemma to "Think" with Tunix and TPUs- Google Developers Blog

5/28/2026 · Wei Wei, Weiren Yu, Tianshu Bao, Lance Wang, Chris Achard

What this post added

This post details the Google Tunix Hackathon, which challenged developers to transform non-reasoning base models (Gemma-2-2B and Gemma-3-1B) into general reasoning models using Tunix and Kaggle TPUs. It highlights winning techniques such as Supervised Fine-Tuning (SFT), Rubric-Based Reinforcement Learning (GRPO) with LLM-as-judge reward systems, SimPO for strict formatting, and TF-IDF rewards. The post also showcases domain-specific reasoning training for medical, chemistry, legal, and robotics applications, and provides resources for developers to start post-training their own reasoning models.

Read the original post ↗