Google Developer Platform
MaxText Expands Post-Training Capabilities: Introducing SFT and RL on Single-Host TPUs- Google Developers Blog

MaxText Expands Post-Training Capabilities: Introducing SFT and RL on Single-Host TPUs- Google Developers Blog

4/16/2026 · Wei Wei, Weiren Yu

What this post added

Introduces Supervised Fine-Tuning (SFT) and Reinforcement Learning (RL) capabilities within the MaxText framework, specifically for single-host TPU configurations (e.g., v5p-8, v6e-8). Highlights seamless integration with Hugging Face datasets and checkpoints, optimized execution via the Tunix library, and the use of vLLM for high-throughput inference in RL algorithms like GRPO and GSPO. Provides installation instructions and command-line examples for running SFT and RL.

Read the original post ↗