Google Developer Platform
TorchTPU: Running PyTorch Natively on TPUs at Google Scale- Google Developers Blog

TorchTPU: Running PyTorch Natively on TPUs at Google Scale- Google Developers Blog

4/7/2026 · Claudio Basile, Kat Ko, Ben Wilson, Lee Howes, Bill Jia, Joe Pamer, Michael Voznesensky, Robert Hundt

What this post added

Introduced TorchTPU, a new engineering effort to enable PyTorch to run natively and efficiently on Google's Tensor Processing Units (TPUs). This includes an 'Eager First' philosophy using PyTorch's 'PrivateUse1' interface with Debug, Strict, and Fused eager modes. For static compilation, it leverages Torch Dynamo, XLA, and StableHLO, with support for custom kernels via Pallas and JAX. TorchTPU also enhances distributed training support for DDP, FSDPv2, and DTensor, and addresses the MPMD challenge for divergent executions. Hardware awareness guidelines are provided for optimal TPU utilization.

Read the original post ↗