
6/8/2026
What this post added
This post introduces the NVFP4 training recipe for JAX, implemented in MaxText, which enables high-throughput, 4-bit mixed-precision pre-training on NVIDIA Blackwell and Rubin platforms. It details the five core techniques used to preserve convergence with negligible accuracy loss: micro block scaling, E4M3 block scale factors, selective Random Hadamard Transform for WGRAD inputs, 2D FP8 scaling per 16x16 weight block, and stochastic rounding for unbiased quantization. The post also provides guidance on enabling NVFP4 in MaxText and presents performance results showing up to 1.73x speedup over FP8 baselines on GB300 hardware.