
7/6/2026
What this post added
Introduces Nonuniform Tensor Parallelism (NTP) as an experimental framework to enhance Goodput in large-scale LLM training. NTP dynamically adjusts the tensor parallelism degree in response to transient GPU unavailability, preventing training stalls and throughput loss. It also proposes dynamic power boosting to compensate for performance loss in affected scale-up domains and employs efficient, overlapped tensor resharding techniques that minimize overhead to less than 1%. This enables resilient operation even as scale-up domains grow to 72 GPUs on NVIDIA Blackwell and Blackwell Ultra systems.