
6/4/2026
What this post added
Introduces NVIDIA Nemotron 3 Ultra, a 550B-parameter Mixture-of-Experts model designed for agent orchestration and long-running agent workflows. Highlights architectural innovations including hybrid Mamba-Transformer layers, NVFP4 quantization for cross-architecture GPU deployment (achieving up to 5x higher throughput), LatentMoE for efficient expert routing, and multi-token prediction for improved generative speed. Details the Multi-Teacher On-Policy Distillation (MOPD) training method for continuous improvement and domain specialization, and outlines expanded training data including domain-specific pre-training data (synthetic legal, Wiki-based, refreshed GitHub tokens) and post-training data (SFT samples, RL tasks, RL environments).