Agentic AI Infrastructure Acceleration with BlueField DPUs
NVIDIA Nemotron 3 Ultra Powers Faster, More Efficient Reasoning for Long-Running Agents | NVIDIA Technical Blog

NVIDIA Nemotron 3 Ultra Powers Faster, More Efficient Reasoning for Long-Running Agents | NVIDIA Technical Blog

6/4/2026

What this post added

Introduces NVIDIA Nemotron 3 Ultra, a 550B-parameter Mixture-of-Experts model designed for agent orchestration and long-running agent workflows. Highlights architectural innovations including hybrid Mamba-Transformer layers, NVFP4 quantization for cross-architecture GPU deployment (achieving up to 5x higher throughput), LatentMoE for efficient expert routing, and multi-token prediction for improved generative speed. Details the Multi-Teacher On-Policy Distillation (MOPD) training method for continuous improvement and domain specialization, and outlines expanded training data including domain-specific pre-training data (synthetic legal, Wiki-based, refreshed GitHub tokens) and post-training data (SFT samples, RL tasks, RL environments).

Read the original post ↗