
8/11/2026
What this post added
Introduces NVIDIA Nemotron 3.5 Lightning, a 30B parameter MoE model with 3B active parameters, optimized for high-volume, low-latency execution in always-on AI agents. Details its features like speculative decoding, harness-optimized training, and quantization (NVFP4 and BF16 checkpoints) for strong accuracy and up to 4x output speed. Also introduces NVIDIA NeMo Switchyard for intelligent model routing to optimize task allocation across different models. The post highlights Nemotron 3.5 Lightning's performance on the Artificial Analysis Intelligence Index and PinchBench, demonstrating its efficiency for agentic tasks.