.png)
3/11/2026
What this post added
This post announces the availability of NVIDIA Nemotron 3 Super on Together AI's Dedicated Inference platform. It details the model's hybrid MoE architecture (Transformer + Mamba), 1M-token context window, and multi-token prediction capabilities, highlighting their benefits for agentic workflows and complex reasoning. The post also explains how Nemotron 3 Super is optimized for single-GPU deployment on H200/H100 GPUs within Together AI's managed infrastructure, emphasizing the use of the Together Inference Engine and custom CUDA kernels for accelerated performance and production-grade isolation with an SLA and SOC 2 compliance.