Real-time Low-Latency Inference Infrastructure
Learn how Cursor partnered with Together AI to deliver real-time, low-latency inference at scale

Learn how Cursor partnered with Together AI to deliver real-time, low-latency inference at scale

1/13/2026

What this post added

This post details the engineering efforts by Together AI in collaboration with Cursor to establish real-time, low-latency inference infrastructure on NVIDIA Blackwell. Key contributions include: - **Hardware Engineering:** Early rollout and reliable delivery of NVIDIA Blackwell GB200 NVL72 and HGX B200, including fast upgrades and replacements. - **System Optimization:** Full throughput achieved on ARM hosts by tuning the inference stack at the kernel and host levels. - **Custom Kernel Development:** Building kernels for Blackwell's new Tensor Core instructions to maximize hardware throughput. - **Parallelism Design:** Designing parallelism meshes for GB200 NVL72 to manage communication and synchronization overhead. - **Quantization Pipeline:** Implementing a quantization pipeline using NVIDIA TensorRT LLM and NVFP4 on Blackwell to balance compression with model quality for real-time serving. - **Production Deployment:** Establishing a repeatable path for moving trained models to production-like endpoints for testing and deployment, including A/B testing and cutovers.

Read the original post ↗