Baseten AI Model Deployment and Serving Platform
Introducing NVIDIA Nemotron 3.5 ASR Streaming

Introducing NVIDIA Nemotron 3.5 ASR Streaming

8/4/2026

What this post added

This post introduces the deployment of NVIDIA Nemotron 3.5 ASR Streaming models on Baseten. It details the architecture of the models (cache-aware FastConformer-RNNT), their performance benchmarks on H100 GPUs (latency, concurrency over WebSocket and gRPC), and accuracy metrics (WER) for both English and multilingual variants. It also highlights the use of NVIDIA NIM for optimized inference and mentions the possibility of fine-tuning via Baseten Training.

Read the original post ↗