8/4/2026
What this post added
This post introduces the deployment of NVIDIA Nemotron 3.5 ASR Streaming models on Baseten. It details the architecture of the models (cache-aware FastConformer-RNNT), their performance benchmarks on H100 GPUs (latency, concurrency over WebSocket and gRPC), and accuracy metrics (WER) for both English and multilingual variants. It also highlights the use of NVIDIA NIM for optimized inference and mentions the possibility of fine-tuning via Baseten Training.