Baseten AI Model Deployment and Serving Platform
Fast, accurate retrieval with NVIDIA Nemotron 3 Embed

Fast, accurate retrieval with NVIDIA Nemotron 3 Embed

7/16/2026

What this post added

This post introduces the availability of NVIDIA Nemotron 3 Embed models (8B and 1B) on the Baseten platform, enhancing its capabilities for retrieval-augmented generation (RAG) and AI agent applications. It details the technical trade-offs between retrieval accuracy and indexing speed offered by the two model sizes. The post also highlights Baseten's support for fine-tuning these models using the Nemotron Embed fine-tuning recipe and deploying them via Truss for production workloads.

Read the original post ↗