8/11/2026
What this post added
This post introduces the Whisper Large V3 model to the Baseten platform, highlighting its implementation achieving 1800x real-time factor for audio transcription. It details recommended hardware configurations (H100 MIG, H100, L4) and concurrency targets for different use cases (balanced, latency-sensitive, cost-sensitive). Example Python code for API usage and expected JSON output for transcription results are provided.