
12/22/2025
What this post added
Introduced new model snapshots for speech-to-text, text-to-speech, and real-time speech-to-speech, focusing on improving reliability and quality for production voice workflows. Key improvements include lower word-error rates in noisy audio, fewer hallucinations during silence or background noise, and more natural and stable voice output. Specifically, the `gpt-realtime-mini` model shows significant gains in instruction-following accuracy (18.6 pp) and tool-calling accuracy (12.9 pp) for real-time agents. Text-to-speech models achieve roughly 35% lower WER on benchmarks like Common Voice and FLEURS. Speech-to-text models produce ~90% fewer hallucinations compared to Whisper v2 in hallucination-with-noise evaluations. Custom Voices also benefit from more natural tones and increased faithfulness.