
2/23/2026
What this post added
This post details research into the failure modes of Automatic Speech Recognition (ASR) systems, specifically their inability to reliably transcribe street names for speakers with diverse linguistic backgrounds. It introduces new benchmarks (SF Streets and US Streets) and quantifies the performance gap, showing an average 39% transcription error rate for street names across 15 state-of-the-art models, with an 18% accuracy gap for non-English primary speakers. The post proposes and demonstrates a synthetic data generation technique using cross-lingual style transfer to improve accuracy by up to 60% with fewer than 1,000 training samples, highlighting a practical method for enhancing ASR robustness for critical applications.