
8/28/2025
What this post added
This post details how Zencastr migrated their GPU-intensive audio transcription workloads from a self-managed Kubernetes cluster to Modal. Key technical contributions include: 1. Demonstrating the cost-effectiveness of Modal's scale-to-zero for spiky AI workloads compared to always-on Kubernetes GPU nodes. 2. Highlighting Modal Images for seamless management of diverse ML model dependencies and CUDA driver versions, reducing infrastructure management overhead. 3. Showcasing the flexibility of Modal for experimenting with different GPU models and concurrency settings without code changes. 4. Illustrating a large-scale batch audio processing architecture using Modal Functions, S3 Gateway endpoints, and Modal Volumes for efficient data handling and parallel GPU utilization (scaling to 1,500 concurrent GPUs).