
6/23/2026
What this post added
Introduces Modal Auto Endpoints, a self-serve system for deploying production-grade LLM inference services. This feature provides users with transparent access to the underlying code, inference engine configurations (e.g., DFlash, SGLang, FlashAttention-4), and detailed engine-level observability metrics (TTFT, ITL, speculative decoding acceptance length). It leverages Modal's existing AI infrastructure, including serverless GPU inference, autoscaling, and Modal Servers for low-latency routing. The system is built on recipes derived from internal optimization efforts and offers a declarative interface based on workloads and SLOs, with an evolving agentic system for generating deployment configurations.