Inference Performance Optimization
Introducing Modal Auto Endpoints: Optimized inference you actually own | Modal Blog

Introducing Modal Auto Endpoints: Optimized inference you actually own | Modal Blog

6/23/2026

What this post added

Introduces Modal Auto Endpoints, a self-serve system for deploying production-grade LLM inference services. This feature provides users with transparent access to the underlying code, inference engine configurations (e.g., DFlash, SGLang, FlashAttention-4), and detailed engine-level observability metrics (TTFT, ITL, speculative decoding acceptance length). It leverages Modal's existing AI infrastructure, including serverless GPU inference, autoscaling, and Modal Servers for low-latency routing. The system is built on recipes derived from internal optimization efforts and offers a declarative interface based on workloads and SLOs, with an evolving agentic system for generating deployment configurations.

Read the original post ↗