Dedicated Container Inference
Introducing Dedicated Container Inference:  Delivering 2.6x faster inference for custom AI models

Introducing Dedicated Container Inference: Delivering 2.6x faster inference for custom AI models

2/12/2026

What this post added

Introduces Dedicated Container Inference, a new capability for deploying custom generative media models. This feature provides production-grade orchestration including autoscaling, queuing, traffic isolation, and monitoring for user-provided Docker containers. It supports job orchestration with independent queues, policy-driven traffic control, and isolation between different traffic types. The architecture treats containers as the unit of execution, utilizes volume mounts for model weights, and offers autoscaling based on queue depth or custom metrics. Observability is built-in.

Read the original post ↗