Model Inference Platform
Serverless 2.0: Three Ways to Run Inference, One API

Serverless 2.0: Three Ways to Run Inference, One API

5/26/2026

What this post added

Introduced Serverless 2.0 with three distinct serving paths: Standard (default, cost-efficient shared fleet), Priority (enhanced admission during congestion, sheds last), and Fast (high-throughput path for faster token generation). Clarified error codes to differentiate between account rate limits (429) and shared fleet overload (503), enabling better error handling and retry strategies. Introduced new model router IDs for Fast models (e.g., `accounts/fireworks/routers/kimi-k2p6-turbo`).

Read the original post ↗