The inference platform now offers a "Deployment Readiness Scorecard" to evaluate models based on task accuracy, cost per successful task, and inference latency. It highlights the "Agent Execution Tax" as a key metric, representing wasted inference due to malformed output or retries. The platform's serving layer contributes to reducing this tax through structured output consistency and predictable latency, enabling more reliable agent deployments. Specific model profiles (GLM-5, MiniMax M2.5, Kimi K2.5) are provided with performance data on accuracy, cost, and latency in agentic tasks.
2026
NVIDIA Nemotron 3 Ultra is live on Fireworks, day zero
6/4/2026
This post announces the day-zero availability of NVIDIA Nemotron 3 Ultra on the Fireworks inference platform. It highlights the platform's optimizations for agentic tasks, including proprietary FireAttention kernels for up to 4x higher throughput. It also details the platform's capabilities for post-training customization (SFT, DPO) and on-demand deployments, emphasizing a unified environment for training and inference.
Serverless 2.0: Three Ways to Run Inference, One API
5/26/2026
Introduced Serverless 2.0 with three distinct serving paths: Standard (default, cost-efficient shared fleet), Priority (enhanced admission during congestion, sheds last), and Fast (high-throughput path for faster token generation). Clarified error codes to differentiate between account rate limits (429) and shared fleet overload (503), enabling better error handling and retry strategies. Introduced new model router IDs for Fast models (e.g., `accounts/fireworks/routers/kimi-k2p6-turbo`).
Fireworks AI
5/20/2026
This post introduces the concept of the "Agent Execution Tax" and a "Deployment Readiness Scorecard" to evaluate LLMs for agentic AI. It details a benchmark of 720 browser agent runs across four LLMs, measuring structured output reliability and step efficiency. The analysis quantifies the "Agent Execution Tax" as the ratio of wasted inference to productive inference, demonstrating how malformed JSON output and retries significantly inflate costs and latency. It also provides a per-site analysis and model profiles (GLM-5, MiniMax M2.5, Kimi K2.5) based on these metrics, emphasizing the importance of reliable execution infrastructure in agent deployments.
Innovative Solutions Rebuilds Enterprise Services Delivery with Fireworks AI
5/4/2026
This post details how Innovative Solutions leveraged Fireworks AI's inference platform to rebuild their enterprise services delivery. Key technical contributions include: 1. Transformation of AI inference from a linear cost center to a predictable scaling layer for multi-agent services. 2. Enabling faster cycles and higher throughput by reducing model integration overhead and stabilizing multi-model execution. 3. Achieving predictable economics at billions of tokens per month, shifting from linear cost growth to controllable economics. 4. Facilitating a redesign of services economics around AI systems, moving from linear delivery models to parallel, agent-driven execution across sales, scoping, and delivery. 5. The platform's ability to handle constant model changes without operational friction (e.g., "works the first time. No tuning, no fiddling.") was critical for their multi-agent workflows. 6. Migration of 90% of Anthropic inference spend to Fireworks within 1-2 weeks of initial deployment, highlighting ease of integration and operational stability.