Inference Performance Optimization
Dollars per token considered harmful

Dollars per token considered harmful

7/16/2025

What this post added

This post introduces a new perspective on cost and performance optimization for LLM inference, advocating for a 'dollars per request' model instead of 'dollars per token'. It explains why this shift is crucial for teams self-hosting inference, impacting latency estimation, replica scaling, and overall cost analysis. The post highlights how this framing aligns engineering efforts with product and business goals, and suggests Modal as a platform to achieve lower costs per request.

Read the original post ↗