
7/15/2026
What this post added
This post details the integration of the Inkling multimodal model on Modal, highlighting its architecture (mixture-of-experts, local attention) and the performance gains achieved through a custom DFlash speculator tuned for this model shape. It specifically details how DFlash was adapted with all-local attention and causal layers for better kernel support to achieve 250 tokens per second per user on agentic workloads. It also announces Inkling's availability as a Managed Endpoint.