Speculative Decoding for LLM Inference
Inkling by Thinking Machines now available on Modal | Modal Blog

Inkling by Thinking Machines now available on Modal | Modal Blog

7/15/2026

What this post added

This post details the integration of the Inkling multimodal model on Modal, highlighting its architecture (mixture-of-experts, local attention) and the performance gains achieved through a custom DFlash speculator tuned for this model shape. It specifically details how DFlash was adapted with all-local attention and causal layers for better kernel support to achieve 250 tokens per second per user on agentic workloads. It also announces Inkling's availability as a Managed Endpoint.

Read the original post ↗