Video Generation Models
Together AI brings Thinking Machines Lab’s new model Inkling on day 0

Together AI brings Thinking Machines Lab’s new model Inkling on day 0

7/15/2026

What this post added

This post announces the availability of the Inkling multimodal model on Together AI's inference platform. It details Inkling's architecture, including query-conditioned relative attention, short causal convolutions, and a shared-sink MoE, and highlights its multimodal input capabilities (text, image, audio) and controllable reasoning effort. The post also emphasizes the optimization of Inkling for production inference on Together AI, specifically mentioning the use of a FlashAttention-4-based kernel to support its attention mechanism.

Read the original post ↗