
7/15/2026
What this post added
This post announces the availability of the Inkling multimodal model on Together AI's inference platform. It details Inkling's architecture, including query-conditioned relative attention, short causal convolutions, and a shared-sink MoE, and highlights its multimodal input capabilities (text, image, audio) and controllable reasoning effort. The post also emphasizes the optimization of Inkling for production inference on Together AI, specifically mentioning the use of a FlashAttention-4-based kernel to support its attention mechanism.