Long Context Task Decomposition
DeepSeek-V4 Pro now available on Together AI

DeepSeek-V4 Pro now available on Together AI

4/29/2026

What this post added

This post announces the availability of DeepSeek V4 Pro on Together AI, highlighting its 512K-token context window (model-level 1M) and 1.6T-parameter MoE architecture. It details the hybrid attention mechanisms (Compressed Sparse Attention and Heavily Compressed Attention) used to optimize serving for long contexts, reducing FLOPs and KV cache usage. The introduction of three controllable reasoning modes (Non-Think, Think High, Think Max) allows for task-specific reasoning depth. The post also emphasizes the practical benefits of cached input pricing for repeated long-context queries, offering a 90% cost reduction, and outlines the deployment path from serverless to dedicated infrastructure for production workloads.

Read the original post ↗