.png)
4/29/2026
What this post added
This post announces the availability of DeepSeek V4 Pro on Together AI, highlighting its 512K-token context window (model-level 1M) and 1.6T-parameter MoE architecture. It details the hybrid attention mechanisms (Compressed Sparse Attention and Heavily Compressed Attention) used to optimize serving for long contexts, reducing FLOPs and KV cache usage. The introduction of three controllable reasoning modes (Non-Think, Think High, Think Max) allows for task-specific reasoning depth. The post also emphasizes the practical benefits of cached input pricing for repeated long-context queries, offering a 90% cost reduction, and outlines the deployment path from serverless to dedicated infrastructure for production workloads.