
7/23/2026
What this post added
This post introduces a major update to Together AI's inference platform, focusing on production-grade deployment capabilities for open-weight and licensed models. It details features for safe model rollouts (canary, blue-green, rolling updates with auto-rollback), traffic testing (A/B, shadow), and autoscaling based on various inference-native metrics. The platform now offers improved model caching for faster warm starts and an organization-level Prometheus endpoint for observability. Additionally, it announces a closed beta for custom training, including RL and SFT, with direct deployment of checkpoints to production inference endpoints.