
8/5/2025
What this post added
This post announces the availability of OpenAI's gpt-oss-120B and gpt-oss-20B models on Together AI's infrastructure. It details their availability as serverless and dedicated endpoints, emphasizing performance, economics, and reliability (99.9% uptime SLA). The post highlights developer tooling, including fine-tuning and OpenAI-compatible APIs. It mentions optimizations with NVIDIA, FlashAttention, and custom kernels. Pricing for input/output tokens is provided, along with mentions of the Batch API for cost savings. Real-world applications and a Python inference example are included.