Together AI launches Provisioned Throughput, reserved inference capacity with a 99% uptime SLA
Together AI introduced Provisioned Throughput on July 8, 2026, a reserved-capacity purchasing option aimed at teams running open-weight models in production who need guaranteed performance rather than shared, on-demand inference.
What's new
Together AI describes the offering as "reserved inference capacity for frontier open models with token-based pricing and a 99% uptime SLA." The core pitch is that customers buy dedicated capacity units — Provisioned Throughput Units, or PTUs — rather than competing for shared inference slots, so performance stays consistent "regardless of traffic patterns."
Key details from the announcement:
- Pricing structure: capacity is billed at "$0.05 per PTU per minute." Together AI estimates that at full utilization on MiniMax M3, this works out to roughly "$0.36 per million input tokens and $2.16 per million output tokens" — which the company frames as "up to 90% lower cost" than comparable proprietary API options.
- Reliability guarantee: the plan carries a "99% uptime SLA," with capacity reserved in advance so throughput doesn't degrade under load.
- Model coverage: available at launch for MiniMax M3 and GLM-5.2 — both open-weight models, consistent with Together AI's positioning as open-source-first inference infrastructure.
- Terms: capacity is offered "in North America, EMEA, and beyond," with a one-month minimum term and discounts for customers who commit to higher capacity levels.
Context
The launch comes about a week after Together AI closed an $800 million Series C explicitly tied to expanding compute commitments to accelerate open-source AI inference. Provisioned Throughput is a direct product expression of that strategy: converting fresh compute capacity into a purchasable, SLA-backed tier rather than only expanding shared on-demand capacity. It also follows Together AI's research push at ICML 2026 on inference optimization, an area the company has repeatedly tied to its commercial roadmap.
Why it matters
Reserved-capacity pricing with an uptime SLA is standard at the hyperscalers (AWS, Azure, GCP all sell some version of committed-capacity compute), but it has been rarer among open-model inference providers, who have mostly competed on raw per-token price rather than guaranteed throughput. By adding an SLA-backed tier on top of already-aggressive token pricing, Together AI is targeting a specific gap: production teams that have already chosen an open-weight model for cost reasons but have been reluctant to commit because shared inference capacity can't promise consistent latency under load. If the economics hold up under real production traffic, it strengthens the case that open-weight models plus specialized inference providers can match not just the price but also the reliability guarantees of closed-model APIs — a claim that's been central to the broader open-source-versus-proprietary argument playing out across the industry this year.
Corroborating sources
- Together
https://www.together.ai/blog/provisioned-throughput
“Reserved inference capacity for frontier open models with token-based pricing and a 99% uptime SLA.”