Together AI launches preemptible GPU compute at half the on-demand price
Together AI introduced preemptible compute for its Together GPU Clusters on September 10, a new capacity tier priced at half the standard on-demand rate for workloads that can tolerate being interrupted.
What's new
Together's blog post lays out the mechanics of the new tier:
- Pricing is set at "the same GPU capacity at a flat 50% of the on-demand rate," with no separate discount tiers to negotiate.
- Nodes can be reclaimed with "a five-minute drain window," after which Together automatically works to refill capacity back toward the customer's target.
- Billing is granular: "sub-hourly, with usage metered every one to two minutes," so customers are not charged for a full hour when a job runs for a few minutes.
- The tier is in public preview now, available across all regions, and can be added directly to existing Kubernetes clusters alongside on-demand nodes.
- Together frames the target workloads as "short experiments, inference bursts, batch jobs," and other temporary capacity bursts — cases where a job can checkpoint, retry, or simply tolerate a mid-run interruption.
Context
Preemptible or "spot" compute is a well-established pattern in general cloud infrastructure — AWS Spot Instances and Google Cloud Preemptible VMs have offered similar discounts for interruptible workloads for years. Together AI's move brings that same model to GPU clusters aimed specifically at AI training and inference, where GPU capacity is scarcer and more expensive than general-purpose compute, and where demand has been running well ahead of supply across the industry.
Together has built its business around undercutting the major hyperscalers on GPU rental price while offering an increasingly full-featured platform — clusters, inference endpoints, fine-tuning, and now a formal spot-pricing tier. Competing neoclouds and inference platforms (CoreWeave, Lambda, Fireworks, and others) have leaned on similar interruptible-capacity offers as a way to monetize otherwise-idle GPU fleets.
Why it matters
A flat 50% discount with a five-minute drain window and sub-hourly billing is a meaningfully more usable spot offering than the coarser hourly discounts common elsewhere in GPU rental markets — the short drain window and granular billing reduce the operational pain of adopting preemptible capacity for teams running short training jobs, batch inference, or bursty workloads.
For AI teams under budget pressure, cheaper GPU access for anything that can be checkpointed or retried is a direct cost lever, and the timing — announced as demand for inference and fine-tuning capacity remains tight industry-wide — suggests Together is trying to capture price-sensitive workloads that might otherwise go to a hyperscaler's own discounted tiers or sit idle waiting for on-demand capacity to free up. It also puts pressure on other GPU-cloud competitors to match on price or drain-window terms, particularly as more workloads shift toward short, bursty inference rather than long-running training runs.
Corroborating sources
- Together
https://www.together.ai/blog/introducing-preemptible-compute-the-same-compute-half-the-price
“the same GPU capacity at a flat 50% of the on-demand rate”