Cerebras partners with Upstage to bring 2,000-token-per-second inference to South Korea
Cerebras announced a partnership with Upstage, one of South Korea's leading enterprise AI companies, to run Upstage's Solar language models on Cerebras' wafer-scale inference infrastructure — targeting inference speeds Cerebras says standard GPU deployments can't match.
What's new
The deal puts Upstage's Solar 31B model on Cerebras Inference Cloud, which runs on Cerebras' Wafer-Scale Engine hardware rather than conventional GPUs. According to Cerebras, Solar 31B running on its infrastructure reaches inference speeds of up to 2,000 tokens per second, exposed through OpenAI-compatible APIs so existing developer tooling can point at it with minimal integration work.
Cerebras illustrates the practical gap with a benchmark comparison: Solar 31B completed a deep-research query pulling from 246 sources in the time Claude Sonnet 4.6 needed roughly eight minutes for a comparable task — Cerebras' framing for why raw inference speed changes what's practical to build, not just how a benchmark scores. As Cerebras puts it in the announcement: "For many AI applications, faster inference is not just a performance benchmark. It changes what a product can do."
Upstage CEO Sung Kim said: "We're excited to work with Cerebras to bring our Solar models to developers with industry-leading inference speeds." Cerebras CMO Julie Choi added: "South Korea is at the forefront of AI innovation, and we're thrilled to partner with Upstage to help developers."
Context
Upstage has built out an enterprise AI portfolio spanning large language models, document intelligence, and workflow automation, with its Solar model family aimed at enterprise-grade language tasks emphasizing speed and groundedness over raw scale. Cerebras has spent 2026 expanding its inference-partner roster beyond its own chips-and-cloud business — it struck a similar speed-focused deployment with Hugging Face on Gemma 4 in early July and separately announced plans for 200MW of AI compute capacity in Europe by the end of 2027. Pairing with a regional model vendor like Upstage, rather than only serving frontier-lab models, extends Cerebras' inference business into markets where a locally-built model already has enterprise traction.
Why it matters
Inference speed has become a competitive axis in its own right as agentic and multi-step AI workflows multiply the number of model calls a single task requires — an 8-minute research task that drops to seconds changes what's viable to ship as a real-time product feature rather than a background job. For Cerebras, landing Upstage gives it a foothold in the South Korean enterprise AI market through a model vendor that already has local customer relationships, rather than trying to sell wafer-scale inference cold. For Upstage, an alliance with a specialized inference vendor is a way to compete on speed against both GPU-cloud incumbents and the frontier labs' own hosted APIs without having to build custom silicon itself.
Corroborating sources
- Cerebras
https://www.cerebras.ai/blog/cerebras-and-upstage-bring-ultra-fast-ai-to-korea
“Running on the Cerebras Wafer-Scale Engine, Upstage's Solar 31B can achieve inference speeds of up to 2,000 tokens per second.”