Qwen 3.8 27B launches on Cerebras at roughly 1,500 tokens per second
Cerebras has added Alibaba's Qwen 3.8 27B to its hosted inference catalog, offering the model at roughly 1,500 tokens per second, according to Cerebras' own model documentation published around September 3, 2026. The listing makes Qwen 3.8 27B available through Cerebras' wafer-scale inference infrastructure rather than only through Alibaba's own hosting.
What's new
Cerebras describes the model as "Alibaba's 27B dense multimodal model for agentic coding, tool use, research, and long-running workflows." On the Cerebras platform it ships with:
- Throughput: approximately 1,500 tokens per second, well above typical GPU-hosted inference speeds for a model this size.
- Context length: 64K tokens on the free trial tier, extending to 128K tokens on paid tiers.
- Pricing: $0.99 per million input tokens and $1.49 per million output tokens.
- Availability: both the free-trial and pay-as-you-go tiers, subject to each tier's own rate limits.
Context
Qwen 3.8 27B is a dense, natively multimodal model from Alibaba's Qwen team, positioned for agentic coding and tool-use workloads rather than as a general chat model. Cerebras has built its business around hosting open-weight models on custom wafer-scale chips that trade GPU flexibility for much higher raw token throughput, and it has previously added other open models — including earlier Qwen releases — to the same catalog shortly after their public release. Listing a 27B-parameter model at 1,500 tokens/second is consistent with the throughput advantage Cerebras has marketed for models in this size range.
Why it matters
Speed is a direct product feature for agentic and coding workloads, where a model iterates through many tool calls and long-running tasks — a 1,500 tokens/second inference speed can turn a multi-minute agent loop into a near-instant one. By hosting Qwen 3.8 27B directly, Cerebras extends its role as a neutral high-speed inference layer for open-weight models rather than a lab building its own frontier model, giving developers a faster, drop-in alternative to running the same open weights on standard GPU infrastructure.
Corroborating sources
- Inference-docs.cerebras
https://inference-docs.cerebras.ai/models/qwen-3.8-27b
“Alibaba's 27B dense multimodal model for agentic coding, tool use, research, and long-running workflows.”