Alibaba releases Qwen3.8-Omni-Flash, cutting audio API pricing by more than 98%
Alibaba's Qwen team launched Qwen3.8-Omni-Flash on September 14, 2026, a native omnimodal model the company says pushes omnimodal AI beyond "understanding" content toward planning tasks, calling tools, and doing creative work autonomously.
What's new
Qwen3.8-Omni-Flash takes text, image, audio, and video input with a 1-million-token context window while holding text performance comparable to a text-only model of the same size. Alibaba says its average score across 29 evaluations improved more than 25% over the prior Qwen3.5-Omni-Plus model, driven by gains including 36.5 points on WildClawBench-MM, 22.3 points on AgenticVBench, and a 69.6 score on UniClawBench, plus double-digit gains on long-form audio and audio-visual understanding benchmarks like LongAudioSpan and OmniVideoBench.
The headline change is pricing: audio-input API costs drop by more than 98%, and audio-visual input costs drop by more than 93%, compared with the prior generation. Alibaba positions the result as matching Gemini 3.8 Flash on audio-visual performance while exceeding it on overall audio performance. The model is live now on the Qianwen AI Platform and through dedicated Qwen3.8-Omni-Flash and Qwen3.8-Omni-Flash-Realtime APIs on Alibaba Cloud Model Studio, aimed at workflows like video editing, music-video creation, film production and commentary, audio-visual summarization, and real-time conversation.
Context
The release lands the same week Alibaba previewed Qwen3.8-Flash-Next, a separate experimental architecture the company has described as an early look at Qwen4 — a different model built for general reasoning and coding rather than omnimodal agent work. Omni-Flash instead extends the agentic capabilities Qwen has been building into its coding and GUI-operation models, now centered specifically on audio and video as inputs an agent can act on rather than just interpret.
Why it matters
The pricing cut is the more consequential number here: audio and audio-visual input have been an expensive tier for any product built on omnimodal APIs, and a 93–98% cost drop from a frontier-adjacent lab puts real pressure on Google and OpenAI's own audio/video API pricing, particularly for use cases like real-time conversation and video summarization at scale. Framing the release around agentic use of audio and video, rather than just multimodal understanding, also tracks the broader industry shift from models that describe media to models that act on it.
Corroborating sources
- Neowin.net
https://www.neowin.net/news/alibabas-qwen38-omni-flash-undercuts-gemini-on-audio/
- Qwen
https://qwen.ai/blog?id=qwen3.8-omni-flash
“Qwen3.8-Omni-Flash supports a 1M-token context window while maintaining text performance comparable to a text-only model of the same size and delivering significant improvements in omnimodal capabilities.”