Qwen shipped an omni-modal model that reads text, images, audio, and video in a single request, holds a one-million-token context, and leans hard on agentic tool use — while slashing the price of audio and audio-visual input over its previous Omni model.
2026-09-18 · by Moe Ameen
Alibaba's Qwen team released Qwen3.8-Omni-Flash on September 18, 2026, describing it as its first omni-modal model built around agentic capabilities. In one request the model takes text, images, audio, and video together and returns text, and it carries a one-million-token context window — long enough to reason over a full recording, a stack of screenshots, and a brief in a single pass. Qwen frames the release around two things: understanding audio and video well, and using tools to act on that understanding.
The efficiency story is the headline. Qwen reports Qwen3.8-Omni-Flash improves on the earlier Qwen3.5-Omni-Plus by roughly 26% across a suite of about 30 audio and video tests, and on the OmniVideoBench benchmark it moves from 63.4 to 67.8 while cutting token consumption by about 45.7%. The pricing moved just as sharply: Qwen says audio input costs are down more than 98% and audio-visual input costs down over 93% versus the prior Omni model, with hosted API rates around $0.15 per million input tokens and $0.47 per million output tokens.
The model is available through the API on QwenCloud, Alibaba Cloud Model Studio, and Qwen Studio. As of launch the weights were not open-sourced — access is API-only, which is a departure from the open-weight releases elsewhere in the Qwen3.8 line. Qwen also points developers who need generated speech to its Qwen3.5-Omni model, so treat Omni-Flash as an understanding-and-reasoning engine over multimodal input rather than a media generator. As with any fresh release, the benchmark and pricing figures are largely vendor-reported at launch — confirm them against Qwen's own materials before depending on a single number.
Read this as a price signal about the front of the content pipeline. The expensive, manual part of repurposing has always been getting raw media into a usable form — transcribing an hour-long interview, pulling the quotable moments, noting where the good clips start. Qwen3.8-Omni-Flash makes that step cheap and fast: it ingests the audio and video and hands back structured text you can actually work from. But understanding a recording is not the same as having posts. The model reads and reasons; it renders no video, no carousels, no images, and it publishes nothing. The distance from "I understand what's in this footage" to "on-brand posts live across every platform" is the whole job, and that is where [Kompozy](/) operates.
The practical loop is model-agnostic by design. Let Omni-Flash read the raw recording and surface the moments; then drop the source into Kompozy and let the engine do the part a text model can't. From one input Kompozy generates roughly 25–35 finished assets across 18 formats — a captioned [Persona Short](/glossary/persona-shorts) fronted by a face-locked HeyGen avatar, [Clipped Shorts](/glossary/clipped-short) cut from the long-form, brand-exact [carousels](/glossary/hyperframes), quote graphics, photo posts, a blog article, and an email newsletter — each governed by a [Persona Brief](/glossary/persona-brief) so the voice stays yours, routed through a per-post review gate, then scheduled and published by [Autopilot](/glossary/autopilot) across the eight social platforms plus blog and email. Cheaper multimodal understanding is a tailwind for that workflow, not a substitute for it — and on the Founding tier you can bring your own Qwen key so the understanding layer runs at cost inside the engine.
It is an omni-modal AI model from Alibaba's Qwen team, released September 18, 2026. It takes text, images, audio, and video in a single request, carries a one-million-token context window, and is built around understanding audio and video and using tools agentically. It returns text — for generated speech Qwen points developers to its Qwen3.5-Omni model.
No. Despite the similar name, Qwen3.8-Omni-Flash is an understanding model — it reads audio and video and returns text. Google's Gemini Omni Flash is a video generation model. Qwen's model is closer to a multimodal reasoning engine you feed media to, not a tool that renders clips.
Qwen reports audio input costs are down more than 98% and audio-visual input costs down over 93% versus the earlier Qwen3.5-Omni-Plus, with hosted API rates around $0.15 per million input tokens and $0.47 per million output tokens. Confirm current pricing on Qwen's own pages, since figures are fresh at release.
Not at launch. Access is API-only through QwenCloud, Alibaba Cloud Model Studio, and Qwen Studio — a departure from the open-weight releases elsewhere in the Qwen3.8 line. Check Qwen's materials for any later open-weight release.