Alibaba's omni-modal understanding model, released September 18, 2026 — it reads text, images, audio, and video together in one request, holds a one-million-token context, and is tuned for cheap, agentic comprehension of long recordings.
Last verified · 2026-09-18 · by Moe Ameen
Qwen3.8-Omni-Flash is an omni-modal understanding model from Alibaba's Qwen team, released on September 18, 2026. In a single request it takes text, images, audio, and video together, reasons over all of it, and returns text — and it carries a native one-million-token context window, long enough to read a full recording, a set of screenshots, and a brief in one pass. Qwen describes it as its first Omni model built around agentic capabilities, meaning it is designed not just to comprehend media but to call tools and act on what it finds.
It is important to get the category right, because the name invites confusion. This is not a video generator like Google's Gemini Omni Flash. Qwen3.8-Omni-Flash is on the intake side of the pipeline: it listens and watches and describes, rather than rendering anything. It returns text, and for generated speech Qwen points developers to its separate Qwen3.5-Omni model.
The pitch is cheap comprehension at scale. Qwen reports the model improves on the earlier Qwen3.5-Omni-Plus by roughly 26% across about 30 audio and video tests, lifts the OmniVideoBench score from 63.4 to 67.8 while cutting token consumption around 45.7%, and drops audio input costs by over 98% and audio-visual input by over 93% versus that prior model — with hosted API rates around $0.15 per million input tokens and $0.47 per million output. It understands text across many languages and audio across a wide range of languages and dialects.
Access is through the API on QwenCloud, Alibaba Cloud Model Studio, and Qwen Studio; at launch the weights were not open-sourced, a departure from the open-weight Qwen3.8-27B and Flash-Next releases. As with any fresh model, treat the benchmark and pricing figures as an early, largely vendor-reported snapshot and confirm them against Qwen's own materials. One thing is not in flux: it comprehends media and reasons over it, but it renders no video, images, or audio — it outputs text.
Omni-Flash is the eyes and ears at the front of a repurposing workflow — and that is exactly the half most solo creators skip because it used to be slow and expensive. Record one long thing: a podcast episode, a client webinar, a Q&A. Omni-Flash reads the entire recording in one pass and hands you a transcript, the quotable moments, and where the good clips start. But a set of notes and timestamps is not content. Nothing is filmed anew, designed, branded, or posted. Turning that understanding into finished, on-brand posts across platforms is a separate job, and it is the one [Kompozy](/) does.
Here is the concrete loop. Let Omni-Flash comprehend the recording and surface the moments, then drop that source into Kompozy and pick your formats. From that single input Kompozy generates the assets a text model can't: [Clipped Shorts](/glossary/clipped-short) cut straight from the long-form, a captioned [Persona Short](/glossary/persona-shorts) fronted by a face-locked HeyGen avatar, brand-exact [Carousel Posts](/glossary/hyperframes), quote graphics pulled from the strongest lines, photo posts, a full blog article, and an email newsletter — each rewritten under a [Persona Brief](/glossary/persona-brief) so the voice reads as yours. [Autopilot](/glossary/autopilot) then schedules and publishes the batch across the eight social platforms plus blog and email, every asset clearing a per-post review gate first. Because Omni-Flash makes the comprehension step so cheap, you can record more and let the engine handle the parts it can't touch — the visuals, the brand identity, and the distribution. On the Founding tier you can bring your own Qwen key so the understanding layer runs at cost inside Kompozy.
Qwen3.8-Omni-Flash is Alibaba Qwen team's omni-modal understanding model, released September 18, 2026. It takes text, images, audio, and video in a single request, carries a one-million-token context, and is built around comprehending audio and video and using tools agentically. It returns text — for generated speech Qwen points to its separate Qwen3.5-Omni model.
No. Despite the similar name, Qwen3.8-Omni-Flash reads audio and video and returns text — it is an understanding model. Google's Gemini Omni Flash generates and edits video. They do opposite jobs: one comprehends clips, the other renders them.
No. It reasons over multimodal input but returns text; it renders no images, video, or audio. To turn what it surfaces from a recording into finished visual posts, you pair it with a generation-and-publishing engine like Kompozy.
Hosted API rates are around $0.15 per million input tokens and $0.47 per million output, and Qwen reports audio input costs cut over 98% and audio-visual over 93% versus Qwen3.5-Omni-Plus. Access is via QwenCloud, Alibaba Cloud Model Studio, and Qwen Studio; weights were not open-sourced at launch. Confirm current figures on Qwen's site.
Use Omni-Flash to comprehend a long recording and surface the quotes and clip moments, then bring the source into Kompozy. Kompozy generates 18 formats from that one input — Clipped Shorts, persona/avatar video, carousels, quote graphics, photo posts, a blog, and a newsletter — holds a consistent face and voice, and schedules and publishes across the eight social platforms plus blog and email.