// AI TOOLS · QWEN3.8-OMNI-FLASH

Qwen3.8-Omni-Flash

Alibaba's omni-modal understanding model, released September 18, 2026 — it reads text, images, audio, and video together in one request, holds a one-million-token context, and is tuned for cheap, agentic comprehension of long recordings.

Last verified · 2026-09-18 · by Moe Ameen

What Qwen3.8-Omni-Flash is

Qwen3.8-Omni-Flash is an omni-modal understanding model from Alibaba's Qwen team, released on September 18, 2026. In a single request it takes text, images, audio, and video together, reasons over all of it, and returns text — and it carries a native one-million-token context window, long enough to read a full recording, a set of screenshots, and a brief in one pass. Qwen describes it as its first Omni model built around agentic capabilities, meaning it is designed not just to comprehend media but to call tools and act on what it finds.

It is important to get the category right, because the name invites confusion. This is not a video generator like Google's Gemini Omni Flash. Qwen3.8-Omni-Flash is on the intake side of the pipeline: it listens and watches and describes, rather than rendering anything. It returns text, and for generated speech Qwen points developers to its separate Qwen3.5-Omni model.

The pitch is cheap comprehension at scale. Qwen reports the model improves on the earlier Qwen3.5-Omni-Plus by roughly 26% across about 30 audio and video tests, lifts the OmniVideoBench score from 63.4 to 67.8 while cutting token consumption around 45.7%, and drops audio input costs by over 98% and audio-visual input by over 93% versus that prior model — with hosted API rates around $0.15 per million input tokens and $0.47 per million output. It understands text across many languages and audio across a wide range of languages and dialects.

Access is through the API on QwenCloud, Alibaba Cloud Model Studio, and Qwen Studio; at launch the weights were not open-sourced, a departure from the open-weight Qwen3.8-27B and Flash-Next releases. As with any fresh model, treat the benchmark and pricing figures as an early, largely vendor-reported snapshot and confirm them against Qwen's own materials. One thing is not in flux: it comprehends media and reasons over it, but it renders no video, images, or audio — it outputs text.

What you can make with it

  • A clean transcript and structured summary of a long recording — an interview, webinar, or course module — read in a single 1M-token pass without chunking
  • A shortlist of clip-worthy moments with timestamps, pulled from watching the video rather than scrubbing it by hand
  • Descriptions, alt text, and captions drafted from images or video frames you feed it
  • Answers and reasoning grounded in mixed inputs at once — audio plus slides plus a document — in one prompt
  • The comprehension layer of an agent that reads multimodal input and then calls tools to act on it
  • Nothing visual as output — Omni-Flash understands and writes text; it generates no images, video, or audio

How Kompozy turns Qwen3.8-Omni-Flash output into content

Omni-Flash is the eyes and ears at the front of a repurposing workflow — and that is exactly the half most solo creators skip because it used to be slow and expensive. Record one long thing: a podcast episode, a client webinar, a Q&A. Omni-Flash reads the entire recording in one pass and hands you a transcript, the quotable moments, and where the good clips start. But a set of notes and timestamps is not content. Nothing is filmed anew, designed, branded, or posted. Turning that understanding into finished, on-brand posts across platforms is a separate job, and it is the one [Kompozy](/) does.

Here is the concrete loop. Let Omni-Flash comprehend the recording and surface the moments, then drop that source into Kompozy and pick your formats. From that single input Kompozy generates the assets a text model can't: [Clipped Shorts](/glossary/clipped-short) cut straight from the long-form, a captioned [Persona Short](/glossary/persona-shorts) fronted by a face-locked HeyGen avatar, brand-exact [Carousel Posts](/glossary/hyperframes), quote graphics pulled from the strongest lines, photo posts, a full blog article, and an email newsletter — each rewritten under a [Persona Brief](/glossary/persona-brief) so the voice reads as yours. [Autopilot](/glossary/autopilot) then schedules and publishes the batch across the eight social platforms plus blog and email, every asset clearing a per-post review gate first. Because Omni-Flash makes the comprehension step so cheap, you can record more and let the engine handle the parts it can't touch — the visuals, the brand identity, and the distribution. On the Founding tier you can bring your own Qwen key so the understanding layer runs at cost inside Kompozy.

  1. Feed a long recording — podcast, webinar, interview, or a set of images — into Qwen3.8-Omni-Flash and have it return a transcript, key moments, and clip timestamps across its 1M-token context.
  2. Drop that recording (or its structured notes) into Kompozy as a source and choose your formats.
  3. Fan the one recording into Clipped Shorts, a persona/avatar short, carousels, quote graphics, photo posts, a blog, and a newsletter — all in one Persona Brief voice with a consistent face.
  4. Review the batch in the per-post queue so nothing off-brand ships.
  5. Let Autopilot schedule and publish across the eight social platforms plus blog and email — and bring your own Qwen key on the Founding tier to keep the ingestion step near cost.

Frequently asked questions

What is Qwen3.8-Omni-Flash?

Qwen3.8-Omni-Flash is Alibaba Qwen team's omni-modal understanding model, released September 18, 2026. It takes text, images, audio, and video in a single request, carries a one-million-token context, and is built around comprehending audio and video and using tools agentically. It returns text — for generated speech Qwen points to its separate Qwen3.5-Omni model.

Is Qwen3.8-Omni-Flash a video generator like Gemini Omni Flash?

No. Despite the similar name, Qwen3.8-Omni-Flash reads audio and video and returns text — it is an understanding model. Google's Gemini Omni Flash generates and edits video. They do opposite jobs: one comprehends clips, the other renders them.

Can Qwen3.8-Omni-Flash generate images or video?

No. It reasons over multimodal input but returns text; it renders no images, video, or audio. To turn what it surfaces from a recording into finished visual posts, you pair it with a generation-and-publishing engine like Kompozy.

How much does Qwen3.8-Omni-Flash cost?

Hosted API rates are around $0.15 per million input tokens and $0.47 per million output, and Qwen reports audio input costs cut over 98% and audio-visual over 93% versus Qwen3.5-Omni-Plus. Access is via QwenCloud, Alibaba Cloud Model Studio, and Qwen Studio; weights were not open-sourced at launch. Confirm current figures on Qwen's site.

How do I turn Qwen3.8-Omni-Flash output into finished, published content?

Use Omni-Flash to comprehend a long recording and surface the quotes and clip moments, then bring the source into Kompozy. Kompozy generates 18 formats from that one input — Clipped Shorts, persona/avatar video, carousels, quote graphics, photo posts, a blog, and a newsletter — holds a consistent face and voice, and schedules and publishes across the eight social platforms plus blog and email.

Related tools

  • Qwen3.8-Flash-NextAlibaba's cheap, fast, open-weight multimodal model, released in late August 2026 as an early architecture preview of the coming Qwen4 family — a mixture-of-experts design tuned for "ultimate cost efficiency" with a long context and a novel N-gram embedding layer.
  • Gemini Omni FlashGoogle's conversational video model — generate a clip, then refine it by chatting instead of re-prompting.
  • Nari Qwen3-TTS & Qwen3-ASRNari Labs' low-latency, low-cost voice models — Qwen3-TTS 1.7B for text-to-speech and Qwen3-ASR 1.7B for transcription — that topped Coval's public voice-AI benchmark in September 2026.
  • Qwen3.8Alibaba's next flagship Qwen model — a ~2.4-trillion-parameter sparse-MoE model announced in July 2026, previewing now as Qwen3.8-Max-Preview and slated to go open-weight. It is the first Qwen flagship above 1T parameters to support multimodal input, and the team pitches it as frontier-class.

← All AI tools · Get started →