// AI NEWS · MODEL RELEASE

Alibaba Releases Qwen3.8-Omni-Flash, a 1M-Context Omni-Modal Model Built for Agentic Audio and Video — With Audio Input Costs Cut Roughly 98%

Qwen shipped an omni-modal model that reads text, images, audio, and video in a single request, holds a one-million-token context, and leans hard on agentic tool use — while slashing the price of audio and audio-visual input over its previous Omni model.

2026-09-18 · by Moe Ameen

What happened

Alibaba's Qwen team released Qwen3.8-Omni-Flash on September 18, 2026, describing it as its first omni-modal model built around agentic capabilities. In one request the model takes text, images, audio, and video together and returns text, and it carries a one-million-token context window — long enough to reason over a full recording, a stack of screenshots, and a brief in a single pass. Qwen frames the release around two things: understanding audio and video well, and using tools to act on that understanding.

The efficiency story is the headline. Qwen reports Qwen3.8-Omni-Flash improves on the earlier Qwen3.5-Omni-Plus by roughly 26% across a suite of about 30 audio and video tests, and on the OmniVideoBench benchmark it moves from 63.4 to 67.8 while cutting token consumption by about 45.7%. The pricing moved just as sharply: Qwen says audio input costs are down more than 98% and audio-visual input costs down over 93% versus the prior Omni model, with hosted API rates around $0.15 per million input tokens and $0.47 per million output tokens.

The model is available through the API on QwenCloud, Alibaba Cloud Model Studio, and Qwen Studio. As of launch the weights were not open-sourced — access is API-only, which is a departure from the open-weight releases elsewhere in the Qwen3.8 line. Qwen also points developers who need generated speech to its Qwen3.5-Omni model, so treat Omni-Flash as an understanding-and-reasoning engine over multimodal input rather than a media generator. As with any fresh release, the benchmark and pricing figures are largely vendor-reported at launch — confirm them against Qwen's own materials before depending on a single number.

Why it matters for creators

  • Understanding audio and video just got cheap. When a model can transcribe, summarize, and reason over an hour of footage for cents, "turn this podcast into content" stops being a manual, expensive first step.
  • The 1M-token context means a creator can hand it a whole long recording — a webinar, an interview, a course module — and get structured notes, quotes, and chapter breaks back in one shot, without chunking.
  • It is an understanding model, not a video generator. Unlike Google's Gemini Omni Flash, which renders clips, Qwen3.8-Omni-Flash reads media and returns text — so it complements a generation tool rather than replacing one.
  • Agentic tool use is the point of the release. That direction matters for anyone building an automated content pipeline where a model reads raw input and decides what to do with it.
  • Weights are API-only at launch and figures are vendor-reported, so treat the specifics as an early snapshot and the model as a hosted service, not something to self-host yet.

How to act on this with Kompozy

Read this as a price signal about the front of the content pipeline. The expensive, manual part of repurposing has always been getting raw media into a usable form — transcribing an hour-long interview, pulling the quotable moments, noting where the good clips start. Qwen3.8-Omni-Flash makes that step cheap and fast: it ingests the audio and video and hands back structured text you can actually work from. But understanding a recording is not the same as having posts. The model reads and reasons; it renders no video, no carousels, no images, and it publishes nothing. The distance from "I understand what's in this footage" to "on-brand posts live across every platform" is the whole job, and that is where [Kompozy](/) operates.

The practical loop is model-agnostic by design. Let Omni-Flash read the raw recording and surface the moments; then drop the source into Kompozy and let the engine do the part a text model can't. From one input Kompozy generates roughly 25–35 finished assets across 18 formats — a captioned [Persona Short](/glossary/persona-shorts) fronted by a face-locked HeyGen avatar, [Clipped Shorts](/glossary/clipped-short) cut from the long-form, brand-exact [carousels](/glossary/hyperframes), quote graphics, photo posts, a blog article, and an email newsletter — each governed by a [Persona Brief](/glossary/persona-brief) so the voice stays yours, routed through a per-post review gate, then scheduled and published by [Autopilot](/glossary/autopilot) across the eight social platforms plus blog and email. Cheaper multimodal understanding is a tailwind for that workflow, not a substitute for it — and on the Founding tier you can bring your own Qwen key so the understanding layer runs at cost inside the engine.

Quick takeaways

  • Qwen released Qwen3.8-Omni-Flash on September 18, 2026 — an omni-modal model taking text, images, audio, and video, with a 1M-token context and an agentic focus.
  • It reportedly beats Qwen3.5-Omni-Plus by ~26% across ~30 audio/video tests, and improves OmniVideoBench from 63.4 to 67.8 while cutting token use ~45.7%.
  • Audio input costs are down over 98% and audio-visual input over 93% versus the prior Omni model; hosted rates are around $0.15/$0.47 per million input/output tokens.
  • Access is API-only via QwenCloud, Alibaba Cloud Model Studio, and Qwen Studio — weights were not open-sourced at launch.
  • It understands and reasons over media but returns text; it does not generate video or publish. Kompozy turns that understanding into finished, scheduled posts.

Frequently asked questions

What is Qwen3.8-Omni-Flash?

It is an omni-modal AI model from Alibaba's Qwen team, released September 18, 2026. It takes text, images, audio, and video in a single request, carries a one-million-token context window, and is built around understanding audio and video and using tools agentically. It returns text — for generated speech Qwen points developers to its Qwen3.5-Omni model.

Can Qwen3.8-Omni-Flash generate video like Gemini Omni Flash?

No. Despite the similar name, Qwen3.8-Omni-Flash is an understanding model — it reads audio and video and returns text. Google's Gemini Omni Flash is a video generation model. Qwen's model is closer to a multimodal reasoning engine you feed media to, not a tool that renders clips.

How much cheaper is Qwen3.8-Omni-Flash?

Qwen reports audio input costs are down more than 98% and audio-visual input costs down over 93% versus the earlier Qwen3.5-Omni-Plus, with hosted API rates around $0.15 per million input tokens and $0.47 per million output tokens. Confirm current pricing on Qwen's own pages, since figures are fresh at release.

Is Qwen3.8-Omni-Flash open-weight?

Not at launch. Access is API-only through QwenCloud, Alibaba Cloud Model Studio, and Qwen Studio — a departure from the open-weight releases elsewhere in the Qwen3.8 line. Check Qwen's materials for any later open-weight release.

Related news

← All AI news · Get started →