The open-weight release of Qwen3.8-Max lands on Hugging Face and ModelScope in August 2026 — a 2.4-trillion-parameter mixture-of-experts model with ~95B active, a long context, and a custom license that asks large commercial users for a cut.
2026-08-12 · by Moe Ameen
Alibaba's Qwen team released the open weights for Qwen3.8-2.4T-A95B — the downloadable version of its Qwen3.8-Max flagship — in August 2026, publishing the model on Hugging Face and ModelScope after making it accessible through Qwen Chat and the Qwen API earlier in the month. It is the first Qwen-Max-class flagship to ship as downloadable weights rather than API-only.
The model is a sparse mixture-of-experts (MoE) design with roughly 2.4 trillion total parameters and about 95 billion active per token, built on the Qwen3.5 architecture — interleaving linear-attention "Gated DeltaNet" layers with standard gated-attention layers. It carries a 262,144-token native context extensible toward roughly one million tokens, output length up to about 128K, and configurable reasoning tiers that trade compute for answer quality per request. Per the official model card, the released checkpoint is text-only and requires thinking mode for all interactions — a detail worth flagging, because earlier announcement coverage described the broader Qwen3.8 line as multimodal. On Alibaba's own benchmark reporting it posts strong research and reasoning scores — figures in the low-90s on suites like PaperBench and GPQA Diamond among them — and is positioned as competitive with leading frontier models; treat those as vendor numbers until independent evaluations land.
The headline caveat is the license. Rather than the permissive Apache 2.0 terms used on smaller Qwen releases, the weights ship under a custom "qwen3.8-max" license, and Alibaba has signaled revenue-sharing terms for large commercial users and cloud providers that offer the model as a service at scale — reportedly aimed at big enterprises and hosts rather than researchers, developers, or small-scale users, with the exact thresholds and percentages still being finalized. It makes Alibaba one of the first major labs to attach a deployment fee to an "open" flagship. On the practical side, a 2.4-trillion-parameter model is a data-center workload: reference serving uses rack-scale Blackwell-class GPU systems, though FP8, GGUF, and NVFP4 quantized builds and serving stacks like SGLang, vLLM, and NVIDIA NIM lower the barrier. Confirm the specs, license terms, and exact release date on Qwen's official model card before building on any single figure.
The fastest move on this release is not to spin up a GPU cluster — it is to publish your read on it. "A Max-class Qwen just went open weight, but under a revenue-share license" is a beat your audience is already seeing, and being early with a sharp take beats being thorough a week late. Feed your angle into [Kompozy](/) as a source and it becomes a [Blog Article](/glossary/output-buckets) on what a licensed "open" flagship means for content teams, a [Carousel](/glossary/output-buckets) laying out where Qwen3.8-2.4T-A95B sits against DeepSeek V4 and Kimi K3, a few captioned shorts, and platform-native posts in your voice through the [Persona Brief](/glossary/persona-brief) — scheduled across the eight social platforms plus blog and email from one queue with [Autopilot](/glossary/autopilot).
The second move is for teams that actually adopt the weights. Self-hosting Qwen3.8-2.4T-A95B gives you private, frontier-grade drafting — and nothing downstream of text, because the checkpoint renders no media, holds no brand voice, and publishes nowhere. Kompozy is the layer that turns that private draft into the finished feed: point your self-hosted Qwen endpoint at Kompozy through bring-your-own-key on the Founding tier, and the angles it reasons out become [Persona Shorts](/glossary/persona-shorts) avatar video, brand-exact carousels, [Quote Graphics](/glossary/output-buckets) of the sharpest lines, a formatted blog, and a newsletter — all held to one look by [HyperFrames](/glossary/hyperframes) and published across nine destinations. Because Kompozy treats every model as an interchangeable input, adopting Qwen for private reasoning is pure upside: you keep your generation stack under your own control right up to the render-and-ship boundary Kompozy owns. One note: Kompozy's own copy generation runs on Claude and OpenAI, so a self-hosted Qwen is the private drafting brain feeding the pipeline, not a swap for it.
Qwen3.8-2.4T-A95B is the open-weight release of Alibaba's Qwen3.8-Max flagship, published on Hugging Face and ModelScope in August 2026. It is a sparse mixture-of-experts model with roughly 2.4 trillion total parameters and about 95 billion active per token, built on the Qwen3.5 architecture, with a 262K-token native context extensible toward ~1M. Per the official model card, the released checkpoint is text-only and runs in thinking mode.
The weights are downloadable, but under a custom "qwen3.8-max" license rather than Apache 2.0. Alibaba has signaled revenue-sharing terms for large commercial users and cloud providers offering it as a service at scale, with thresholds and rates still being finalized. Small-scale and research use is far less restricted, but confirm the license on the official model card before commercial deployment.
No. The released checkpoint generates and reasons over text but produces no video, images, or designs, enforces no brand voice, and publishes to no platform. To turn its drafts into finished, on-brand posts across platforms, pair it with a content engine like Kompozy — which can call your own Qwen key on the Founding tier — that renders the media and handles scheduling and publishing.