// AI NEWS · MODEL RELEASE

Z.ai Launches GLM-5.3-Flash — the Stealth "Ox Alpha" Model, Now a Cheap Multimodal Open-Weight Release

The anonymous "Ox Alpha" model that appeared on OpenRouter in August is officially GLM-5.3-Flash: a 320B-parameter mixture-of-experts model, natively multimodal, priced near a tenth of the flagship, with MIT-licensed open weights.

2026-08-26 · by Moe Ameen

What happened

Z.ai — the Chinese lab formerly known as Zhipu AI — released GLM-5.3-Flash on August 26, 2026, and confirmed it is the model that had been circulating anonymously on OpenRouter as "Ox Alpha." Ox Alpha appeared around August 20 as a free, unbranded preview reasoning model; community fingerprinting tools that probe a model's tokenizer and serving behavior had already pointed hard at Z.ai's GLM family, and the launch now makes that official.

GLM-5.3-Flash is the small, fast, cheap sibling to the heavier GLM-5.3 flagship. It is a mixture-of-experts model with roughly 320 billion total parameters and about 18 billion active per token, with a claimed context window of 1,048,576 tokens. Z.ai describes it as the first natively multimodal model in the GLM-5 line — it accepts image input alongside text (its guides also document video and file input) and returns text. It exposes forced-thinking reasoning that is always on and cannot be disabled, and independent benchmarking from Artificial Analysis placed its Intelligence Index around 57, strong for an open-weight model of its size, while noting slower-than-median output throughput.

The headline is price. At launch Z.ai listed API pricing around $0.15 per million input tokens and $0.50 per million output, with cached input near $0.03 and a launch promotion that temporarily halved those rates — roughly a tenth of what the heavier GLM tiers cost. The weights and inference code were published under an MIT license on Hugging Face, so the model can be self-hosted. On coding and agent benchmarks Z.ai reported figures such as Terminal-Bench 2.1 at 84.3 and DeepSWE v1.1 at 63.4, close to far pricier frontier models. Prices, promo dates, and benchmark numbers are early and vendor-reported — confirm them on Z.ai before depending on any single figure.

Why it matters for creators

  • Drafting is now nearly free. At a fraction of a cent per typical generation, the cost of producing scripts, hooks, and caption variants for an entire content calendar rounds to nothing — the economic bottleneck moves off writing.
  • Multimodal input for pennies. Flash reads screenshots, product photos, and whiteboards, so creators can caption and analyze the raw material they already have without a separate vision model.
  • Open weights, MIT license. Anyone can self-host and wire it into a pipeline, which keeps pushing per-token prices down across the whole market.
  • The moat shifts to production and distribution. When the words are commodity-cheap, the hard, differentiated work is turning them into on-brand video, images, and carousels and getting them published everywhere on a schedule.
  • Vendor numbers, no third-party audit yet. The benchmark and pricing claims are Z.ai's own; treat them as preliminary until independent testing settles.

How to act on this with Kompozy

The useful way to read this launch is as a price signal: the drafting layer of content just got commodity-cheap, so the part that still takes work — and still separates a real content operation from a pile of text — is everything after the draft. GLM-5.3-Flash reasons and writes for pennies and can even read your raw footage and photos, but it stops at text on a screen. It renders no video, no carousels, no images; it holds no brand voice; it publishes nothing. That gap between "I have a week of cheap drafts" and "I have on-brand posts live across every platform" is exactly where [Kompozy](/) operates.

You can act on this today. Point Flash at your source material and let it draft a batch cheaply, then drop the strongest into Kompozy as a source. From that one input Kompozy generates roughly 25–35 finished assets across 18 formats — a captioned [Persona Short](/glossary/persona-shorts) with a face-locked HeyGen avatar, brand-exact [carousels](/glossary/hyperframes), quote graphics, photo posts, a blog article, and an email newsletter — each rewritten under a [Persona Brief](/glossary/persona-brief) so the voice stays yours, then routed through a per-post review gate and published by [Autopilot](/glossary/autopilot) across the eight social platforms plus blog and email. Because Flash is so cheap, you can over-generate drafts and let Kompozy spend its effort on the visuals, identity, and distribution the model can't touch. On the Founding tier you can bring your own Z.ai key so Flash stays your low-cost drafting front end inside the engine.

Quick takeaways

  • GLM-5.3-Flash = the officially-revealed "Ox Alpha" stealth model.
  • Cheap (about a tenth of the flagship), natively multimodal, MIT open weights.
  • It drafts and reasons; it does not generate visuals or publish.
  • The new bottleneck is production + distribution — Kompozy turns cheap drafts into finished, scheduled posts across platforms.

Frequently asked questions

Is GLM-5.3-Flash the same as Ox Alpha?

Yes. Ox Alpha was the anonymous stealth model that appeared on OpenRouter around August 20, 2026, and community fingerprinting had linked it to Z.ai's GLM family. Z.ai confirmed at the August 26 launch that it is GLM-5.3-Flash.

How much does GLM-5.3-Flash cost?

At launch Z.ai listed API pricing around $0.15 per million input tokens and $0.50 per million output, with cached input near $0.03 and a launch promotion that temporarily halved those rates — roughly a tenth of the heavier GLM tiers. The weights are MIT-licensed on Hugging Face for self-hosting. Confirm current numbers on Z.ai.

Can GLM-5.3-Flash make video or images for social?

No. It reads images and other inputs but returns text and code only — no generated visuals or audio. To turn its cheap drafts into finished posts you pair it with a generation-and-publishing engine like Kompozy, which produces the video, carousels, and images and publishes them across platforms.

Related news

← All AI news · Get started →