The anonymous "Ox Alpha" model that appeared on OpenRouter in August is officially GLM-5.3-Flash: a 320B-parameter mixture-of-experts model, natively multimodal, priced near a tenth of the flagship, with MIT-licensed open weights.
2026-08-26 · by Moe Ameen
Z.ai — the Chinese lab formerly known as Zhipu AI — released GLM-5.3-Flash on August 26, 2026, and confirmed it is the model that had been circulating anonymously on OpenRouter as "Ox Alpha." Ox Alpha appeared around August 20 as a free, unbranded preview reasoning model; community fingerprinting tools that probe a model's tokenizer and serving behavior had already pointed hard at Z.ai's GLM family, and the launch now makes that official.
GLM-5.3-Flash is the small, fast, cheap sibling to the heavier GLM-5.3 flagship. It is a mixture-of-experts model with roughly 320 billion total parameters and about 18 billion active per token, with a claimed context window of 1,048,576 tokens. Z.ai describes it as the first natively multimodal model in the GLM-5 line — it accepts image input alongside text (its guides also document video and file input) and returns text. It exposes forced-thinking reasoning that is always on and cannot be disabled, and independent benchmarking from Artificial Analysis placed its Intelligence Index around 57, strong for an open-weight model of its size, while noting slower-than-median output throughput.
The headline is price. At launch Z.ai listed API pricing around $0.15 per million input tokens and $0.50 per million output, with cached input near $0.03 and a launch promotion that temporarily halved those rates — roughly a tenth of what the heavier GLM tiers cost. The weights and inference code were published under an MIT license on Hugging Face, so the model can be self-hosted. On coding and agent benchmarks Z.ai reported figures such as Terminal-Bench 2.1 at 84.3 and DeepSWE v1.1 at 63.4, close to far pricier frontier models. Prices, promo dates, and benchmark numbers are early and vendor-reported — confirm them on Z.ai before depending on any single figure.
The useful way to read this launch is as a price signal: the drafting layer of content just got commodity-cheap, so the part that still takes work — and still separates a real content operation from a pile of text — is everything after the draft. GLM-5.3-Flash reasons and writes for pennies and can even read your raw footage and photos, but it stops at text on a screen. It renders no video, no carousels, no images; it holds no brand voice; it publishes nothing. That gap between "I have a week of cheap drafts" and "I have on-brand posts live across every platform" is exactly where [Kompozy](/) operates.
You can act on this today. Point Flash at your source material and let it draft a batch cheaply, then drop the strongest into Kompozy as a source. From that one input Kompozy generates roughly 25–35 finished assets across 18 formats — a captioned [Persona Short](/glossary/persona-shorts) with a face-locked HeyGen avatar, brand-exact [carousels](/glossary/hyperframes), quote graphics, photo posts, a blog article, and an email newsletter — each rewritten under a [Persona Brief](/glossary/persona-brief) so the voice stays yours, then routed through a per-post review gate and published by [Autopilot](/glossary/autopilot) across the eight social platforms plus blog and email. Because Flash is so cheap, you can over-generate drafts and let Kompozy spend its effort on the visuals, identity, and distribution the model can't touch. On the Founding tier you can bring your own Z.ai key so Flash stays your low-cost drafting front end inside the engine.
Yes. Ox Alpha was the anonymous stealth model that appeared on OpenRouter around August 20, 2026, and community fingerprinting had linked it to Z.ai's GLM family. Z.ai confirmed at the August 26 launch that it is GLM-5.3-Flash.
At launch Z.ai listed API pricing around $0.15 per million input tokens and $0.50 per million output, with cached input near $0.03 and a launch promotion that temporarily halved those rates — roughly a tenth of the heavier GLM tiers. The weights are MIT-licensed on Hugging Face for self-hosting. Confirm current numbers on Z.ai.
No. It reads images and other inputs but returns text and code only — no generated visuals or audio. To turn its cheap drafts into finished posts you pair it with a generation-and-publishing engine like Kompozy, which produces the video, carousels, and images and publishes them across platforms.