Z.ai's cheap, fast, natively multimodal model, launched August 26, 2026 — the model previously teased on OpenRouter as the stealth "Ox Alpha." A 320B-parameter MoE (18B active) with a claimed 1M-token context, MIT open weights, and API pricing near a tenth of the flagship.
Last verified · 2026-08-26 · by Moe Ameen
GLM-5.3-Flash is a large language model from Z.ai — the Chinese lab formerly known as Zhipu AI — launched on August 26, 2026. It is the production release of the anonymous "Ox Alpha" model that had appeared on OpenRouter around August 20 with no branding, which community fingerprinting had already traced to Z.ai's GLM family. Where the full GLM-5.3 flagship is a large text-and-code model aimed at heavyweight coding and agent work, Flash is the small, fast, cheap sibling — and, per Z.ai, the first natively multimodal model in the GLM-5 line.
Architecturally it is a mixture-of-experts model with roughly 320 billion total parameters and about 18 billion active per token, with a claimed context window of 1,048,576 tokens. It accepts text and image input (Z.ai's guides also document video and file input) and returns text. It exposes forced-thinking reasoning that is always on and cannot be disabled. Independent benchmarking from Artificial Analysis placed its Intelligence Index around 57 — strong for an open-weight model of its size — while noting output throughput on the slower side and a verbose style.
The headline is the economics. At launch Z.ai listed API pricing around $0.15 per million input tokens and $0.50 per million output tokens, with cached input near $0.03, and a launch promotion that temporarily halved those rates — roughly a tenth of what the heavier GLM tiers cost. The weights and inference code were published under an MIT license on Hugging Face, so it can be self-hosted. On coding and agent tests Z.ai reported figures such as Terminal-Bench 2.1 at 84.3 and DeepSWE v1.1 at 63.4, putting it within striking distance of far pricier frontier models. Treat exact prices, promo dates, and benchmark numbers as an early snapshot and confirm them on Z.ai before you depend on one.
The reason to reach for Flash over a pricier model is volume and vision: it is cheap enough to draft your whole week and it can actually look at the raw material you already have. Point it at a folder of screenshots, a whiteboard photo, a product shot, or a rough transcript, and it will read them and come back with hooks, captions, and script variants for pennies. That is a genuinely good front end. What it hands back is still text on a screen — no face, no video, no carousel, no schedule, no post. Turning that text into finished, on-brand, published content is a different job, and it is the one [Kompozy](/) does.
Here is the concrete loop. Have Flash read your existing assets and draft a batch of angles, then drop the strongest into Kompozy as a source. From that single input Kompozy generates roughly 25–35 finished assets across 18 formats — a captioned [Persona Short](/glossary/persona-shorts) fronted by a face-locked HeyGen avatar, a brand-exact [Carousel](/glossary/hyperframes), quote graphics pulled from the copy, photo posts, a full blog article, and an email newsletter — each rewritten under a [Persona Brief](/glossary/persona-brief) so the voice reads as yours and not as raw model output. [Autopilot](/glossary/autopilot) then schedules and publishes the set across the eight social platforms plus blog and email, every asset clearing a per-post review gate first. Because Flash is so cheap, you can afford to over-generate drafts and let Kompozy do the expensive parts — the visuals, the brand identity, and the distribution — that the model can't touch. On the Founding tier you can even bring your own Z.ai key so Flash stays your low-cost drafting layer inside Kompozy.
GLM-5.3-Flash is a fast, low-cost, natively multimodal model from Z.ai (formerly Zhipu AI), launched August 26, 2026. It is the production release of the stealth "Ox Alpha" model, a 320B-parameter mixture-of-experts design with about 18B active per token, a claimed 1M-token context, and MIT-licensed open weights.
The full GLM-5.3 flagship is a larger, text-and-code model built for heavyweight coding and agent work. Flash is the smaller, faster, much cheaper sibling — roughly a tenth the per-token price — and, per Z.ai, the first natively multimodal model in the GLM-5 line, taking image input as well as text. It trades some peak capability for cost and speed.
No. It can read images and other inputs, but it returns text and code only — no generated images, video, or audio. To turn its drafts and image analysis into finished visual posts, you pair it with a generation-and-publishing engine like Kompozy.
At launch Z.ai listed API pricing around $0.15 per million input tokens and $0.50 per million output, with cached input near $0.03 and a launch promotion that temporarily halved those rates. The weights are MIT-licensed on Hugging Face, so it can also be self-hosted. Confirm current pricing on Z.ai.
Draft cheaply in Flash, then bring the best output into Kompozy as a source. Kompozy generates 18 formats from that one input — persona/avatar video, carousels, quote graphics, photo posts, a blog, and a newsletter — holds a consistent face and voice, and schedules and publishes across the eight social platforms plus blog and email.