On September 18, 2026, Shanghai AI lab StepFun previewed Step 5, a reasoning model that scores near the top of the mid-price tier on the Artificial Analysis Intelligence Index at $1 in / $2.70 out per million tokens.
2026-09-19 · by Moe Ameen
On September 18, 2026, StepFun — the Shanghai AI lab founded in 2023 by former Microsoft researchers and counted among China's so-called "AI Tiger" companies — previewed Step 5, a new reasoning model. Artificial Analysis, the independent benchmarking service, listed Step 5 Preview with a score of roughly 44 on its composite Intelligence Index, placing it around #25 of the 200 models tracked and well above the median (about 25) for models in a similar price tier.
Step 5 is a reasoning model, meaning it works through problems with extended chain-of-thought before answering. StepFun's API prices it at $1.00 per million input tokens and $2.70 per million output tokens, with a 95% discount on cached input. It carries a one-million-token context window and accepts both text and image inputs, generating text output. On throughput, Artificial Analysis measured output around 99.8 tokens per second with a time-to-first-token near 3 seconds — above-average speed, though the model is notably verbose, generating a high volume of output tokens on its way to an answer.
The "Preview" label matters: this is an early, evolving release, not a finalized product, and independent figures are a snapshot that will shift as the model is refined and more evaluations land. As always with vendor and third-party benchmarks, treat the specific numbers as indicative rather than final, and confirm current pricing and specs on StepFun's own documentation before building on them.
The useful way to read a launch like Step 5 is as another drop in the cost of raw intelligence — a strong reasoning model, a huge context window, and a low per-token price. What it does not change is the part that actually eats a creator's week: a reasoned draft is not a set of published posts. Step 5 will read your long transcript and write you a sharp brief; it will not cut the clip, design the carousel, film the avatar, caption the short, or schedule any of it. That gap between "good text out of a model" and "finished content live on eight platforms" is exactly where [Kompozy](/) lives.
Here is how a creator acts on this today. Use a model like Step 5 for what it is best at — reasoning over a long source to produce a tight draft or outline — then bring that source into Kompozy and let it manufacture the rest. Under a [Persona Brief](/glossary/persona-brief) that locks your voice and banned words, Kompozy fans one input into the formats a text model can't: [Clipped Shorts](/glossary/clipped-short) and captioned [Persona Shorts](/glossary/persona-shorts) fronted by a face-locked avatar, brand-exact [Carousel Posts](/glossary/hyperframes), quote graphics, photo posts, a blog article, and an email newsletter. Then [Autopilot](/glossary/autopilot) schedules and publishes the batch across the eight social platforms plus blog and email behind a per-post review gate. Cheaper reasoning lowers the cost of the first draft; Kompozy is what turns that draft into a week of on-brand, published content.
Step 5 is a reasoning model previewed by the Shanghai AI lab StepFun on September 18, 2026. It uses chain-of-thought reasoning, carries a one-million-token context window, and accepts text and image inputs while generating text output. Independent benchmarking from Artificial Analysis put its Intelligence Index score near 44, high for its price tier.
StepFun's API prices Step 5 at $1.00 per million input tokens and $2.70 per million output tokens, with a 95% discount on cached input. Because it is a preview, confirm current pricing on StepFun's documentation before relying on it.
No. Step 5 outputs text. It can read images and reason over them, but it does not generate video, images, carousels, or captions, and it publishes nothing. To turn its output into finished, published content you pair it with a generation-and-publishing engine like Kompozy.
Use it for the drafting and reasoning step — outlining a piece, analyzing a long transcript, synthesizing research across its million-token context. Then bring that source into Kompozy, which fans one input into clips, persona video, carousels, quote graphics, a blog, and a newsletter, and schedules them across the eight social platforms plus blog and email.