StepFun's reasoning model, previewed September 18, 2026 — a chain-of-thought LLM with a one-million-token context and text + image input that scores near the top of its price tier on the Artificial Analysis Intelligence Index at $1 in / $2.70 out per million tokens.
Last verified · 2026-09-19 · by Moe Ameen
Step 5 Preview is a reasoning large language model from StepFun, the Shanghai AI lab founded in 2023 by former Microsoft researchers and counted among China's "AI Tiger" companies. StepFun previewed it on September 18, 2026. Being a reasoning model, it works through problems with extended chain-of-thought before answering, which favors multi-step tasks — analysis, synthesis, structured drafting — over quick single-turn replies.
Its two headline traits are context and price. It carries a one-million-token context window, so you can feed it a full transcript, a long document, or a stack of research and have it reason over the whole thing at once. It accepts both text and image inputs and generates text output. StepFun's API prices it at $1.00 per million input tokens and $2.70 per million output tokens, with a 95% discount on cached input. Independent benchmarking from Artificial Analysis put its composite Intelligence Index score near 44 — around #25 of the 200 models tracked and well above the median for models at a similar price — with measured output near 99.8 tokens per second, though it tends to be verbose, generating a high volume of tokens on the way to an answer.
The clean framing for a creator: Step 5 is a text-and-reasoning engine, not a content engine. It reads and writes — including reading an image and reasoning about it — but it produces no video, no designed image, no carousel, no captions, and it publishes nothing. It is one of many strong, cheap models a creator can use for the thinking-and-drafting step. As a preview, its scores and pricing are a snapshot and will shift — confirm current details on StepFun's documentation before building on it.
Step 5's real edge for a creator is the front of the pipeline: a million-token context and cheap, capable reasoning make it very good at digesting a large source — a full webinar transcript, a quarter of newsletter archives, a pile of customer calls — and handing back a tight, structured brief. That is genuinely useful, and it is also where most creators stop, because the leap from "a good brief in a chat window" to "a week of finished, on-brand posts across every platform" is the actual work. Step 5 does not make that leap: it writes no captions on a video, designs no carousel, films no avatar, and posts to nothing. [Kompozy](/) is the engine that takes over exactly there.
The concrete workflow: use Step 5 to reason over your long source and produce the brief or draft, then bring that same source into Kompozy and pick your formats. From one input, and governed by a [Persona Brief](/glossary/persona-brief) that locks your voice and banned words so nothing reads like raw model output, Kompozy generates what a text model structurally can't — [Clipped Shorts](/glossary/clipped-short) cut at the strong moments, captioned [Persona Shorts](/glossary/persona-shorts) fronted by a face-locked HeyGen avatar, brand-exact [Carousel Posts](/glossary/hyperframes), quote graphics, photo posts, a blog article, and an email newsletter. Then [Autopilot](/glossary/autopilot) schedules and publishes the whole batch across the eight social platforms plus blog and email, each asset clearing a per-post review gate first. Step 5 turns a mountain of source material into a sharp draft; Kompozy turns that draft into published content in every format your audience actually sees.
It is a reasoning large language model previewed by the Shanghai AI lab StepFun on September 18, 2026. It uses chain-of-thought reasoning, has a one-million-token context window, and accepts text and image inputs while outputting text. Artificial Analysis scored its Intelligence Index near 44, high for its price tier.
StepFun prices it at $1.00 per million input tokens and $2.70 per million output tokens, with a 95% cached-input discount. Artificial Analysis measured output near 99.8 tokens per second with time-to-first-token around 3 seconds, though it is verbose. As a preview, confirm current figures on StepFun's documentation.
No. Step 5 outputs text. It can read an image and reason about it, but it generates no video, designed images, carousels, or captions, and it publishes nothing. To turn its text into finished, published content, pair it with a generation-and-publishing engine like Kompozy.
Use Step 5 for the reasoning and drafting step over a long source, then bring that source into Kompozy. Kompozy fans one input into clips, a persona/avatar short, carousels, quote graphics, a blog, and a newsletter under one Persona Brief, and schedules and publishes them across the eight social platforms plus blog and email.