StepFun Step 5 Preview review: honest scoring on intelligence, its 1M-token context, pricing, speed, verbosity, and who this reasoning model actually fits.
Step 5 Preview is a strong-value reasoning model. StepFun scores near the top of its price tier on the Artificial Analysis Intelligence Index (around 44), pairs that with a one-million-token context and text-plus-image input, and charges only $1.00 / $2.70 per million tokens with a 95% cached-input discount. Scored as a reasoning LLM, it is excellent for the price. The honest caveats: it is a preview whose numbers will move, it is verbose (which inflates real output cost), and it outputs text only — it makes no content and publishes nothing.
Step 5 Preview is the reasoning model StepFun — the Shanghai AI lab founded in 2023 by former Microsoft researchers, one of China's "AI Tiger" companies — previewed on September 18, 2026. It arrived with a genuinely interesting pitch for the mid-price tier: a composite Intelligence Index score near 44 on Artificial Analysis, a one-million-token context window, and pricing of $1.00 per million input tokens and $2.70 per million output. This review scores it on the things that matter for a reasoning LLM: intelligence, context length, multimodal input, speed, verbosity, price, and availability.
I run Kompozy, a content generation and publishing engine, so I'll be upfront that Kompozy is not a foundation-model lab and doesn't compete with Step 5 — one is a model you call by API, the other is the layer that turns any model's text into finished, published content. That's exactly why I can score the model on its own terms without an axe to grind. Where the two touch is narrow and I'll be explicit about it: a reasoning model produces text, and text is the input to a content operation, not the operation itself.
Two things anchor the verdict. First, the value is real: near-top-of-tier intelligence at $1/$2.70 per million tokens, plus a huge context and image input, is a lot for the money. Second, the caveats are the kind you expect from a preview: the benchmark figures are a snapshot that will shift, the model is notably verbose so your real bill runs higher than the sticker output rate implies, and — like any LLM — it outputs text and nothing else.
Everything below reflects Step 5 Preview as described around its September 18, 2026 preview, verified against Artificial Analysis and StepFun's materials. Because it is a preview, availability, pricing, and scores will evolve — confirm current details on StepFun's documentation before you build on it.
Step 5 Preview is a reasoning large language model: it works through problems with extended chain-of-thought before answering, which suits multi-step analysis and structured drafting over quick single-turn replies. It carries a one-million-token context window, so it can reason over a full transcript, a long report, or a stack of documents in one pass, and it accepts both text and image inputs while generating text output. StepFun exposes it through its API at $1.00 per million input tokens and $2.70 per million output tokens, with a 95% discount on cached input. On independent benchmarks, Artificial Analysis placed its composite Intelligence Index near 44 — around #25 of the 200 models it tracks, and well above the median for models in a similar price tier — with output throughput measured near 99.8 tokens per second and time-to-first-token around three seconds. The consistent knock is verbosity: it generates a high volume of output tokens on its way to an answer, which is a cost and latency consideration in practice. What it does is turn text (and images) into reasoned text, cheaply and over a very long context — that is the whole of its job.
Step 5 Preview fits anyone whose bottleneck is reasoning over text: developers and teams that need capable analysis or structured drafting at a low per-token price, researchers synthesizing long documents, and creators who want a cheap model for the thinking-and-outlining step. The million-token context makes it especially good for long-source work — digesting a full recording or a large archive rather than a fragment — and image input helps when a chart or screenshot is part of the source. It fits poorly for anyone who needs a finished product rather than a draft: it produces no video, images, carousels, or captions, offers no brand-voice layer, no clipping, and no scheduler, and its verbosity makes it a weaker pick where output cost or tight latency dominate. As a preview, it is also a weak fit for anyone who needs a stable, benchmarked model to standardize a production workflow on today.
| Dimension | Score | Why |
|---|---|---|
| Reasoning / intelligence | 4.3 / 5 | A composite Intelligence Index near 44 on Artificial Analysis is high for its price tier — front-of-pack among comparably priced reasoning models, though the figures are a preview snapshot. |
| Context length | 4.6 / 5 | A one-million-token window lets it reason over a full transcript or document set in a single pass — a genuine strength for long-source work. |
| Multimodal input | 3.8 / 5 | Accepts text and image inputs, useful for reasoning over charts and screenshots — but input only; it does not generate images. |
| Speed / throughput | 3.9 / 5 | Output near 99.8 tokens/second with time-to-first-token around three seconds — above average, though the verbosity works against effective latency. |
| Verbosity / output efficiency | 3.2 / 5 | Generates a high volume of output tokens to reach an answer, which raises real cost and latency beyond what the headline rate suggests. |
| Cost & value | 4.4 / 5 | At $1.00 / $2.70 per million tokens with a 95% cached-input discount, it is aggressively priced for its intelligence level. |
| Availability / maturity | 3.0 / 5 | A preview release accessed via StepFun's API — figures and behavior will change, and there is no consumer app surface. |
| Content-workflow scope | 1.5 / 5 | Text output only — no clipping, feed captions, images, video, scheduling, or publishing. Not what the model is for. |
On sticker price, Step 5 Preview is one of the better deals in the reasoning tier: $1.00 per million input tokens and $2.70 per million output, plus a 95% discount on cached input, for a model scoring near the top of its price bracket on the Artificial Analysis Intelligence Index. If you were paying more for comparable reasoning, that is a real saving, and the cached-input discount rewards workloads that reuse a long context repeatedly — which is exactly the case its million-token window invites.
The caveat that changes the math is verbosity. A reasoning model that generates a high volume of output tokens on its way to an answer bills you for all of them, so the effective cost per finished task runs above what the $2.70 output rate implies on its own. Artificial Analysis flags the high output-token count directly. For output-heavy or latency-sensitive workloads, price the model on tokens actually generated for your tasks, not on the headline rate.
Set against the wider market, Step 5 Preview competes in a crowded, fast-cheapening reasoning tier alongside models from DeepSeek, Qwen, and others, and its combination of price, context length, and score makes it a reasonable pick to test. But it is priced and sold as a model, not a product — the cost of turning its output into finished content sits entirely downstream, and no per-token discount touches that. Verify current pricing on StepFun's documentation before you commit, since preview pricing can move.
| Use case | Fit | Why |
|---|---|---|
| Reasoning over long documents or transcripts | Strong | The million-token context lets it analyze a full source in one pass rather than a fragment — its clearest strength. |
| Low-cost drafting and outlining at volume | Strong | Near-top-of-tier intelligence at $1/$2.70 per million tokens makes it efficient for the thinking-and-drafting step. |
| Analyzing charts, slides, or screenshots | OK | It accepts image input and reasons over it, but only outputs text — useful for interpretation, not generation. |
| Latency- or output-cost-sensitive production apps | Weak | Its verbosity raises effective cost and latency; benchmark it on your own tasks before relying on it. |
| Standardizing a stable production workflow | Weak | As a preview, its scores, pricing, and behavior will change. |
| Producing video, images, or social posts | Weak | It outputs text only — no clips, designed images, carousels, or captions. |
| Publishing and scheduling content across platforms | Weak | It has no distribution layer at all; that is a separate job entirely. |
To review this honestly: Kompozy is not an alternative to Step 5, because it is not a foundation model — you don't choose between them, and implying otherwise would be spin. Step 5 is one of many capable, cheap models a creator might use for the reasoning-and-drafting step. Kompozy is the layer that operationalizes whatever text a model produces. The two sit on opposite ends of the same pipeline: Step 5 reads a long source and returns a reasoned draft; Kompozy takes a source and returns finished, on-brand posts in every format, then publishes them.
The practical point for anyone weighing Step 5 is that the model is the cheap, easy part now — capable reasoning at a low per-token price is increasingly commodity. The expensive part is everything after the draft: cutting clips, designing carousels, filming a consistent avatar, captioning, keeping a brand voice, and getting all of it out across platforms on schedule and reviewed. Kompozy does that part from one input under a single Persona Brief and publishes across the eight social platforms plus blog and email. Use Step 5 (or any strong reasoning model) for the thinking; use Kompozy to turn the thinking into content that ships.
It is a reasoning large language model previewed by the Shanghai AI lab StepFun on September 18, 2026. It uses chain-of-thought reasoning, has a one-million-token context window, and accepts text and image inputs while outputting text. Artificial Analysis scored its Intelligence Index near 44, high for its price tier.
As a reasoning model for the price, yes — it scores near the top of its tier on Artificial Analysis, has a million-token context, and is cheap at $1.00 / $2.70 per million tokens. The honest caveats are that it is a preview whose numbers will move, it is verbose (raising real cost), and it outputs text only.
StepFun prices it at $1.00 per million input tokens and $2.70 per million output tokens, with a 95% discount on cached input. Because it tends to be verbose, price it on the tokens your tasks actually generate, and confirm current rates on StepFun's documentation since preview pricing can change.
It competes directly in the low-cost reasoning tier with models from DeepSeek and Qwen. Its draw is the combination of a near-top-of-tier Intelligence Index score, a million-token context, and low pricing. As with any preview, benchmark it on your own tasks rather than relying on leaderboard scores alone.
No. It outputs text and can reason over images you supply, but it generates no video, designed images, carousels, or captions, and it publishes nothing. To turn its output into finished, published content you pair it with a generation-and-publishing engine like Kompozy.
They sit on opposite ends of one pipeline and don't compete. Step 5 is a model that returns a reasoned draft; Kompozy is the engine that turns a source into finished clips, persona video, carousels, quote graphics, a blog, and a newsletter under one brand voice, then publishes them across the eight social platforms plus blog and email.
See StepFun Step 5 Preview vs Kompozy comparison → · Get Started →