// AI TOOLS · STEPFUN STEP 5 PREVIEW

StepFun Step 5 Preview

StepFun's reasoning model, previewed September 18, 2026 — a chain-of-thought LLM with a one-million-token context and text + image input that scores near the top of its price tier on the Artificial Analysis Intelligence Index at $1 in / $2.70 out per million tokens.

Last verified · 2026-09-19 · by Moe Ameen

What StepFun Step 5 Preview is

Step 5 Preview is a reasoning large language model from StepFun, the Shanghai AI lab founded in 2023 by former Microsoft researchers and counted among China's "AI Tiger" companies. StepFun previewed it on September 18, 2026. Being a reasoning model, it works through problems with extended chain-of-thought before answering, which favors multi-step tasks — analysis, synthesis, structured drafting — over quick single-turn replies.

Its two headline traits are context and price. It carries a one-million-token context window, so you can feed it a full transcript, a long document, or a stack of research and have it reason over the whole thing at once. It accepts both text and image inputs and generates text output. StepFun's API prices it at $1.00 per million input tokens and $2.70 per million output tokens, with a 95% discount on cached input. Independent benchmarking from Artificial Analysis put its composite Intelligence Index score near 44 — around #25 of the 200 models tracked and well above the median for models at a similar price — with measured output near 99.8 tokens per second, though it tends to be verbose, generating a high volume of tokens on the way to an answer.

The clean framing for a creator: Step 5 is a text-and-reasoning engine, not a content engine. It reads and writes — including reading an image and reasoning about it — but it produces no video, no designed image, no carousel, no captions, and it publishes nothing. It is one of many strong, cheap models a creator can use for the thinking-and-drafting step. As a preview, its scores and pricing are a snapshot and will shift — confirm current details on StepFun's documentation before building on it.

What you can make with it

  • A reasoned brief, outline, or first draft synthesized from a long source — an interview transcript, a report, a research folder — using the million-token context
  • Analysis of a large body of text in one pass, rather than a summary of a fragment
  • Structured extraction: pulling themes, quotes, timestamps, or angles out of a long recording for you to build content around
  • Reasoning over an image you supply — reading a chart, a screenshot, or a slide and explaining or restructuring it as text
  • Draft copy for a post, email, or article that you then refine and design elsewhere
  • Note: outputs are text only — no video, images, carousels, captions, or published posts

How Kompozy turns StepFun Step 5 Preview output into content

Step 5's real edge for a creator is the front of the pipeline: a million-token context and cheap, capable reasoning make it very good at digesting a large source — a full webinar transcript, a quarter of newsletter archives, a pile of customer calls — and handing back a tight, structured brief. That is genuinely useful, and it is also where most creators stop, because the leap from "a good brief in a chat window" to "a week of finished, on-brand posts across every platform" is the actual work. Step 5 does not make that leap: it writes no captions on a video, designs no carousel, films no avatar, and posts to nothing. [Kompozy](/) is the engine that takes over exactly there.

The concrete workflow: use Step 5 to reason over your long source and produce the brief or draft, then bring that same source into Kompozy and pick your formats. From one input, and governed by a [Persona Brief](/glossary/persona-brief) that locks your voice and banned words so nothing reads like raw model output, Kompozy generates what a text model structurally can't — [Clipped Shorts](/glossary/clipped-short) cut at the strong moments, captioned [Persona Shorts](/glossary/persona-shorts) fronted by a face-locked HeyGen avatar, brand-exact [Carousel Posts](/glossary/hyperframes), quote graphics, photo posts, a blog article, and an email newsletter. Then [Autopilot](/glossary/autopilot) schedules and publishes the whole batch across the eight social platforms plus blog and email, each asset clearing a per-post review gate first. Step 5 turns a mountain of source material into a sharp draft; Kompozy turns that draft into published content in every format your audience actually sees.

  1. Use Step 5 Preview to reason over a long source — a transcript, report, or research set — inside its million-token context, and get back a structured brief or draft.
  2. Bring the original source into Kompozy as a source and set your Persona Brief so voice and banned words govern every output.
  3. Fan the one source into Clipped Shorts, a captioned persona/avatar short, carousels, quote graphics, a blog, and a newsletter — all in one brand voice with a consistent face.
  4. Review the batch in the per-post queue so nothing off-brand ships, and let Kompozy reframe each clip for TikTok, Reels, and Shorts.
  5. Let Autopilot schedule and publish across the eight social platforms plus blog and email.

Frequently asked questions

What is StepFun Step 5 Preview?

It is a reasoning large language model previewed by the Shanghai AI lab StepFun on September 18, 2026. It uses chain-of-thought reasoning, has a one-million-token context window, and accepts text and image inputs while outputting text. Artificial Analysis scored its Intelligence Index near 44, high for its price tier.

How much does Step 5 cost and how fast is it?

StepFun prices it at $1.00 per million input tokens and $2.70 per million output tokens, with a 95% cached-input discount. Artificial Analysis measured output near 99.8 tokens per second with time-to-first-token around 3 seconds, though it is verbose. As a preview, confirm current figures on StepFun's documentation.

Can Step 5 create videos, images, or social posts?

No. Step 5 outputs text. It can read an image and reason about it, but it generates no video, designed images, carousels, or captions, and it publishes nothing. To turn its text into finished, published content, pair it with a generation-and-publishing engine like Kompozy.

How do I turn Step 5 output into finished content?

Use Step 5 for the reasoning and drafting step over a long source, then bring that source into Kompozy. Kompozy fans one input into clips, a persona/avatar short, carousels, quote graphics, a blog, and a newsletter under one Persona Brief, and schedules and publishes them across the eight social platforms plus blog and email.

Related tools

  • Qwen3.8-Omni-FlashAlibaba's omni-modal understanding model, released September 18, 2026 — it reads text, images, audio, and video together in one request, holds a one-million-token context, and is tuned for cheap, agentic comprehension of long recordings.
  • DeepSeek V4 Pro 0813The general-availability build of DeepSeek's flagship model — a 1.6-trillion-parameter mixture-of-experts LLM with a 1M-token context, MIT-licensed weights, and API pricing well under Western frontier models.
  • K2 HorizonMBZUAI's Institute of Foundation Models: a family of six fully open text models from 0.9B to 375B, released with weights, code, and training data under Apache 2.0.
  • Writer Palmyra X6Writer's enterprise flagship agentic model, launched August 13, 2026 as a post-trained variation of the open-source GLM-5.2 — built to run governed, multi-step business tasks at a lower token cost alongside a rebuilt Agent harness.

← All AI tools · Get started →