Here is the part most "make videos with ChatGPT" guides bury: ChatGPT does not render video. It is a language model with an image generator bolted on, so it writes scripts, plans scenes, and draws stills — but there is no button that turns a prompt into a moving clip. OpenAI's actual video model, Sora, was the thing that did that, and OpenAI is shutting it down: the app and website closed on April 26, 2026 and the developer API is scheduled to end on September 24, 2026. So in 2026 there is no native ChatGPT video generation at all, and any workflow that claims otherwise is either using an old Sora integration that no longer exists or quietly handing the render step to a different tool.
That does not make ChatGPT useless for video — it makes it a pre-production tool. It is genuinely fast at the parts before the camera rolls: the script, the shot list, the storyboard, reference images, and the prompts you feed a real video generator. This guide walks the honest end-to-end pipeline — what ChatGPT does, where you hand off to a render tool, and how to caption, size, and publish the finished clip — so you leave with a video that posts, not a chat transcript. For the background on the capability itself, see [Can ChatGPT make videos?](/guides/can-chatgpt-make-videos) and the [Sora shutdown write-up](/news/openai-sora-shutdown).
The steps
Set the expectation: ChatGPT plans the video, it does not render it. Before you start, be clear on the boundary so you don't waste an hour prompting for a clip that will never come. ChatGPT writes text and generates static images through GPT-Image; it cannot produce a moving video, and with Sora being wound down there is no OpenAI consumer video generator to fall back on. Every workflow that "makes a video with ChatGPT" is really ChatGPT doing pre-production and a separate tool doing the render.
Write the script in ChatGPT — as spoken lines, not an essay. Give ChatGPT the audience, angle, platform, target length, and the one action you want, then ask for short spoken lines under 20 words each with a hook in the first few seconds. The specificity you put in is the specificity you get back; a one-line prompt yields generic filler. Our [ChatGPT video-script guide](/how-to/use-chatgpt-to-write-video-scripts) has the reusable prompt template and the word-count-to-runtime math.
Turn the script into a shot list and storyboard. Ask ChatGPT to break the script into scenes: one line per shot with the on-screen action, the b-roll or visual, and any text overlay. This is where it earns its keep — it converts a wall of narration into a shootable plan in seconds. Request a storyboard table (scene number, spoken line, visual, duration) so you know exactly what footage each moment needs.
Generate reference stills and thumbnails with GPT-Image. ChatGPT can render static images, so use it for the things that are actually images: a thumbnail, concept frames, a title card, or reference art to hand a video model. It will not draw a consistent character across frames — image models drift — so treat these as references and mood, not final video frames. A clear thumbnail is often the single highest-leverage still you make here.
Have ChatGPT write the render prompts for your video tool. Whatever generator you use next, it needs a prompt, and ChatGPT is good at translating a scene into one. Ask it to write per-shot prompts in the format your tool expects — a text-to-video scene description for a cinematic model, or a spoken script for an avatar tool. Tell it the tool by name so it matches the right shape: a talking-head avatar needs a clean script, while a model like [Runway](/ai-tools/runway), [Veo 3](/ai-tools/veo-3), or [Kling](/ai-tools/kling-ai) needs a visual scene description.
Render the actual video in a real generator. This is the hand-off ChatGPT can't do. For a talking-head or faceless narration video, an avatar tool like [HeyGen](/ai-tools/heygen) turns your script into a presenter clip with a voice; for cinematic b-roll, a text-to-video model renders the scenes from your prompts. Pick the tool that matches your format — avatar for explainers and faceless shorts, text-to-video for visual, scene-driven pieces — and render each shot from the plan ChatGPT built.
Caption, reframe, and finish the clip. A raw render is not a postable video. Add burned-in captions (most short-form is watched on mute), reframe to 9:16 for TikTok, Reels, and Shorts or 1:1 and 16:9 where those fit, and cut the dead air. This post-production step is entirely outside ChatGPT — it wrote the words and drew a still, but sizing, subtitles, and the final cut are on you or your editing tool.
Publish it — and fan the same script into the rest of the week. Upload to each platform, or use a scheduler so one clip lands everywhere on cadence. While you're there, the script you wrote is also a carousel, a text post, a blog section, and a newsletter — so don't let it produce a single video and stop. Getting more than one asset out of one scripting session is the difference between a one-off and a content habit.
Common gotchas
ChatGPT cannot make a video, full stop. If a tutorial shows a "generate video" button inside ChatGPT, it is either out of date (an old Sora integration) or the render is silently happening in a third-party plugin, not in ChatGPT.
Sora is not a fallback anymore. The app closed April 26, 2026 and the API ends September 24, 2026 — do not build a 2026 workflow that depends on it.
GPT-Image makes stills, not frames. It will not hold the same face or scene across shots, so images from ChatGPT are references and thumbnails, never a substitute for a video model.
Fact-check every specific ChatGPT writes into the script. It fabricates confident stats, studies, and product details exactly where you most need them right, and those go on camera if you don't catch them.
The render tool renders exactly what you feed it. A vague script or a lazy shot list produces a vague video — the quality is decided in the ChatGPT pre-production step, not at render time.
Skipping captions and reframing is the most common reason an AI video underperforms. The clip can be great and still flop if it's letterboxed and silent in a vertical feed.
Where Kompozy fits
Count the tools this guide actually needed: ChatGPT for the script, ChatGPT again for the shot list, ChatGPT again for the render prompts, a separate avatar or text-to-video generator, a captioning tool, a reframing tool, and a scheduler. That is the real cost of "making a video with ChatGPT" — a chain of six or seven handoffs where you re-explain your brand at every link, and where the video generation you started out looking for turns out to be the one thing ChatGPT can't do at all.
Kompozy collapses steps two through eight of this workflow into one pass. You give it the same topic you'd have briefed ChatGPT with, and the engine writes the script under a [Persona Brief](/glossary/persona-brief) that already holds your voice, audience, and banned words — then renders the actual video. A [Persona Short](/glossary/persona-shorts) becomes a captioned talking-head clip delivered by a face-locked avatar; a Listicle or Marketing Short composites footage and cards; and if you bring your own generated clip from Runway, Veo, or Kling, Kompozy clips, captions, reframes, and finishes it. The render step ChatGPT hands off to a fourth tool is native here, and the captions and 9:16/1:1/16:9 reframing that were manual are automatic.
Then it publishes — to nine platforms plus blog and email behind a per-post review gate — and fans the same script into a carousel, quote graphic, text post, blog article, and newsletter, so one scripting session becomes a week of posts instead of a single clip. Honest framing: if you make a video every now and then and like directing each tool by hand, ChatGPT plus a render app gives you full manual control for less. Kompozy earns its place when you want that script rendered, captioned, on-brand, and published without stitching the chain together yourself every time — Starter ($99/mo for 5,500 credits) for a solo creator on a steady cadence, Pro ($299/mo for 18,000 credits) for high-volume multi-format output, Enterprise custom for teams.
Frequently asked questions
Can ChatGPT make videos in 2026?
No — not on its own. ChatGPT is a language model that writes text and generates static images; it does not render video. OpenAI's video model, Sora, was the tool that did that, and it is being discontinued (app closed April 26, 2026; API ends September 24, 2026). To make a video you use ChatGPT for the script and plan, then a separate video generator for the render.
Does ChatGPT have a video generator built in?
No. There is no native video-generation feature in ChatGPT as of 2026. It generates images via GPT-Image but has no text-to-video model of its own available to consumers now that Sora is winding down. Any "ChatGPT video" pipeline hands the actual rendering to another tool.
What is the best tool to turn a ChatGPT script into a video?
It depends on the format. For talking-head or faceless narration, an avatar generator like HeyGen turns the script into a presenter clip with a voice. For cinematic, scene-driven video, a text-to-video model like Runway, Veo, or Kling renders from ChatGPT-written prompts. A content engine like Kompozy can also render the script as an avatar video and then caption and publish it in one pass.
Can I use ChatGPT and Sora together to make a video?
Not reliably going forward. Sora's consumer app already closed on April 26, 2026 and the API is scheduled to end on September 24, 2026, so a Sora-based workflow has a hard deadline. Use a live video generator for the render instead of building on a product that is being shut down.
How long does it take to make a video this way?
The ChatGPT parts — script, shot list, render prompts, a thumbnail — take about ten to twenty minutes with a good prompt. The render and post-production (captions, reframing, cutting) is the longer half and depends on your tool. Collapsing the render, captioning, and publishing into one engine is where most of the time is saved.