The August 1, 2026 update lets creators lock up to seven elements — a face, a product, a location — across a generated scene, and generate straight from a text prompt.
2026-08-05 · by Moe Ameen
On August 1, 2026, xAI shipped a significant update to Grok Imagine Video 1.5, the video half of Grok Imagine inside Grok. The headline addition is reference-based generation: you can now pass in up to seven reference images in a single generation, and each one locks a single thing in place — a person's face, a product, or a location. That gives the model the control layer generic text-to-video lacks. Keep a character and swap the scene, keep the scene and swap the character, or hold both and change only the action.
The update also adds a voice reference. Supply a voice sample alongside a character image and the same face and the same voice carry across every scene, which is the piece that turns a one-off clip into a repeatable on-camera identity. Alongside references, xAI added native 1080p output and prompt-only text-to-video, so you can generate a clip straight from a written description with no starting image. Grok Imagine Video 1.5 already generates synchronized native audio — dialogue, sound effects, and music — in the same pass rather than adding it afterward.
Image references, text-to-video, and native 1080p are available through both the Grok Imagine website and the xAI API, where the model id is `grok-imagine-video-1.5`; voice references require a request to xAI. The consumer rollout started in the United States for SuperGrok Heavy and SuperGrok Plus subscribers on the Imagine site and apps, with xAI saying the tools would reach all tiers over the following days. The model first launched as a preview in late May 2026 and reached general availability in mid-June, where it topped the public Image-to-Video Arena leaderboard and undercut higher-end rivals on price. xAI iterates quickly, so treat the tiers, resolutions, and per-second API pricing as a launch snapshot and confirm them on xAI's own pages before quoting.
The real unlock here is repeatability. Before this update, an AI-generated character drifted between runs, so you could not build a recurring on-camera identity out of Grok. Now a locked face plus a voice reference holds steady across scenes, and text-to-video means you can spin those scenes up without ever filming. What you still end up with, though, is a stack of 1080p clips with baked-in audio sitting in a downloads folder — no captions, no vertical reframe, no hook, and nothing scheduled. That gap between "a consistent clip" and "a published series" is exactly what [Kompozy](/) closes. Bring your reference-locked clips in and [Clipped Shorts](/glossary/output-buckets) cuts them into 9:16 segments with word-synced branded captions, while your [Persona Brief](/glossary/persona-shorts) governs the copy so every post reads in one voice. Because the on-screen identity is now stable, the same character can anchor a recurring format that ships weekly — and Kompozy fans each idea into a Carousel rendered pixel-exact through HyperFrames, Quote Graphics, a Photo Post, a Blog Article, and an Email Newsletter, then schedules and publishes the whole set across the eight social platforms plus blog and email on [Autopilot](/glossary/autopilot). Grok holds the character; Kompozy runs the operation.
There is also a same-week content play for anyone who covers AI video, creator tooling, or agency work. "Grok Imagine can now hold a face and a voice across a whole scene, in 1080p, from a text prompt" is a high-intent question right now, and the nuance most quick recaps skip — that references are what separate this from generic text-to-video, and that the rollout is tier-gated — is the useful part. Drop your take on the update into Kompozy as a source and it fans out a blog post for the searchers, a carousel breaking down reference control versus plain text-to-video, a captioned explainer clip, and platform-native posts in your voice, queued across your channels in one pass while the story is still current. See our deeper [Grok Imagine Video 1.5 breakdown](/ai-tools/grok-imagine-video-1-5) for how the model fits a full production workflow.
On August 1, 2026, xAI added reference-based generation (up to seven reference images per generation, each locking a face, product, or location), a voice reference for face-and-voice consistency across scenes, native 1080p output, and prompt-only text-to-video. The model already generated synchronized native audio in the same pass.
You pass in up to seven reference images and each one locks a single element in place. That lets you keep a character and change the scene, keep the scene and change the character, or hold both and change only the action — the control layer generic text-to-video lacks. A voice reference can travel alongside a character image so the same face and voice carry across scenes.
The consumer rollout started in the United States for SuperGrok Heavy and SuperGrok Plus subscribers on the Grok Imagine site and apps, with xAI saying the tools would reach all tiers over the following days. Image references, text-to-video, and 1080p are also available through the xAI API using the grok-imagine-video-1.5 model; voice references require a request to xAI. Confirm current tiers on xAI before budgeting.
Grok generates the reference-consistent clip; Kompozy is the finishing and distribution layer. Bring the clip in and Kompozy cuts vertical captioned shorts, reframes for each feed, spins the same idea into a carousel, quote graphics, a blog, and a newsletter in your Persona Brief voice, then schedules and publishes across the eight social platforms plus blog and email.