// AI NEWS · AI VIDEO

xAI Updates Grok Imagine Video 1.5 With Seven-Reference Scene Control, Voice Consistency, Native 1080p, and Text-to-Video

The August 1, 2026 update lets creators lock up to seven elements — a face, a product, a location — across a generated scene, and generate straight from a text prompt.

2026-08-05 · by Moe Ameen

What happened

On August 1, 2026, xAI shipped a significant update to Grok Imagine Video 1.5, the video half of Grok Imagine inside Grok. The headline addition is reference-based generation: you can now pass in up to seven reference images in a single generation, and each one locks a single thing in place — a person's face, a product, or a location. That gives the model the control layer generic text-to-video lacks. Keep a character and swap the scene, keep the scene and swap the character, or hold both and change only the action.

The update also adds a voice reference. Supply a voice sample alongside a character image and the same face and the same voice carry across every scene, which is the piece that turns a one-off clip into a repeatable on-camera identity. Alongside references, xAI added native 1080p output and prompt-only text-to-video, so you can generate a clip straight from a written description with no starting image. Grok Imagine Video 1.5 already generates synchronized native audio — dialogue, sound effects, and music — in the same pass rather than adding it afterward.

Image references, text-to-video, and native 1080p are available through both the Grok Imagine website and the xAI API, where the model id is `grok-imagine-video-1.5`; voice references require a request to xAI. The consumer rollout started in the United States for SuperGrok Heavy and SuperGrok Plus subscribers on the Imagine site and apps, with xAI saying the tools would reach all tiers over the following days. The model first launched as a preview in late May 2026 and reached general availability in mid-June, where it topped the public Image-to-Video Arena leaderboard and undercut higher-end rivals on price. xAI iterates quickly, so treat the tiers, resolutions, and per-second API pricing as a launch snapshot and confirm them on xAI's own pages before quoting.

Why it matters for creators

  • Consistency was the wall. Generic text-to-video re-draws your subject slightly differently every run, which is fatal for a brand. Locking a face, a product, or a style across generations is what makes AI video usable for recurring content, not just a novelty clip.
  • A voice reference plus a character image gives you a repeatable on-camera persona — the same face and voice across every scene — without filming. That is the difference between one clip and a series.
  • Text-to-video means you no longer need source footage or a starting image to generate. A written prompt is enough, which lowers the barrier for creators who do not shoot.
  • Native 1080p closes part of the quality gap that kept earlier Grok clips out of a polished feed, though a raw generated clip still has no captions, no hook frame, and no platform formatting.
  • The update is gated at launch — US-first, top SuperGrok tiers, voice by request — so what you can actually do this week depends on your tier. Plan around today, not the rollout roadmap.

How to act on this with Kompozy

The real unlock here is repeatability. Before this update, an AI-generated character drifted between runs, so you could not build a recurring on-camera identity out of Grok. Now a locked face plus a voice reference holds steady across scenes, and text-to-video means you can spin those scenes up without ever filming. What you still end up with, though, is a stack of 1080p clips with baked-in audio sitting in a downloads folder — no captions, no vertical reframe, no hook, and nothing scheduled. That gap between "a consistent clip" and "a published series" is exactly what [Kompozy](/) closes. Bring your reference-locked clips in and [Clipped Shorts](/glossary/output-buckets) cuts them into 9:16 segments with word-synced branded captions, while your [Persona Brief](/glossary/persona-shorts) governs the copy so every post reads in one voice. Because the on-screen identity is now stable, the same character can anchor a recurring format that ships weekly — and Kompozy fans each idea into a Carousel rendered pixel-exact through HyperFrames, Quote Graphics, a Photo Post, a Blog Article, and an Email Newsletter, then schedules and publishes the whole set across the eight social platforms plus blog and email on [Autopilot](/glossary/autopilot). Grok holds the character; Kompozy runs the operation.

There is also a same-week content play for anyone who covers AI video, creator tooling, or agency work. "Grok Imagine can now hold a face and a voice across a whole scene, in 1080p, from a text prompt" is a high-intent question right now, and the nuance most quick recaps skip — that references are what separate this from generic text-to-video, and that the rollout is tier-gated — is the useful part. Drop your take on the update into Kompozy as a source and it fans out a blog post for the searchers, a carousel breaking down reference control versus plain text-to-video, a captioned explainer clip, and platform-native posts in your voice, queued across your channels in one pass while the story is still current. See our deeper [Grok Imagine Video 1.5 breakdown](/ai-tools/grok-imagine-video-1-5) for how the model fits a full production workflow.

Quick takeaways

  • xAI updated Grok Imagine Video 1.5 on August 1, 2026 with reference-based generation, voice consistency, native 1080p, and prompt-only text-to-video.
  • Up to seven reference images per generation each lock one element — a face, product, or location — for consistent characters and scenes.
  • A voice reference paired with a character image keeps the same face and voice across scenes; voice references require a request to xAI.
  • Image references, text-to-video, and 1080p are available via the Grok Imagine site and the xAI API (model `grok-imagine-video-1.5`); the consumer rollout started US-first for top SuperGrok tiers.
  • Grok generates the consistent clip; Kompozy captions, reframes, repurposes, and publishes it across eight social platforms plus blog and email.

Frequently asked questions

What did the August 2026 Grok Imagine Video 1.5 update add?

On August 1, 2026, xAI added reference-based generation (up to seven reference images per generation, each locking a face, product, or location), a voice reference for face-and-voice consistency across scenes, native 1080p output, and prompt-only text-to-video. The model already generated synchronized native audio in the same pass.

How do reference images work in Grok Imagine Video 1.5?

You pass in up to seven reference images and each one locks a single element in place. That lets you keep a character and change the scene, keep the scene and change the character, or hold both and change only the action — the control layer generic text-to-video lacks. A voice reference can travel alongside a character image so the same face and voice carry across scenes.

Who can use the new Grok Imagine Video 1.5 features?

The consumer rollout started in the United States for SuperGrok Heavy and SuperGrok Plus subscribers on the Grok Imagine site and apps, with xAI saying the tools would reach all tiers over the following days. Image references, text-to-video, and 1080p are also available through the xAI API using the grok-imagine-video-1.5 model; voice references require a request to xAI. Confirm current tiers on xAI before budgeting.

How does Kompozy fit with Grok Imagine Video 1.5?

Grok generates the reference-consistent clip; Kompozy is the finishing and distribution layer. Bring the clip in and Kompozy cuts vertical captioned shorts, reframes for each feed, spins the same idea into a carousel, quote graphics, a blog, and a newsletter in your Persona Brief voice, then schedules and publishes across the eight social platforms plus blog and email.

Related news

← All AI news · Get started →