// AI NEWS · AI RESEARCH

Google DeepMind Argues Video Generators Are Becoming Foundational Models for Understanding Reality

A DeepMind research team says generative video models like Veo 3 already solve vision tasks they were never trained for — evidence, they argue, that video generation is learning an implicit world model on the same trajectory LLMs took for language.

KompozyTurn one idea into a week of content — across every platform, published for you.
Get Started →

2026-07-19 · by Moe Ameen

What happened

Google DeepMind researchers have been making a sustained case that generative video models are not just tools for making clips — they are becoming foundational models for machine perception, learning an implicit "world model" of how reality looks and behaves. The clearest statement of the argument is a paper titled "Video models are zero-shot learners and reasoners," posted to arXiv in September 2025 by a DeepMind team (lead author Thaddäus Wiedemer, with co-authors including Robert Geirhos and Been Kim). Its thesis: video models are on the same path for vision that large language models took for text.

The evidence in the paper is that Veo 3, DeepMind's video generation model, can solve a broad range of visual tasks it was never explicitly trained to do — segmenting objects, detecting edges, editing images, reasoning about physical properties, recognizing what objects can be used for, and simulating tool use. It can even handle early forms of visual reasoning like solving mazes and completing symmetry puzzles. The authors describe the mechanism as "chain-of-frames" reasoning: the model works a problem out step by step across generated frames, the visual analog of the chain-of-thought prompting that unlocked reasoning in language models. The conclusion they draw is that video models are heading toward becoming "unified, generalist vision foundation models."

This sits inside a broader DeepMind position on "world models." DeepMind CEO Demis Hassabis has argued through 2025 and 2026 that language models alone do not truly understand physics, causality, or space, and has pushed world models — systems that simulate and predict real-world dynamics — as a necessary step toward more capable AI. DeepMind's Genie line is the product face of that idea: Genie 3, unveiled in 2025, generates interactive, navigable environments in real time rather than fixed clips. The through-line across the paper, the Genie models, and Hassabis's public comments is that learning to generate video forces a model to learn something real about the structure of the world.

Two honest caveats belong on this. First, these are research findings and internal capabilities, not features a creator can go use — the paper studies what Veo 3 can do in a lab, not a shipped "world model" product. Second, the claim is contested: skeptics argue that scoring well on vision benchmarks is not the same as genuine physical understanding, and that impressive demos can still mask brittle, statistical behavior underneath. The direction is a strong signal about where video AI is going, not a settled fact about what it already understands.

Why it matters for creators

  • It reframes what a video model is. If DeepMind is right, the tools creators already use for clips are early versions of general-purpose vision systems — which is why they keep getting more physically coherent release over release, not just prettier.
  • Physical coherence is the practical payoff. The same world-model learning that lets a model reason about depth and object interactions is what stops generated motion from glitching — better hands, gravity, and continuity in the avatar and B-roll footage creators actually publish.
  • It is a research claim, not a product. There is nothing here to sign up for. The creator takeaway is directional: bet on video models improving fast underneath your workflow, not on any single one being "the" world model today.
  • The debate itself is content. "Is AI video actually starting to understand the world?" is a high-interest question your audience is asking — explaining it credibly, in your voice, is a way to ride the topic without hype.
  • It rhymes with the wider industry pivot. From Genie 3 to OpenAI redirecting Sora research toward "world models," the frontier labs are converging on the same idea — so this framing will keep showing up and is worth understanding now.

How to act on this with Kompozy

There is nothing to install here — this is a research argument about what video models are becoming, not a tool. The move a creator can actually make today is to publish a sharp take on it. "Google DeepMind thinks AI video is quietly learning how reality works — here's what that means" is exactly the kind of timely, high-interest angle that earns reach, and Kompozy turns that one take into a full run of content in your voice. Drop your point of view into the engine and it drafts a Blog Article explaining the world-model idea, a Carousel Post that walks through chain-of-frames reasoning slide by slide, captioned Persona Shorts where your avatar breaks it down to camera, and platform-native Text Posts, then schedules and publishes the set across all nine social platforms plus your blog and newsletter. You cover the story fast, everywhere, without writing each version by hand.

The deeper connection is that Kompozy already runs on the video models this research is about. When a creator generates a HeyGen talking-head Persona Short, a Persona VFX HeyGen hook, or Pexels-backed B-roll, they are using exactly the class of model DeepMind is describing — and as those models inherit the physical coherence the world-model work points at, the finished videos Kompozy assembles get cleaner without the creator changing anything. Kompozy sits above the model layer as the generation-and-publishing engine, so you get the upside of the trend in your output while it plays out in the labs.

Quick takeaways

  • DeepMind's paper "Video models are zero-shot learners and reasoners" (arXiv, September 2025) argues Veo 3 solves vision tasks it was never trained for.
  • "Chain-of-frames" reasoning is the video analog of chain-of-thought — the model reasons across generated frames.
  • It ties into DeepMind's broader "world models" push (Genie 3, Hassabis) toward AI that simulates real-world dynamics.
  • These are research capabilities, not a creator product — the takeaway is directional, not something to sign up for.

Frequently asked questions

What is Google DeepMind arguing about video generation models?

DeepMind researchers argue that generative video models are becoming foundational, general-purpose vision models — learning an implicit "world model" of how reality looks and behaves, rather than just synthesizing clips. Their September 2025 paper "Video models are zero-shot learners and reasoners" shows Veo 3 solving vision tasks it was never trained for (segmentation, edge detection, physical reasoning, even maze-solving), which they read as evidence video models are on the same trajectory large language models took for text.

What is "chain-of-frames" reasoning?

It is DeepMind's term for how a video model appears to reason through a visual problem step by step across the frames it generates — for example, tracing a path through a maze frame by frame. The authors present it as the visual counterpart to "chain-of-thought" reasoning in language models, where the model works through intermediate steps rather than jumping to an answer.

Does this mean AI video now understands the real world?

Not settled. DeepMind's evidence — a single model handling many vision tasks it was never trained for — is a strong signal that video pre-training learns real structure about the world rather than surface patterns. But the claim is contested: skeptics note that passing vision benchmarks is not proof of genuine physical understanding, and that demos can hide brittle behavior. Read it as a credible direction, not a finished fact.

Can creators use these world-model capabilities?

Not directly — these are research findings and internal model capabilities, not a product you can sign up for. The practical benefit reaches creators indirectly: as video models get better at modeling physics and geometry, the avatar video, B-roll, and generated footage in everyday tools become more coherent. A generation-and-publishing engine like Kompozy already turns today's video models into finished, scheduled posts across platforms, so you inherit those improvements as they ship.

Related news

← All AI news · Get started →