On September 1, 2026, Google added two AI capabilities under its "image and video tools" banner — Pics, a prompt-first image generator, and agentic video understanding in Gemini. Only one of them actually makes new visual content.
2026-09-02 · by Moe Ameen
On September 1, 2026, Google rolled out two new AI capabilities it grouped together as image and video tools. The first is Pics, an image generation and editing tool tied to Google's paid AI Pro tier (and Workspace/Ultra). The second is agentic video understanding, an upgrade to how Gemini reads video, which launched via the Gemini API and AI Studio at standard developer token pricing rather than behind a subscription tier. Both began a phased rollout that reaches more users over the following weeks.
Pics generates and edits still images from a prompt, powered by Google's Nano Banana model. It first appears inside Google Workspace — Docs and Slides — with Drive to follow, and is available to Workspace customers plus people who pay for Google AI Pro or Ultra. Instead of starting from a template, you describe what you want and Pics returns options; editing is prompt- and click-driven, treating each element as its own object so you can select, transform, or even translate the text inside an image. Google positions it against Canva and Adobe Express. (For the full breakdown, see our report on the [Pics launch](/news/google-pics-launch).)
Agentic video understanding is the video half, and it is where the "video tools" label needs care. It does not generate video. It changes how Gemini analyzes footage: rather than scanning a clip at a fixed frame rate, the model decides what to watch, at what speed, and through which channel — frames, audio, or transcript — and fetches only the segments it needs to answer a question. Google reports up to 88% fewer tokens, up to 66% lower cost, and up to 7% better accuracy versus fixed-rate scanning, with the gains most pronounced on long-form video from 10-minute how-tos to multi-hour recordings. It runs on Gemini 3.7 Flash, 3.6 Flash, and 3.5 Flash-Lite. At launch it is live via the Gemini API in Google AI Studio and the Gemini Enterprise Agent Platform as a Generative AI Preview at standard token pricing with no extra feature fee, with the Gemini app coming soon and YouTube's Ask YouTube feature slated for the coming months.
So the honest summary is that Google shipped one creation tool and one comprehension tool on the same day. Pics makes images. Agentic video understanding reads video faster and cheaper. Neither generates video, clips a long recording into shorts, or produces a talking-head or avatar clip. Treat specific availability, tiers, and figures as launch-day snapshots and confirm them on Google's own pages before relying on them.
This rollout is a clean illustration of where the platforms stop. Google can now help you find the right 30 seconds inside a two-hour recording, and it can spin up a still image from a sentence. What it can't do is turn either into finished video posts across your feeds — there is no clipping, no captioning, no avatar or talking-head generation, no scheduling. That is the whole job that sits between "I understand my footage" and "the clips are live," and it's exactly what [Kompozy](/) is built for.
Point Kompozy at a long-form video and it does the part agentic understanding stops short of: cuts it into vertical [Clipped Shorts](/glossary/content-repurposing), burns in captions, reframes for each platform, and queues the set. When you need net-new video that Google's tools can't produce at all, Kompozy generates it — [Persona Shorts](/glossary/persona-shorts) and HeyGen avatar cuts that deliver your script on camera, Listicle Video over a portrait clip. And a Pics still becomes raw material for a brand-exact [Carousel](/glossary/hyperframes), Quote Graphics, a blog, and a newsletter, all held to one [Persona Brief](/glossary/persona-brief) so the batch reads as one brand. Then [Autopilot](/glossary/autopilot) schedules and publishes the whole set across the eight social platforms plus blog and email behind a per-post review. Google reads your video and draws your picture; Kompozy turns both into a week of published content.
No. The video half of the rollout is agentic video understanding, which analyzes and retrieves moments from existing footage more cheaply and accurately. It does not generate video. The image half, Pics, does create new still images from a prompt.
It is a mode where Gemini decides what parts of a video to watch — frames, audio, or transcript — instead of scanning at a fixed frame rate, cutting cost and improving accuracy on long footage. Google reports up to 88% fewer tokens, 66% lower cost, and 7% better accuracy versus fixed-rate scanning. It runs on Gemini 3.7 Flash, 3.6 Flash, and 3.5 Flash-Lite.
Pics is tied to Google Workspace subscriptions plus Google AI Pro and Ultra. Agentic video understanding launched via the Gemini API and AI Studio at standard Gemini token pricing with no additional feature fee. Confirm current tiers and rates on Google's own pages.
Not on their own. Agentic video understanding can find the right moment in a long recording, but it does not cut, caption, reframe to vertical, or publish it. A content engine like Kompozy handles that clipping-to-published workflow, and can also generate net-new avatar and listicle video the Google tools do not make.