// AI NEWS · FEATURE

Google Rolls Out AI Image and Video Tools: Pics Creates, Agentic Video Understands

On September 1, 2026, Google added two AI capabilities under its "image and video tools" banner — Pics, a prompt-first image generator, and agentic video understanding in Gemini. Only one of them actually makes new visual content.

2026-09-02 · by Moe Ameen

What happened

On September 1, 2026, Google rolled out two new AI capabilities it grouped together as image and video tools. The first is Pics, an image generation and editing tool tied to Google's paid AI Pro tier (and Workspace/Ultra). The second is agentic video understanding, an upgrade to how Gemini reads video, which launched via the Gemini API and AI Studio at standard developer token pricing rather than behind a subscription tier. Both began a phased rollout that reaches more users over the following weeks.

Pics generates and edits still images from a prompt, powered by Google's Nano Banana model. It first appears inside Google Workspace — Docs and Slides — with Drive to follow, and is available to Workspace customers plus people who pay for Google AI Pro or Ultra. Instead of starting from a template, you describe what you want and Pics returns options; editing is prompt- and click-driven, treating each element as its own object so you can select, transform, or even translate the text inside an image. Google positions it against Canva and Adobe Express. (For the full breakdown, see our report on the [Pics launch](/news/google-pics-launch).)

Agentic video understanding is the video half, and it is where the "video tools" label needs care. It does not generate video. It changes how Gemini analyzes footage: rather than scanning a clip at a fixed frame rate, the model decides what to watch, at what speed, and through which channel — frames, audio, or transcript — and fetches only the segments it needs to answer a question. Google reports up to 88% fewer tokens, up to 66% lower cost, and up to 7% better accuracy versus fixed-rate scanning, with the gains most pronounced on long-form video from 10-minute how-tos to multi-hour recordings. It runs on Gemini 3.7 Flash, 3.6 Flash, and 3.5 Flash-Lite. At launch it is live via the Gemini API in Google AI Studio and the Gemini Enterprise Agent Platform as a Generative AI Preview at standard token pricing with no extra feature fee, with the Gemini app coming soon and YouTube's Ask YouTube feature slated for the coming months.

So the honest summary is that Google shipped one creation tool and one comprehension tool on the same day. Pics makes images. Agentic video understanding reads video faster and cheaper. Neither generates video, clips a long recording into shorts, or produces a talking-head or avatar clip. Treat specific availability, tiers, and figures as launch-day snapshots and confirm them on Google's own pages before relying on them.

Why it matters for creators

  • The "image and video tools" framing hides an asymmetry: image creation got easier, but the video side is analysis, not generation. If you were hoping this rollout lets you make video from a prompt, it does not — that gap is still yours to fill.
  • Agentic video understanding is genuinely useful for creators sitting on long footage: it can find a specific moment, count things, or spot an anomaly in a 90-minute recording for a fraction of the old token cost, which is the hard part of turning raw video into clips.
  • But finding the moment is not shipping the clip. Understanding where the good 30 seconds live in your webinar does not cut, caption, reframe to vertical, or publish it — that is still separate, manual work.
  • Pics output is a single static image inside a Google doc or slide. Getting it onto TikTok, Reels, LinkedIn, or a blog as a captioned, scheduled post is downstream work the tool does not do.
  • The pattern to read: the big platforms are commoditizing generation and comprehension of single assets inside the apps you already pay for. The remaining edge for creators is consistency across formats and actually distributing the output.

How to act on this with Kompozy

This rollout is a clean illustration of where the platforms stop. Google can now help you find the right 30 seconds inside a two-hour recording, and it can spin up a still image from a sentence. What it can't do is turn either into finished video posts across your feeds — there is no clipping, no captioning, no avatar or talking-head generation, no scheduling. That is the whole job that sits between "I understand my footage" and "the clips are live," and it's exactly what [Kompozy](/) is built for.

Point Kompozy at a long-form video and it does the part agentic understanding stops short of: cuts it into vertical [Clipped Shorts](/glossary/content-repurposing), burns in captions, reframes for each platform, and queues the set. When you need net-new video that Google's tools can't produce at all, Kompozy generates it — [Persona Shorts](/glossary/persona-shorts) and HeyGen avatar cuts that deliver your script on camera, Listicle Video over a portrait clip. And a Pics still becomes raw material for a brand-exact [Carousel](/glossary/hyperframes), Quote Graphics, a blog, and a newsletter, all held to one [Persona Brief](/glossary/persona-brief) so the batch reads as one brand. Then [Autopilot](/glossary/autopilot) schedules and publishes the whole set across the eight social platforms plus blog and email behind a per-post review. Google reads your video and draws your picture; Kompozy turns both into a week of published content.

Quick takeaways

  • On Sep 1, 2026, Google rolled out two things under "image and video tools": Pics (image creation, tied to AI Pro/Ultra/Workspace) and agentic video understanding (video analysis, launched via the Gemini API).
  • Pics generates and edits still images from a prompt via the Nano Banana model, inside Google Workspace Docs and Slides first, for Workspace, AI Pro, and Ultra users.
  • Agentic video understanding does NOT generate video — it lets Gemini scan footage selectively, reporting up to 88% fewer tokens, 66% lower cost, and 7% better accuracy on long-form video.
  • It runs on Gemini 3.7 Flash, 3.6 Flash, and 3.5 Flash-Lite; live now via the Gemini API and AI Studio, with the Gemini app and Ask YouTube to follow.
  • Neither tool clips long video into shorts, makes avatar or talking-head video, or publishes across platforms — Kompozy does that end to end.

Frequently asked questions

Did Google launch a video generator on September 1, 2026?

No. The video half of the rollout is agentic video understanding, which analyzes and retrieves moments from existing footage more cheaply and accurately. It does not generate video. The image half, Pics, does create new still images from a prompt.

What is agentic video understanding and which models run it?

It is a mode where Gemini decides what parts of a video to watch — frames, audio, or transcript — instead of scanning at a fixed frame rate, cutting cost and improving accuracy on long footage. Google reports up to 88% fewer tokens, 66% lower cost, and 7% better accuracy versus fixed-rate scanning. It runs on Gemini 3.7 Flash, 3.6 Flash, and 3.5 Flash-Lite.

How much do these tools cost?

Pics is tied to Google Workspace subscriptions plus Google AI Pro and Ultra. Agentic video understanding launched via the Gemini API and AI Studio at standard Gemini token pricing with no additional feature fee. Confirm current tiers and rates on Google's own pages.

Can these tools turn my long video into short clips for TikTok or Reels?

Not on their own. Agentic video understanding can find the right moment in a long recording, but it does not cut, caption, reframe to vertical, or publish it. A content engine like Kompozy handles that clipping-to-published workflow, and can also generate net-new avatar and listicle video the Google tools do not make.

Related news

← All AI news · Get started →