Rolling out globally from September 24, 2026, the Search type filter now splits "Web" into Text-based and Multimodal, so you can finally see impressions and clicks from visual searches — Lens, Circle to Search, camera and screenshot lookups. There is one catch: no query data.
2026-09-25 · by Moe Ameen
On September 24, 2026, Google began rolling out globally a new Multimodal filter in Search Console's Performance report. Opening the Search type filter now splits the "Web" option into two: Text-based and Multimodal. Per Google's documentation, "Web-multimodal tracks search results triggered by a query that uses an image, photo, or screenshot. Text-only queries are tracked as Web text-based." In other words, the searches where someone pointed a camera, uploaded a screenshot, or circled something on their screen are now broken out from ordinary typed queries.
The multimodal bucket covers searches made with Google Lens, Circle to Search on Android, image uploads to Google Search, and Chrome's right-click "Search this image." Google product managers Harsh Kharbanda and Moshe Samet framed the change as visibility into a discovery surface that was previously invisible: "This update is designed to give you insights into how your content is surfaced when users search using images (such as with a smartphone camera)." The split appears in the standard Search results performance report and also feeds the generative AI performance report.
There is a real limitation to understand before reading anything into the numbers. You get impressions, clicks, and position for multimodal traffic, sliceable by page, country, and device — but you do not get the queries. Google's stated reason: "Because multimodal searches mostly use images rather than text, specific text query data isn't available for this traffic." So the classic keyword-to-content loop that text search enables does not work here. The generative AI report, separately, still lacks click and query data of its own. Treat availability as a rollout-window snapshot and confirm what has landed in your own Search Console before drawing conclusions from a single week.
This filter quietly confirms a shift most content strategies still ignore: a growing share of discovery is visual, not textual. People aim a camera at a product, screenshot a post, or circle an object on screen — and Google matches that against the actual images on your pages. The uncomfortable implication is that being findable in multimodal search is a production problem, not a keyword problem. There is nothing to optimize for a phrase, because there is no phrase; there is only whether you have published enough strong, distinctive, on-brand imagery for a visual engine to surface. That is the exact ceiling [Kompozy](/) is built to lift. It is not a text-first tool bolted onto images — it generates visual content as a first-class output: scene photo posts, infographic posters, face-locked [Persona Photos](/glossary/output-buckets), brand-exact [Carousels](/glossary/hyperframes), quote graphics, and short video, all in your look, then [Autopilot](/glossary/autopilot) schedules and publishes them across the eight social platforms plus blog and email behind a per-post review gate.
The angle that fits this specific news is visual volume with brand consistency. Because multimodal search returns no query data, you cannot tune toward keywords — the winning move is to produce many clear, distinctive, first-party images across your topics so a visual match has something to find, and to keep every one of them recognizably yours so exposure compounds into a brand people recall. A [Persona Brief](/glossary/persona-brief) plus HyperFrames hold that look across the whole set, which is what separates a match-worthy original image from interchangeable stock. Then the new Multimodal filter becomes your feedback loop: watch which pages earn image-driven impressions and make more of what works. The honest boundary — Kompozy cannot make Lens rank you, and no tool can fake a distinctive visual; matches are earned by publishing genuinely good, original imagery. What it removes is the production ceiling that keeps most creators from ever making enough visual content to be found this way. For the deeper playbook, see [image SEO for AI-powered search](/guides/image-seo-for-ai-powered-search), [AI search impressions in Google](/guides/ai-search-impressions-in-google), and [how to set up AI search performance reporting in Search Console](/how-to/set-up-ai-search-performance-reporting-in-search-console).
It is a new option in the Search Console Performance report, rolling out globally from September 24, 2026, that splits the "Web" search type into Text-based and Multimodal. Google defines it as tracking "search results triggered by a query that uses an image, photo, or screenshot," while text-only queries stay under Web text-based. It covers Google Lens, Circle to Search on Android, image uploads to Google Search, and Chrome's "Search this image."
Because these searches are made with images rather than typed text. Google's stated reason is that "because multimodal searches mostly use images rather than text, specific text query data isn't available for this traffic." You can see impressions, clicks, and position for multimodal traffic, sliceable by page, country, and device — but not the queries that produced them, so the usual keyword-to-content loop does not apply.
Since there are no keywords to target, the lever is producing a high volume of clear, distinctive, on-brand first-party images that a visual engine like Lens can match against your pages, and then watching which pages earn multimodal impressions in the new filter. Original, consistent imagery beats generic stock or templated graphics. A content engine like Kompozy generates that visual set — photo posts, carousels, infographics, persona images, and short video — under one brand brief and publishes it across platforms.
Yes. The Text-based and Multimodal split appears in both the standard Search results performance report and the generative AI performance report. Note that the generative AI report separately still lacks click and query data of its own, so read the multimodal split there as visibility into which content surfaces, not as a complete performance picture.