An analysis of 1,249 ChatGPT answers found OpenAI's in-house search index served small, unlicensed websites the same way it served licensed publishing partners — same formatting, length, and freshness — walking back an earlier claim that the index favored an allowlist of established names.
2026-08-12 · by Moe Ameen
New data suggests ChatGPT Search can surface and index content from smaller websites, not just established publishers. The French SEO consultancy Resoneo analyzed 1,249 ChatGPT answers captured in July 2026 and reviewed the cited pages behind them. For a window before OpenAI changed the behavior, ChatGPT's server stream tagged each web result with the name of the pipeline that fetched it — and one label, reported as "labrador," corresponded to OpenAI's own in-house search index rather than a third-party search API. That tag let analysts see which citations came from OpenAI's index versus an external fetch.
The finding: there was no meaningful difference in how that index treated unlicensed sites versus licensed publishing partners. Pages from small, independent sites — including small Italian sites flagged by one reader using a free account — were stored and served the same way as pages from outlets with formal OpenAI content deals: comparable formatting, length, and freshness. The index appears to store a page's title and a short snippet (on the order of a couple hundred characters) of its content.
The data backs a public correction. In June, analyst Suganthan Mohanadasan had described OpenAI's index as an "allowlist of established publishers." On July 14, 2026, he retracted that characterization after receiving captures showing every publisher citation ran through the same pipeline, including small sites, and said he had "over-reached" with the tiered-index claim. Caveats travel with this: the original counts were "directional" from a single account, and around July 21, 2026, OpenAI reportedly stopped tagging search results with pipeline identifiers, so this exact visibility into the mechanics may be temporary. Treat the pipeline-label specifics as a snapshot and confirm current behavior before relying on it.
The timely post writes itself: "New data says ChatGPT Search indexes small sites like yours — here's what actually gets you cited." Your audience is anxious that AI search is a closed shop for big publishers, and a clear-eyed explainer that separates *indexed* from *cited* is exactly the useful, early take that travels. Drop the facts into [Kompozy](/) and they become a [Blog Article](/glossary/output-buckets) breaking down the study, a [Carousel](/glossary/output-buckets) of the "indexed ≠ cited" checklist, [Quote Graphics](/glossary/output-buckets) of the sharpest lines, and a batch of captioned shorts — all held to your voice by the [Persona Brief](/glossary/persona-brief) and scheduled across the eight social platforms plus blog and email with [Autopilot](/glossary/autopilot).
The strategic move is bigger than one post. If a small site is genuinely eligible to be indexed and surfaced, the constraint becomes output: a thin site with three pages gives an AI index almost nothing to pull from, while a site publishing specific, answer-shaped content on a real cadence gives it a lot. That volume-without-a-team problem is what Kompozy exists to solve — it generates net-new blogs, carousels, images, persona/avatar video, and newsletters (not just repurposed clips) and publishes them to an owned blog plus eight social platforms, building the consistent, structured footprint that [Generative Engine Optimization](/glossary/generative-engine-optimization) rewards. Kompozy won't force ChatGPT to cite you, but it removes the reason most small sites never get considered: they don't publish enough useful pages, often enough, for any index to notice.
New data suggests yes. A July 2026 analysis of 1,249 ChatGPT answers by the SEO consultancy Resoneo found OpenAI's in-house search index served small, unlicensed sites the same way it served licensed publishing partners — comparable formatting, length, and freshness. It corrected an earlier claim that the index favored an allowlist of established publishers.
Not automatically. Being eligible to be indexed is necessary but not sufficient. Separate citation studies still find AI answers lean toward high-authority domains, so a small site earns citations by publishing genuinely useful, specific, answer-shaped content — not just by being crawlable.
For a window in mid-2026, ChatGPT's server stream tagged each web result with the pipeline that fetched it, and one label reported as "labrador" corresponded to OpenAI's own in-house search index rather than a third-party search API. OpenAI reportedly stopped exposing those pipeline labels around July 21, 2026, so treat the specifics as a snapshot.
Publish specific, well-structured content consistently on pages you own, with clear titles and a tight opening the index can store as a snippet. Kompozy helps by generating net-new blogs, carousels, images, video, and newsletters and publishing them across an owned blog plus eight social platforms, building the footprint AI search indexes can pick up.