A controlled experiment across six large language models — 252,000 trials — found that topical relevance and list position, not trust badges or polish, decide which of two competing pages gets cited first in an AI answer.
2026-09-18 · by Moe Ameen
Researchers at Sprinklr published a study, "What Gets Cited: Competitive GEO in AI Answer Engines," that asks a narrow but consequential question: when two retrieved pages compete to answer the same query, what makes an AI engine cite one first instead of the other? The paper was posted to arXiv on May 25, 2026 and accepted to SIGIR 2026, the ACM information-retrieval conference held in Melbourne in July 2026.
The method is deliberately controlled. The authors built a two-document retrieval-augmented generation (RAG) testbed that injects exactly two candidate sources into a model's context and records which one is referenced by the first citation marker in the output. In each trial the two sources differ in exactly one of 18 content factors — spanning topical match, completeness, trustworthiness, readability, competitive standing, and freshness — so the effect of that single factor can be isolated. To keep familiarity and ordering from contaminating the result, they anonymized brands and counterbalanced the source order (running each pair in both orders). They executed 252,000 trials in total across six models: Gemini-2.5-Flash, GPT-5-Nano, GPT-5-Mini, GPT-5.2, Claude-3.5-Sonnet, and Kimi-K2-Thinking, built from 100 anonymized product-review articles across 50 categories.
The headline result: mixed-effects models show that topical relevance and list position are the two biggest drivers of being cited first. In other words, where a page sits in the list of retrieved sources handed to the model — not just how good the page is — strongly predicts whether it earns the citation. Including explicit price information and a recent timestamp also helped consistently. Completeness and trust cues added smaller gains, and formatting-only edits had little impact.
Two caveats worth stating plainly. This is a controlled lab experiment with two synthetic sources per trial, not a measurement of live ChatGPT or Perplexity behavior on the open web, so the exact magnitudes won't map one-to-one to production answer engines. And "list position" is largely set by the retrieval and ranking layer, which a publisher influences indirectly (by ranking well) rather than controls outright.
The study points at three controllable levers — tight topical relevance, explicit price, and a fresh timestamp — and one uncomfortable truth: you have to earn a top spot in the retrieved set before any of that matters. Both jobs come down to publishing enough genuinely on-topic, current content that answer engines keep finding you for the query. That volume is exactly what most solo creators and lean teams cannot sustain by hand, and it is what [Kompozy](/) is built to produce. From one source — a transcript, a doc, a product update — it generates a Blog Article that answers the query directly, a Text Post, brand-exact [Carousel Posts](/glossary/hyperframes), Quote Graphics, an Email Newsletter, and short video, all governed by one [Persona Brief](/glossary/persona-brief) that locks your voice and banned words so the on-topic depth the study rewards stays consistent across every piece.
Freshness and price are workflow problems, and this is where an engine beats a one-off. Because Kompozy is a generation-plus-publishing system, you can spin up a genuinely new, correctly dated piece for a query on a regular cadence instead of letting a page go stale — the "recent timestamp" signal the study found helps, earned by actually republishing, not faking a date. Put a product's real price and specifics into the source and they carry through into the copy verbatim. Then [Autopilot](/glossary/autopilot) schedules and publishes the whole set across the eight social platforms plus blog and email behind a per-post review gate, so you are building topical coverage everywhere answer engines retrieve from — your own blog, plus the social and community surfaces they lean on. You cannot hand-set your list position, but you can flood the query with fresh, specific, on-brand answers until being retrieved near the top stops being luck. That is the work Kompozy does at volume.
In a controlled two-source experiment across six large language models — 252,000 trials in total — topical relevance and list position were the two biggest factors in which page an AI engine cited first. Where a source sat in the list of retrieved candidates strongly predicted whether it won the citation, alongside how well it matched the query. Explicit price information and a recent timestamp helped consistently, while formatting-only changes had little effect.
It was authored by researchers at Sprinklr — "What Gets Cited: Competitive GEO in AI Answer Engines" — posted to arXiv on May 25, 2026 and accepted to SIGIR 2026, the ACM information-retrieval conference held in Melbourne in July 2026. It tested Gemini-2.5-Flash, GPT-5-Nano, GPT-5-Mini, GPT-5.2, Claude-3.5-Sonnet, and Kimi-K2-Thinking.
Not directly. List position is largely decided by the retrieval and ranking layer that feeds sources to the model, so you influence it indirectly by ranking well for the query rather than setting it yourself. What the study shows you can control is the content: be clearly on-topic, include explicit price and specifics when relevant, and keep the page genuinely fresh. Publishing enough on-topic, current content is how you improve the odds of being retrieved near the top.
No. It is a controlled lab experiment that injects exactly two synthetic candidate sources per trial to isolate one factor at a time, with brands anonymized and source order counterbalanced. That design cleanly separates content effects from position bias, but the exact magnitudes will not map one-to-one to live answer engines on the open web. Treat the direction of the findings — relevance and position dominate, price and recency help, cosmetics do not — as the takeaway.