We finally have real research on what makes AI tools recommend one brand over another, and some of it contradicts the advice being sold right now.
Two studies published recently are worth your attention because they were actually measured rather than asserted. The first is a peer-reviewed paper presented at SIGIR 2026, “What Gets Cited: Competitive GEO in AI Answer Engines,” from a research team at Sprinklr. The second is a large observational study from The Digital Bloom covering 105,000 ChatGPT prompts. They approach the problem from opposite ends, and together they form the clearest picture we have.
Here is what they found, in plain terms.
There Are Two Separate Problems, Not One
Almost every article about AI visibility treats it as a single challenge. It isn’t. There are two distinct stages, and failing at either one makes you invisible:
- Retrieval: does the AI find your page at all when it goes looking?
- Citation: once your page is in front of the model alongside competitors, does it get chosen?
The Digital Bloom study measures the first. The SIGIR paper measures the second. Keeping them separate is the single most useful idea in either study, because the fixes are completely different. If you are never retrieved, writing better content will not help you. If you are retrieved but never cited, more backlinks will not help you.
Stage One: Getting Retrieved
The Digital Bloom analysis looked at 145 industries, 1,595 buyer personas, and 29,562 domains, then correlated who got recommended against thirteen external signals. The strongest correlations:
| Signal | Correlation | Variance explained |
|---|---|---|
| Appearances in search engines | +0.241 | 5.8% |
| Best search engine rank | +0.238 | 5.7% |
| Backlink count | +0.204 | 4.2% |
| Backlink authority | +0.200 | 4.0% |
| PageRank | +0.194 | 3.7% |
| Wikidata presence | +0.120 | 1.4% |
| Reddit comments | +0.111 | 1.2% |
The headline: traditional search visibility is the strongest known predictor of AI recommendation. Ranking well in Google and having real backlink authority remain the best-measured inputs to being recommended by an AI. Anyone telling you SEO is dead is contradicted by the only large dataset we have.
Two things make this more interesting than it first looks.
First, the industry variation is enormous. Wikidata presence explains 49.9% of the variance for furniture retailers and 42.9% for ERP software, while Reddit posts explain 27.9% for enterprise AI platforms and 25.3% for live entertainment. The signal that matters most in your category may be almost irrelevant in another. This is the strongest argument I have seen for researching your own market rather than following a generic checklist, which is the same point I make about knowing the prompts your buyers actually use.
Second, and be honest about this one: all thirteen signals combined explain well under 20% of the outcome. The authors put 80 to 85% of AI recommendation behavior down to model internals, meaning training data, fine-tuning and reinforcement learning that nobody outside the labs can observe or influence. Anyone selling you a complete formula for AI rankings is selling you the small visible slice and not mentioning the rest.
Stage Two: Getting Cited
This is where the SIGIR paper is genuinely valuable, because it is a controlled experiment rather than a correlation study.
The method is clever. The researchers took 100 real product review articles, stripped out every brand and publisher name so familiarity could not skew the results, then created pairs of pages that were identical except for exactly one characteristic. They showed each pair to six different language models (GPT-5 variants, Claude, Gemini and Kimi), swapped the order to control for position bias, and recorded which one got cited first. In total: 252,000 trials across 18 content factors.
One incidental finding is worth pausing on. In 86.4% of answers, the model cited exactly one source. Not three, not five. One. AI answers are closer to winner-take-all than a search results page, which changes how you should think about “visibility” entirely.
The four gatekeepers
Four factors were significant across all six models with very large effects. Fail any one of these and your citation odds collapse regardless of how good the rest of your page is:
- Topic match. The page has to actually be about the thing being asked. Obvious, but it was the single biggest effect measured.
- Price is stated. Pages that include explicit pricing beat otherwise identical pages that omit it, decisively. “Contact us for pricing” is a citation killer.
- A recent timestamp. Content dated 2026 crushed identical content dated 2019. Same words, different date, dramatically different outcome.
- Position in the retrieved list. Being the first source rather than the second was close to deterministic across models. This is why stage one matters so much: retrieval rank carries into citation.
The seven differentiators
Once the gatekeepers are satisfied, these provided meaningful additional advantage:
- Technical specifications included rather than omitted, one of the largest secondary effects in the study
- Comprehensive coverage rather than shallow treatment of the topic
- Confident language rather than hedging with “might,” “possibly” and “could”
- Claims backed by evidence such as tests or certifications
- Internal consistency with no contradictory statements
- Query terms present rather than absent from the content
- Comparisons with alternatives rather than describing your product in isolation
Notice how much of this is simply publishing complete, specific, verifiable information. Price, specs, evidence, comparisons. That is not a trick, it is a thorough product page, and it is the same conclusion I reached from client work before this research existed.
What Did Not Matter
Seven of the eighteen factors showed no consistent effect, and this is the part that should change some behavior.
Formatting had no measurable impact. Structured sections versus dense paragraphs, and organized versus scattered information, made no reliable difference to which source got cited. The researchers concluded that language models parse content regardless of visual organization.
I want to be careful and precise about this, because it partially cuts against advice I have given. Formatting still matters for retrieval, for extraction, for accessibility, and for the humans who read your pages. What this study shows is narrower and still important: when two pages containing the same information compete for a citation, prettier structure does not win it. If you have been reformatting existing content into bullet lists and headings and calling that AI optimization, this is evidence that the effort is better spent adding information that is not there yet.
Also showing no consistent effect: promotional tone, stronger value proposition claims, and social proof such as ratings and review counts. Marketing language does not move the needle. Facts do.
The Models Do Not All Behave the Same
Sensitivity to content quality varied a lot. Kimi-K2 responded to 83% of the tested factors, the GPT family sat in the middle, Claude responded to 50% and Gemini to just 33%. Gemini and Claude also tended toward near-binary decisions, either strongly preferring a source or not distinguishing at all.
All six agreed on the four gatekeepers. So the gatekeepers are safe to treat as universal, while the finer differentiators matter more in some assistants than others.
What I Would Actually Do With This
- Diagnose which stage is failing first. Ask the assistants your buyers use, and look at what gets cited. If your content is not cited at all, you have a retrieval problem, which is SEO work. If you are cited but not recommended, you have a content problem, which is the factor list above. Fixing the wrong one wastes months.
- Put prices on the page. Of everything here, this is the easiest change with the largest measured effect, and it is the one most B2B companies refuse to make. If you genuinely cannot publish a price, publish ranges, starting points, or the model by which price is determined.
- Publish complete specifications. The evidence for specs, comparisons and evidence-backed claims is strong, and it matches what I have seen produce real client results.
- Update your dates, and mean it. A recent timestamp is a gatekeeper. That does not mean changing the date on stale content, because the models are also weighing the content itself. It means genuinely maintaining your important pages, which is the habit I argue for anyway.
- Stop hedging. Replace “may help improve” with what your product does, and back it with evidence. Hedged language measurably loses citations.
- Keep doing SEO. Search visibility is the strongest known retrieval signal, and retrieval position is a citation gatekeeper. The two stages are connected, and traditional rankings feed both.
- Do not trust anyone claiming certainty. Over 80% of this is unexplained by any observable signal. Treat confident formulas accordingly.
The Honest Summary
The measurable part of AI visibility comes down to being findable through conventional search, and then being the most complete, specific, current and factual answer in the retrieved set. Price, specifications, evidence, comparisons, recency. No formatting trick, no file you can upload, no schema hack substitutes for publishing better information than your competitors.
That is unglamorous, which is probably why so much of the advice in this space is about anything else. It is also the part you can actually control, and the research now supports it. If you want help working out which of the two stages is holding your brand back, that diagnosis is the starting point of most of my AEO and AI SEO engagements.
Sources
- Vishwakarma, R., Kumar, S., and Jamidar, R. (2026). What Gets Cited: Competitive GEO in AI Answer Engines. SIGIR ’26, Melbourne.
- The Digital Bloom. LLM Ranking Factors 2026.