The research on generative engine optimization has grown from one Princeton paper into a real literature, and its headline is more sobering than the industry selling GEO would like: across the dozens of studies now reviewed, no tactic has demonstrated a stable, cross-platform effect on whether engines discover you organically, while the levers that do replicate are topical relevance and where your passage sits in what the engine reads. For anyone selling secret tactics that’s bad news. My read is that the boring parts were always the job.

I read these papers because their subject is my day job. I lead Innovation and AI at Tandem Interactive, where the AI-visibility tracking I built samples engines in production every day, so when arXiv started filling up with GEO studies I wanted to know whether the field agreed with what the dashboards had been telling me. Mostly it does, and the disagreements taught me more.

What does the GEO research corpus look like?

The corpus splits into three branches: studies of what makes engines cite a source, methodology papers about how to measure visibility at all, and a newer adversarial branch treating GEO as an attack surface. The four papers below are the ones I would hand a colleague first, because each settles a question practitioners have been arguing about from anecdotes.

Paper Scale The finding
A critical survey of GEO, 2023 to 2026 (Martinez) 45 studies reviewed No technique shows a stable cross-platform effect on organic discoverability; relevance and context position replicate
Don’t Measure Once (Schulte, Bleeker, Kaufmann) Methodology Visibility in AI search is a distribution across runs, prompts, and time, never a single observation
GEO at scale (Kumar) 100,000+ responses, 100+ brands Brand visibility stacks in tiers by stature, with roughly thirty points between rungs
What Gets Cited (Vishwakarma, Kumar, Jamidar) 252,000 trials, six LLMs Topical relevance and list position drive citation choice; formatting-only edits barely register

What holds up across 45 studies?

Across the 45 studies in Martinez’s survey, the tactics that replicate are topical relevance and the position your content holds in what the engine reads, while the famous percentage gains from the original GEO paper hold only under a condition most coverage skipped: the source was already inside the engine’s context window. Nothing reviewed demonstrates a stable, longitudinal, cross-platform effect on organic discoverability, which is the thing every GEO invoice implicitly promises.

The distinction between being cited once retrieved and being retrieved at all is the spine of my GEO playbook, and the survey landing there saved me rewriting half of what I tell clients. The survey also documents that citation-oriented rewrites can impair retrieval, which I’ve watched teams walk into: a page gets restructured so aggressively for quotability that it stops ranking, and the citation never comes because the engine never fetches the page.

Do you have to measure AI visibility as a distribution?

Yes, and there is now a paper to hand anyone who wants the weekly screenshot back. Schulte, Bleeker, and Kaufmann make the case that a single query gives you a representative snapshot in classic search and nothing of the kind in AI search, where the same prompt returns different answers across runs, sessions, and weeks, so a brand’s visibility only means something as a distribution over repeated measurements.

I built the tracking in Nexus around repeat sampling and Wilson intervals because clients were close to making budget decisions on sampling noise, and the full method is written up in my measurement essay. What the paper adds is standing: when a vendor report shows one confident number with no sample count behind it, the objection is no longer mine alone.

How much of AI visibility is brand stature?

Most of it, on current evidence. Kumar’s tracking of 100,000+ answers across 100+ brands through spring 2026 found household names appearing in roughly 73% of relevant answers, established mid-market brands in 44%, and niche brands in 11%, which stacks the market into tiers separated by about thirty points each. The 252,000-trial study from Vishwakarma, Kumar, and Jamidar points the same direction from the content side: citation choice runs on topical relevance and list position, timestamps and concrete details help at the margin, and formatting-only edits do close to nothing.

  • Household names 73%
  • Mid-market brands 44%
  • Niche and small brands 11%
Share of relevant AI answers in which brands appeared, by tier, from Kumar's tracking of 100,000+ responses in spring 2026. Roughly thirty points separate each rung.

Read together, those results say on-page tactics matter at the margin while brand tier decides most of the outcome, and tiers move slowly, through coverage you don’t control and entity data that agrees with itself everywhere it appears. That’s the same conclusion my entity SEO essay reaches from inside one contested name, so on this evidence I’d put the next dollar into corroboration before putting it into another round of on-page polish.

What about GEO manipulation research?

A second branch of the literature treats GEO as an attack surface, and it’s growing fast. There are benchmarks measuring how far adversarial content rewriting pushes engine answers (GEO-Bench), studies of recommendation agents degrading when product pages are poisoned with GEO-style manipulation (SafeGEO), and a proposal for platform incentives that reward verifiable content over citation-bait (Xu, Guo, and Xiong). The engines’ operators can read these papers as easily as I can, and the history of SEO is a history of tactics working loudly right up until they became detection signatures.

The adversarial branch didn’t change what I run. Nothing in those papers makes me want to own a strategy that depends on engines staying naive, and every defense they prototype raises the value of the patient route the rest of this essay argues for.

What changes in practice?

The research changes less of my week than I expected:

  1. Keep funding the fundamentals, because retrieval still gates everything. The survey’s discoverability gap is technical SEO’s whole argument.
  2. Structure for the passage, not the page. Relevance and position are the replicated levers, and they’re cheap.
  3. Stop buying formatting tricks. A quarter-million trials say formatting-only edits barely register.
  4. Put real budget into third-party corroboration and entity consistency, because the stature ladder is the biggest coefficient in the system.
  5. Publish original numbers worth citing. Being the source is the one on-page move that also builds authority off the page.
  6. Accept only measurements with sample counts and intervals, from vendors and from me. The methodology literature is on your side now.
  7. Skip anything that smells adversarial, since the defenses are already being built.

The corpus is young, and some of these findings will sharpen or break as the engines change underneath them, which is a normal thing to say about a literature three years old. The primary sources in the table above are where I’d start anyone who wants to check my read against their own.

FAQ

Is generative engine optimization backed by research?

Partly, and the split matters. A real research corpus now exists on arXiv, and it supports passage-level structure, topical relevance, and distribution-based measurement while finding no stable evidence that any tactic improves organic discoverability across platforms. It supports the fundamentals and undercuts the tricks.

What GEO tactics does the research support?

Topical relevance to the engine’s sub-questions, content positioned and structured so context assembly includes it, and being retrievable through ordinary search infrastructure in the first place. Concrete details like timestamps help at the margins. Formatting-only edits showed almost no effect across 252,000 trials.

Does brand authority matter more than GEO tactics?

On current evidence, yes. Large-scale tracking found household names in roughly 73% of relevant answers and niche brands in 11%, a gap no on-page tactic in the literature comes close to closing. Corroboration, coverage, and consistent entity data are what move a brand between tiers.

How should you measure whether GEO is working?

As a distribution, never as a single answer. Sample each prompt repeatedly per engine, report rates with confidence intervals, and be suspicious of any tool that shows one number without a sample count. The measurement methodology now has peer company in the literature.

Where should you start reading GEO research?

Start with Martinez’s critical survey for the map, then the Don’t Measure Once paper for methodology, then the two large empirical studies on citation drivers and brand tiers. All four are linked in the table near the top of this essay and readable in an afternoon.