What the Latest GEO Research Actually Tells Us
The GEO industry promises 40% visibility gains. Olivier Martinez's July 2026 survey of 45 studies (November 2023 through July 2026) offers something more useful: a map of what generative engine optimization can prove today, where the evidence thins out, and what that means for your content program.
If you sell software or services in search, you have heard the pitch: rewrite a page, add a few statistics, and ChatGPT will cite you 40% more often. That number comes from real academic work. Aggarwal et al. published "GEO: Generative Engine Optimization" at KDD 2024. It belongs in every serious GEO conversation. It does not belong on a slide as a guaranteed outcome.
Martinez's A Critical Survey of Generative Engine Optimization (arXiv:2607.14035, July 2026) reviews 45 studies on generative visibility, from the foundational GEO paper through commercial audits of Google AI Overviews, Perplexity, and ChatGPT. The conclusion is not that GEO is fake. The conclusion is that GEO is multistage, partially observable, and easy to misread when vendors collapse citation, retrieval, and revenue into one score.
Primary source
Olivier Martinez, Optimizing Visibility in Generative Engines: A Critical Survey of Generative Engine Optimization (arXiv:2607.14035v1, July 15, 2026). This guide summarizes the survey for practitioners; read the paper for the full evidence tables and study matrix.
This guide translates that survey for digital marketers and SEO practitioners. For the full category definition, see our GEO explainer and how to do GEO workflow .
What does the 2026 GEO research actually establish?
Direct answer
Martinez grades GEO claims by evidence strength. High confidence: content already in an AI system's retrieved context can change citation share; relevance and context position dominate; commercial engines differ and drift over time. Moderate: extractable evidence and learned optimization help in controlled settings. Low: durable organic discoverability. Very low: traffic and conversion impact. The widely cited 40% figure is a conditional lab result, not a universal promise.
Table 5 in the survey summarizes confidence levels. Three claims sit at the top. First, a document already placed in the generator's context can causally alter its rank, citation rate, or use. Second, query-document relevance and position inside the context window are major determinants of what gets cited first. Third, commercial engines cite different source ecosystems and change decisions between runs.
| Confidence | Claim |
|---|---|
| High | Retrieved content can change citation and prominence inside the context window |
| High | Relevance and context position dominate most rewrites |
| Moderate | Extractable stats, definitions, and structure help when truthful |
| Low | Stable cross-engine gains in organic discoverability |
| Very low | Citation scores predict clicks, conversions, or revenue |
| Rejected | "GEO increases visibility by 40%" as a general claim |
Aggarwal et al. tested nine content modifications on a fixed five-document context drawn from Google top results. Quotation Addition raised position-adjusted word count (pawc) from 19.3 to 27.2, about a 41% relative increase on that metric. Keyword stuffing lowered pawc to 17.7. Those numbers describe redistribution of attributed answer share among sources already in the window. They do not prove your page will be retrieved tomorrow on Perplexity, Gemini, or ChatGPT browse mode.
The gap between conditional citation effects and organic discoverability is the central story of the survey. Most GEO tooling optimizes downstream text. Most real-world visibility problems sit upstream: crawlability, index coverage, topical fit, and whether the engine activates search at all.
Win retrieval before you rewrite for citation
Martinez shows that conditional citation gains are well supported. Organic retrieval is not. GeoCopy helps on both sides: campaigns produce depth across topic clusters so more URLs can enter retrieval pools, and every article ships with answer capsules, sourced statistics, and citation-ready structure for the generation stage. Free trial: 10 free articles, no credit card required.
How does the GEO pipeline actually work?
Direct answer
Generative visibility is a chain: search activation, retrieval, reranking into the context window, answer generation and citation, factual absorption, fidelity checks, and user behavior. An intervention can raise citation while lowering retrieval. Martinez models this as a visibility vector (discoverability, context exposure, citation, prominence, absorption, fidelity, behavior), not a single rank score.
Figure 1 in the survey labels the causal pipeline marketers should picture:
Activation → Crawling / indexing → Retrieval → Reranking / context → Generation / citation → Absorption / fidelity → Attention / click / conversion
Most vendor playbooks focus on the middle: rewrite the body, add quotes, insert statistics. SAGEO Arena (Kim et al., 2026) reintroduced retrieval and reranking on 171,003 documents. Body-only optimization cut average top-20 presence by about 9%, top-10 presence after reranking by 16%, and final citation by 6%. AutoGEO applied to the body alone produced even larger upstream losses in some configurations. A page can win a lab citation test and lose ground in a full pipeline test.
| Stage | What you optimize | Typical mistake |
|---|---|---|
| Activation | Query patterns that trigger search | Scoring citations only when search ran |
| Retrieval | Coverage, crawlability, relevance | Polishing copy on pages the index never selects |
| Context position | Authority, intent match, internal links | Assuming formatting beats rank in the stack |
| Citation | Extractable claims, quotes, stats | Keyword stuffing (negative in Aggarwal et al.) |
| Fidelity | Accurate, verifiable statements | Chasing mentions that misquote your page |
| Behavior | Offer, UX, brand demand | Reporting AI mentions as revenue proof |
Martinez separates conditional citation probability Pr(cited | retrieved) from discoverability Pr(retrieved). High conditional citation with low retrieval still yields low commercial visibility. That identity alone explains why dashboards full of green citation badges can coexist with flat referral traffic.
For a plain-language walkthrough of retrieval and synthesis, see how LLMs cite sources and what is RAG .
Which GEO tactics hold up under scrutiny?
Direct answer
Prioritize topical relevance, retrieval strength, and context position. Add verifiable statistics, definitions, prices, and dates where they match intent. Treat document structure as stage-specific. Deprioritize generic recipe lists (C-SEO Bench: 3 of 54 method-domain pairs positive), body-only rewrites without upstream tests, and keyword stuffing.
1. Topical relevance beats generic authority cues
Wan et al. (ACL 2024, ConflictingQA) use counterfactual passages to show models strongly favor text that explicitly aligns with the question, sometimes ahead of scientific references or neutral tone. Write pages that resolve one user intent per URL. Stack secondary keywords only when they support the same answer.
2. Context position beats most rewrites
Puerto et al. (C-SEO Bench, NeurIPS 2025) find that moving a source higher in the context window outperforms many transformation recipes. Vishwakarma et al. (SIGIR 2026) ran 252,000 trials across six LLMs and eighteen factors. Relevance and position drive first citation more reliably than fluency tweaks or authoritative tone alone.
3. Extractable evidence helps when it is true
Statistics, definitions, comparisons, prices, and dates give models discrete units to quote. Aggarwal et al., AutoGEO (Wu et al., ICLR 2026), and FeatGEO (Liu and Xu, ACL 2026) report gains for several informational features. Martinez adds a hard qualification: fabricated statistics can increase reuse while damaging fidelity. The rule is verifiable evidence with attribution, not "add numbers."
4. Structure is not a universal talisman
Headings, tables, and HTML fields can improve passage extraction, but SAGEO shows structural optimization may help retrieval on one benchmark and hurt reranking on another. Test each stage instead of assuming one template wins everywhere.
What the survey tells you to stop doing
- Copying fixed GEO recipes at scale. C-SEO Bench found only three significant positives out of 54 method-domain pairs in the main experiment and none in question answering. Broad adoption drives gains toward zero.
- Body-only optimization without retrieval checks. SAGEO Arena documents upstream losses that erase downstream citation wins.
- Keyword stuffing. Aggarwal et al. measured lower pawc versus baseline, aligning with older SEO failure modes.
Martinez's conservative recommendation matches what high-quality editorial teams already do: publish relevant, comprehensive, verifiable, clearly structured pages that crawlers can fetch. Measure retrieval, citation, and fidelity separately. For execution detail, see GEO optimization techniques .
Ship evidence-backed GEO structure on every publish
GeoCopy bakes in what the survey supports: question-format headings, 40–60 word answer capsules, named expert quotes, inline citations, comparison tables, and freshness-friendly updates. It avoids what the survey rejects: keyword stuffing, opaque recipe rewrites, and body-only edits with no retrieval strategy. Run a free GEO score check on an existing page, then generate a citation-ready replacement in one workflow.
What do platform audits reveal about real-world GEO?
Direct answer
Commercial generative surfaces cite different domains, activate search on different query shapes, and change sources between runs. Kirsten et al. (ACL 2026) and Grossman et al. (SIGIR 2026) document low cross-surface overlap. Xu et al. (2026) report 13.7% overall AI Overview activation versus 64.7% for question-phrased queries. Traffic impact remains the weakest part of the evidence base.
There is no global GEO rank. Li and Sinnamon (2024) found only 26% domain overlap between Bing Chat and Perplexity in one audit phase. Kirsten et al. report that 53% of Google AI Overview domains sit outside the organic top 10 and 27% sit outside the top 100. Two-month page overlap for AI Overviews was 18% compared with 45% for organic Google. Grossman et al. measured URL-level Jaccard similarity of 0.11 to 0.18 among organic Google, AI Overviews, and Gemini.
| Signal | Finding | Source |
|---|---|---|
| Cross-engine overlap | Jaccard often 0.11–0.18 | Grossman et al. 2026 |
| AIO vs organic stability | 18% page overlap over two months | Kirsten et al. 2026 |
| AIO domains outside top 10 | 53% | Kirsten et al. 2026 |
| AIO activation rate | 13.7% overall / 64.7% for questions | Xu et al. 2026 |
| ChatGPT without web search | 57.8% of repetitions | Schulte et al. 2026 |
| Run-to-run stability | Daily Jaccard ~0.34–0.42 | Schulte et al. 2026 |
Sharma (2026) highlights a discovery gap for startups: models recognize named products at high rates but surface them on only 3.32% to 8.29% of organic discovery queries in that preprint sample. Brand awareness in weights or by name is not the same as recommendation for generic intent.
Citation also does not equal support. Liu et al. (EMNLP 2023) measured historical engines where only 51.5% of sentences were fully supported and 74.5% of citations matched their associated claims. Xu et al. classify about 11% of atomic AI Overview claims as insufficiently supported in their 2026 audit. Track fidelity alongside mention rate.
Platform guides: optimize for Perplexity , Google AI Mode , and LLM visibility metrics .
Why does competition change GEO outcomes?
Direct answer
GEO share is often zero-sum inside a fixed context window. Aggarwal et al. show large gains for lower-ranked sources and losses for leaders under some strategies. Puerto et al. show method gains eroding as adoption spreads. Early depth in a topic cluster still wins while optimization density stays low.
When attributed word share must sum to 100% across sources, one page's lift is another page's loss. Under Aggarwal et al.'s Cite Sources strategy, the fifth source in a five-document context gained 115.1% in pawc while the first lost 30.3%. Puerto et al. model congestion as more documents adopt the same conversational SEO tactics. A playbook that works for an early adopter can flatten once competitors copy it.
Nestaas et al. (ICLR 2025) and follow-on ranker studies show retrieved content is also an attack surface: indirect instructions in pages can shift recommendations. White-hat optimization and adversarial manipulation share a channel. Martinez proposes normative tests: semantic preservation, evidentiary authenticity, no hidden model instructions, and fair disclosure.
Build the library before the niche gets crowded
Competition erodes generic tricks faster than it erodes depth. GeoCopy campaigns turn keyword clusters into dozens of retrieval-ready articles with consistent GEO structure, internal links, and CMS publish. Start while your category still has room to own the context window.
How should marketers measure GEO without fooling themselves?
Direct answer
Define the estimand before you collect data: conditional rewrite effect, total pipeline effect, observational visibility, or business outcome. Track activation, retrieval, citation, fidelity, and null answers. Repeat prompts across paraphrases, engines, and dates. Schulte et al. suggest seven to eight repetitions per prompt as a starting point for stable estimates.
Martinez's Section 11 reads like a study protocol, but the marketing takeaway is simple. Most teams measure what is easy (one ChatGPT query, one citation badge) instead of what matters (whether search activated, whether you were retrieved, whether the citation was accurate, whether traffic moved).
Monthly checklist derived from the survey:
- Retrieval presence when search activates
- Citation rate conditional on retrieval
- Position proxy via competitor citation order on target queries
- Fidelity spot checks on cited claims
- Share of runs with no search or no citations (keep these rows)
- Referral trend with a control for platform growth (Watanabe and Nakayashiki 2026)
Use the identity Pr(cited) = Pr(search) × Pr(retrieved | search) × Pr(cited | retrieved, search). High conditional citation with low activation still means low visibility. For tooling options, see best GEO tools and GEO tool stack guide .
Frequently asked questions
Is the 40% GEO lift real?
Yes, in Aggarwal et al. (KDD 2024) when five pages were already in context and position-adjusted word count was measured. No, as a universal claim about rankings, traffic, or stable cross-platform citation. Martinez (2026) rejects the general marketing version of the figure.
What should marketers optimize first for GEO?
Topical relevance and retrieval presence first, then position inside the retrieved context window, then extractable evidence (stats, definitions, dates). Puerto et al. (2025) and Vishwakarma et al. (2026) show position and relevance beat most rewrites once a page is in the candidate set.
How long until GEO results show up?
Controlled citation effects can appear quickly when a page is already retrieved. The survey does not support one timeline for organic discoverability or revenue. Kirsten et al. and Schulte et al. show high run-to-run and cross-date variability, so plan for repeated measurements over months.
Does GEO work at all?
Yes, with limits. Martinez (2026) rates it high-confidence that content already in context can change citation and prominence. No reviewed technique yet proves durable, cross-platform causal gains in organic retrieval or conversions.
Does an AI citation mean the model trusts my page?
Not automatically. Liu et al. (EMNLP 2023) found only 51.5% of sentences fully supported in historical generative search engines and 74.5% of citations correct. Track fidelity, not just mention count.