GEO Beyond the Hype: What the Princeton Generative Engine Optimization Study Actually Measured
Roth Miklós

Few acronyms have spread through marketing circles as quickly as GEO — generative engine optimization. Agencies sell it, conference talks promise it, and much of the commentary traces back to a single academic paper. Reading that paper carefully, rather than the slide decks built on top of it, turns out to be the more useful exercise. What the researchers measured is genuinely interesting; what some vendors claim it proves is considerably more ambitious than the data.
What the study actually did
The paper, “GEO: Generative Engine Optimization” by Aggarwal and co-authors, was published on arXiv in November 2023 and presented at KDD 2024. The authors built a benchmark called GEO-bench — a large set of user queries across diverse domains — and evaluated how different content modifications affected a source’s visibility inside generative-engine answers. Visibility was measured through impression-style metrics: whether a source was cited, how prominently, and how much of the answer it supported.
The headline finding most often quoted: certain optimization methods improved source visibility by up to 40 per cent. That figure deserves its qualifier. It is a best-case result measured on a simulated generative engine built by the researchers for the benchmark — not a controlled experiment run inside ChatGPT, Gemini or Perplexity at production scale. The authors themselves frame the work as a first systematic study of a new optimization problem, not a finished playbook.
Which tactics moved the numbers
The practically relevant part of the paper is the ranking of tactics. Three families of interventions performed best:
Citing sources. Adding references to authoritative external material made content more likely to be surfaced and credited. This aligns with how retrieval-augmented systems work: passages that themselves contain verifiable claims are easier for an engine to reuse with attribution.
Adding quotations. Incorporating relevant quotes from identifiable people or documents improved measured visibility, plausibly because quotations add extractable, attributable substance.
Adding statistics. Concrete, sourced numbers gave engines something specific to lift into an answer.
Just as instructive is what did not work well. Classic keyword stuffing — repeating query terms — performed poorly, and in some configurations worse than doing nothing. The finding supports a conclusion many practitioners had suspected: generative engines reward informational density and verifiability, not term frequency.
What the study does not establish
An honest reading also requires listing the limits. The benchmark was run against a research prototype, so absolute effect sizes may not transfer to any specific commercial engine. The paper measures visibility within answers, not downstream business outcomes such as traffic, leads or revenue. And it does not compare GEO tactics against simply having strong, authoritative content to begin with. Treating “40 per cent” as a guaranteed uplift for any website would misread the methodology.
A practitioner perspective from Central Europe
Miklós Róth, founder of the Budapest-based AI Marketing & SEO Agency, which also positions for the Vienna and Zurich markets through German-language service sites, has been applying citation-oriented content tactics in client programs and comments regularly on the shift from rankings to answer surfaces. According to his published profile, he describes himself as an AI-driven marketing and SEO strategist with long-standing experience in search. His practitioner reading of the study is deliberately conservative: the measured tactics — real citations, named sources, verifiable statistics — are not tricks layered onto weak content, but properties that make content worth citing in the first place. On his personal-brand site, mymarketingworld.at, he frames GEO as an extension of editorial quality rather than a replacement for it, and cautions against vendors who present a single research result as a universal ranking formula.
That framing is consistent with the paper’s own tone. The authors position GEO as an open research direction and call for further work on real-world engines and user behaviour.
What marketing teams can reasonably take away
First, audit your most important pages for citability: do they contain sourced claims, named experts and concrete data, or only generic statements? Second, treat citations as a two-way street — pages that reference credible sources are structurally easier for generative engines to credit. Third, keep measurement honest: track whether your brand and pages actually appear in AI answers for your target queries over time, rather than trusting a projected percentage. Finally, expect the ground to keep moving. The engines change monthly, and the research literature is only beginning to catch up.
GEO, stripped of the hype, is a useful discipline with a promising early evidence base. The study behind it measured real effects under defined conditions — and defined conditions are exactly what marketing teams should demand from anyone selling them certainty.
Useful references for this topic: AI Marketing & SEO Agency Budapest/Vienna website, Service details, Authority guidance, Industry context, Further official reference.