What is GEO (generative engine optimization)? Origin, the study behind it and how it differs from SEO
Published 7 October 20269 min readBy the SmoothSeen editorial team
GEO, short for generative engine optimization, is the work of getting a generative engine such as ChatGPT search, Perplexity or Google's AI Overviews to use your page as a source when it answers. It has nothing to do with geotargeting. The term comes from a study presented at KDD 2024 that measured which text changes increase that visibility.
Key points
- GEO (generative engine optimization) is the work of getting a generative engine to use your page as a source when it writes its answer. It has nothing to do with geotargeting.
- The term comes from a paper by Aggarwal and colleagues (arXiv, November 2023), presented at KDD 2024, with a 10,000-query test bench called GEO-bench.
- Adding quotations was the most effective technique (27.2 against 19.3 on its main metric); keyword stuffing scored below leaving the page alone (17.7).
- The techniques helped sources ranked fifth a great deal (up to 115.1%) and took visibility away from those ranked first (up to 30.3%).
- The study does not show how to get into the list of sources: that still depends on crawling, indexing and SEO.
To check it on your own site: AI visibility audit
On this page
What does GEO mean, and where does the term come from?
GEO is the optimisation of a website's content so that an answer engine built on generative AI uses it and cites it when it replies. In marketing, "geo" usually refers to location: geotargeted ads, geofencing or local SEO. Not here. If you are trying to show up in Google Maps, what you need is the guide to local SEO in the age of AI search.
The term was coined by Pranjal Aggarwal, Vishvak Murahari, Tanmay Rajpurohit, Ashwin Kalyan, Karthik Narasimhan and Ameet Deshpande, from Princeton University, the Indian Institute of Technology Delhi and independent research. They posted "GEO: Generative Engine Optimization" on arXiv on 16 November 2023 and presented it at KDD 2024, held in Barcelona from 25 to 29 August 20241.
The authors define a generative engine as a system that searches for information and generates an answer from several sources, with inline attributions to each. Their examples were Bing Chat, Google's then-experimental SGE and Perplexity. Their argument is that in these engines a website's visibility is no longer a position in a list: it depends on how much of the answer comes from it and where in the answer it appears.
GEO is one of the two halves of AI search optimization. The other half, AEO, deals with format; the comparison of all three acronyms is in AEO vs GEO vs SEO.
How did the study measure visibility in a generative engine?
As there was no public dataset for this, the authors built one, along with their own test engine. The set-up has three parts1.
The test bench, GEO-bench. It holds 10,000 queries from nine sources, including real anonymised Bing and Google searches (MS MARCO, ORCAS-1 and Natural Questions), trending queries from Perplexity, questions from Reddit's ELI5 forum and queries generated with GPT-4. They span 25 topics; 80% are informational, 10% transactional and 10% navigational.
The engine. For each query they took the text of Google's top five results and asked GPT-3.5 Turbo for an answer citing those sources. Each test was repeated with five random seeds to reduce variance.
The intervention. For each query they picked one of the five sources at random and rewrote it with each technique separately, again using GPT-3.5 Turbo. That way they could compare the same source before and after each change.
They measured the effect with two metrics:
- Position-adjusted word count. It counts the words in the answer attributed to the source and gives more weight to those near the top. This is the main metric.
- Subjective impression. GPT-3.5 scores seven aspects of how the citation is used: relevance to the query, influence on the answer, uniqueness of the material, perceived position, perceived amount, likelihood that the user clicks and diversity of the material.
Which techniques worked, and which did not?
The nine techniques split into those that add content (quotations, sources, statistics) and those that change how existing content is presented. This is Table 1 of the paper, sorted from largest to smallest effect on the main metric1:
- Technique
- No changes (baseline)
- What it changes in the text
- Nothing
- Position-adjusted word count
- 19.3
- Subjective impression
- 19.3
- Technique
- Quotation addition
- What it changes in the text
- Adds relevant quotations from credible sources
- Position-adjusted word count
- 27.2
- Subjective impression
- 24.7
- Technique
- Statistics addition
- What it changes in the text
- Replaces qualitative discussion with quantitative data
- Position-adjusted word count
- 25.2
- Subjective impression
- 23.7
- Technique
- Fluency optimisation
- What it changes in the text
- Makes the text read more fluently
- Position-adjusted word count
- 24.7
- Subjective impression
- 21.9
- Technique
- Cite sources
- What it changes in the text
- Adds references to credible sources
- Position-adjusted word count
- 24.6
- Subjective impression
- 21.9
- Technique
- Technical terms
- What it changes in the text
- Adds technical terms where they fit
- Position-adjusted word count
- 22.7
- Subjective impression
- 21.4
- Technique
- Easy to understand
- What it changes in the text
- Simplifies the language
- Position-adjusted word count
- 22.0
- Subjective impression
- 20.5
- Technique
- Authoritative tone
- What it changes in the text
- Makes the text more persuasive and assertive
- Position-adjusted word count
- 21.3
- Subjective impression
- 22.9
- Technique
- Unique words
- What it changes in the text
- Adds uncommon terms
- Position-adjusted word count
- 20.5
- Subjective impression
- 20.4
- Technique
- Keyword stuffing
- What it changes in the text
- Adds more words from the query
- Position-adjusted word count
- 17.7
- Subjective impression
- 20.2
The figures are the overall column of each metric. Some write-ups quote a different sub-column, the unadjusted word count, where the baseline is 19.5 and keyword stuffing 17.8: they do not contradict each other, they measure different things.
What the authors highlight from the table:
- Cite sources, quotation addition and statistics addition achieved a relative improvement of 30 to 40% on the main metric and 15 to 30% on subjective impression. The best technique beat the baseline by 41% and 28%.
- Style changes help too. Fluency optimisation and simpler language gave gains of 15 to 30%, which suggests these engines also value how information is presented.
- An authoritative tone brought no significant improvement. The authors conclude that the engines are already fairly robust to persuasion without data.
- Keyword stuffing did not work. According to the paper, it offered little to no improvement.
There is a useful nuance about combinations. Citing sources works less well on its own than adding quotations, but combined with other techniques it improved visibility by 31.4% on average. The best pair was fluency plus statistics, which beat any single technique by more than 5.5%.
The effect depends on topic and position
The paper lists the topics where each technique worked best. Statistics: law and government, debate and opinion. Quotations: people and society, explanations and history. Citing sources: statements, facts, and law and government.
Starting position matters even more. The study measured the change in visibility according to where the source ranked on Google:
- Technique
- Cite sources
- Source ranked 1st
- −30.3%
- Source ranked 5th
- +115.1%
- Technique
- Quotation addition
- Source ranked 1st
- −22.9%
- Source ranked 5th
- +99.7%
- Technique
- Statistics addition
- Source ranked 1st
- −20.6%
- Source ranked 5th
- +97.9%
The reading is that these techniques share out an answer of limited size among five sources. A site in fifth place can gain a lot; the one already dominating the answer has little to gain and, in the experiment, lost ground.
The Perplexity test
To check whether the results carried over to a real engine, the authors repeated part of the experiment on Perplexity with 200 queries from the test set. As Perplexity does not let you choose sources by URL, they uploaded the source text as files1.
- Technique on Perplexity
- No changes
- Position-adjusted word count
- 24.1
- Subjective impression
- 24.7
- Technique on Perplexity
- Quotation addition
- Position-adjusted word count
- 29.1
- Subjective impression
- 32.1
- Technique on Perplexity
- Cite sources
- Position-adjusted word count
- 26.8
- Subjective impression
- 19.0
- Technique on Perplexity
- Statistics addition
- Position-adjusted word count
- 26.2
- Subjective impression
- 33.9
- Technique on Perplexity
- Fluency optimisation
- Position-adjusted word count
- 26.0
- Subjective impression
- 30.0
- Technique on Perplexity
- Authoritative tone
- Position-adjusted word count
- 25.9
- Subjective impression
- 30.6
- Technique on Perplexity
- Unique words
- Position-adjusted word count
- 23.6
- Subjective impression
- 24.1
- Technique on Perplexity
- Keyword stuffing
- Position-adjusted word count
- 21.9
- Subjective impression
- 28.1
Quotations came out on top again, with a 22% gain on the main metric, and keyword stuffing scored 10% below the baseline on that same metric. Two results do not match the lab: citing sources dropped on subjective impression and the authoritative tone rose. With 200 queries, it is a sign that the effect changes from one engine to another.
What the study does not show
The authors acknowledge some limits, and others follow from the set-up:
- The sources were already chosen. The five sources for each query were the same for every technique, which share visibility among pages that had already made the list; none brought a new page in. The LLM SEO guide explains what decides that entry.
- The engine dates from 2023. GPT-3.5 Turbo in a custom set-up is not ChatGPT or AI Mode in 2026. The authors warn that the methods will need to adapt as engines evolve, as happened with SEO.
- They did not measure the effect on Google. Because the search engine is a black box, they did not evaluate how these changes affect rankings.
- It measures how much a source is used, not whether it is true. The rewrites were done by a model. An added statistic raises visibility on the metric, but on a real website it has to be true and carry its source, or the problem becomes one of trust.
GEO vs SEO
GEO does not replace SEO: it assumes it. Google says its generative AI features are rooted in its core ranking and quality systems and that, from its point of view, optimising for generative AI search "is optimizing for the search experience, and thus still SEO"2. ChatGPT works similarly through a different door: OpenAI says sites that opt out of OAI-SearchBot will not be shown in ChatGPT search answers, though they can still appear as navigational links3.
- What competes
- SEO
- The whole page, in a list of results
- GEO
- Passages from the page, inside a written answer
- How visibility is measured
- SEO
- Position, impressions and clicks
- GEO
- How much of the answer comes from you, where it appears and whether you are cited
- What moves it
- SEO
- Relevance, links, crawling, indexing, page experience
- GEO
- All of that, plus sourced data, attributed quotations and trust signals
- Prerequisite
- SEO
- The search engine can crawl and index the page
- GEO
- The engine retrieves the page, which depends on SEO and on access for its crawlers
Which part of GEO depends on your page?
Not everything that decides a citation is within your reach. This list separates what you can change by editing from what you cannot:
- Depends on the page
- Figures with a source, a date and a sample
- Does not depend on the page alone
- Whether the engine decides to search the web for that question
- Depends on the page
- Links to primary sources and attributed quotations
- Does not depend on the page alone
- Which pages the engine retrieves, and in what order
- Depends on the page
- A visible author, a publication or update date and an About page
- Does not depend on the page alone
- The domain's reputation and mentions on other sites
- Depends on the page
- Structured data that identifies the publisher
- Does not depend on the page alone
- What the model learnt about you in training
- Depends on the page
- Access for AI search crawlers in robots.txt
- Does not depend on the page alone
- Sources rotating from one week to the next
The left-hand column is the one to review first. Author, dates and the About page are covered in E-E-A-T author signals, and writing with data and quotations in how to get cited by AI.
A declaration of interest: this blog belongs to SmoothSeen, a web audit tool that measures visibility in search engines and AI assistants. In its AI ranking analysis, the Credibility (GEO) block reviews that left-hand column: whether there are concrete figures, dates or prices, links to external sources, attributed quotations, authorship, a date, an About page, signs of experience and structured data, plus the domain's authority and whether AI search crawlers can get in. It audits the causes on the page; it does not measure whether Perplexity, Claude or AI Overviews cite you.
Frequently asked questions
Does GEO have anything to do with geolocation?
No. In this context, GEO stands for generative engine optimization: getting an AI assistant to use your page as a source when it answers. In digital marketing, "geo" is also used for location targeting, geofencing or local search, which is where the confusion comes from. If your goal is to appear in local results or on Google Maps, what you are looking for is local SEO.
Who coined the term GEO?
A team of six researchers whose first author is Pranjal Aggarwal, with authors from Princeton University and the Indian Institute of Technology Delhi. They posted "GEO: Generative Engine Optimization" on arXiv in November 2023 and presented it at KDD 2024, an ACM conference on data mining. Alongside the paper they released GEO-bench, a bench of 10,000 queries for measuring visibility in generative engines.
Does GEO replace SEO?
No. For a generative engine to use your page, it first has to retrieve it, and that depends on crawling, indexing and access for its crawlers. Google says explicitly that optimising for its generative AI features is still SEO. GEO adds a layer on top: sourced data, attributed quotations and trust signals that tip the choice between pages already in contention.
Does keyword stuffing help you appear in AI answers?
Not in the study that coined the term. Adding more words from the query left visibility at 17.7 against 19.3 for the untouched page, and on Perplexity it scored 10% below the baseline. The techniques that worked added verifiable information: quotations, statistics and references to sources. Google also treats keyword stuffing as spam4.
What to do next
Pick a page that already gets traffic and review what depends on it: figures with sources, attributed quotations, author and date. Then move on to the other blocks in the AI search optimization guide.
Sources
- 1GEO: Generative Engine Optimization, Aggarwal, Murahari, Rajpurohit, Kalyan, Narasimhan and Deshpande, KDD 2024 (arXiv), updated 28 June 2024.
- 2Optimizing your website for generative AI features on Google Search, Google Search Central, updated 10 July 2026.
- 3Overview of OpenAI Crawlers, OpenAI, accessed 7 October 2026.
- 4Spam policies for Google web search, Google Search Central, updated 28 August 2026.
How to cite this article
SmoothSeen. (2026, October 7). What is GEO (generative engine optimization)? Origin, the study behind it and how it differs from SEO. https://smoothseen.com/en/blog/what-is-geo/
Keep reading
AI search optimization: a guide to AEO and GEO for getting cited by ChatGPT, Gemini and Google
What AI search optimization (AEO and GEO) is, how ChatGPT, Gemini and Google pick sources, which bots to allow, plus a block-by-block checklist.
AEO vs GEO vs SEO: what each one is and how they differ
AEO vs GEO vs SEO: what each one aims for, where the result appears, which signals matter and how each is measured. With a table and an example.
AI crawlers and robots.txt: how to block GPTBot without dropping out of AI answers
Which AI crawlers OpenAI, Anthropic, Google, Perplexity, Apple, Meta and Amazon use, which to block in robots.txt and how to see if your server stops them.