How to get your content cited by Google AI Overviews, ChatGPT, Perplexity & Claude

Marketing Matters
10 min read September 8, 2026
sid
Siddhant SEO Specialist

A buyer types "best Drupal migration agency" into ChatGPT. It answers in seconds, complete with a citation, and the buyer never opens a browser tab.

AI Overviews, ChatGPT, and Perplexity now answer many buyer questions before anyone reaches a website. In this playbook, we shall discuss what's changed in Google's AI search stack, how AI crawlers actually work, and what research says earns a citation. AEO (Answer Engine Optimization) wins visibility inside Google's AI Overviews and AI Mode; GEO (Generative Engine Optimization) wins citations from standalone LLMs like ChatGPT, Perplexity, and Claude. Use it to plan your content and technical work for both.

Key Takeaways

  • AEO wins visibility in Google's AI Overviews and AI Mode; GEO wins citations from standalone LLMs like ChatGPT, Perplexity, and Claude.
  • Different mechanisms, same winning content pattern.
  • The pattern that works everywhere: answer-first structure, real evidence (stats, quotes, sources), clean access for AI crawlers, and demonstrable authorship.
  • Keyword-stuffed, thin, or anonymous content is now actively excluded from AI answers, not just demoted in rankings.
  • Referral traffic from Search has already dropped roughly 60% for small sites and 22% for large ones since AI Overviews launched. Being cited inside the answer is the new page-one ranking.
  • Google's own documentation says no special AEO tricks are required, but page structure and authority still clearly shape which pages get pulled into synthesized answers.

Why does this matter now?

Traditional SEO optimizes for a ranked list of links. AEO and GEO optimize for something different: whether an AI system quotes, paraphrases, or cites your content directly inside a synthesized answer, often with no click required at all. This is already happening. Google's AI Mode alone passed 1 billion monthly users in 2026, with query volume doubling every quarter, and referral traffic to publishers from Search has already dropped roughly 60% for small sites and 22% for large ones since AI Overviews launched.

What's changed recently?

A few dated, concrete developments that change how we should plan work right now:

Google AI Mode redesign (May 2026 I/O)

Search now defaults to Gemini 3.5 Flash in AI Mode. Google explicitly uses "query fan-out": firing multiple related sub-queries behind the scenes to build a fuller answer. Practical implication: a page that only targets one exact query is no longer enough; comprehensive coverage of a topic's adjacent questions gets pulled into more answers.

Search Console's new Generative AI performance report (launched June 3, 2026)

Performance → there's now a "Generative AI" tab alongside "Web," showing impressions from AI Overviews and AI Mode by page, country, and device. Limitation: impressions only; no clicks, CTR, or query-level data yet. Still the only first-party visibility signal Google gives us, and worth checking weekly.

Google's official line: no special AEO tricks required

Google's own developer documentation states there are "no additional requirements to appear in AI Overviews or AI Mode, nor other special optimizations necessary." It comes down to standard indexability, crawlability, structured data, and helpful-content fundamentals. True, but incomplete: which pages get pulled into the synthesized answer clearly correlates with authority and structure, which is exactly what the GEO research below investigates.

March 2026 core update: AI content and E-E-A-T tightened

A clear pattern emerged: "AI-assisted + heavy human editing + real examples" content holds or gains; "AI-drafted + light editing + generic coverage" declines sharply. Author credentials now matter more: anonymous content and generic bios are losing visibility, especially on YMYL-adjacent topics. "Information gain" (original data, first-hand experience, real case studies) is now weighted above content that merely rephrases what already ranks. Thin, purely summarized content is being excluded from AI Overviews specifically, beyond being demoted in blue links.

Publisher "Subscribed" labels in AI Overviews (May 2026)

Google now visually flags citations from publications a user is subscribed to, with measurably higher click-through than unlabeled citations. Confirms Google's own data that a citation does not automatically mean a click: visibility and traffic are becoming two separate things to track.

How do AI crawlers work?

This is the part most teams get wrong. There is no single "AI bot." Each company runs at least two, doing fundamentally different jobs:

Training crawlers: scrape content to build the next model version. One-off; doesn't affect what the current live model can cite today.

Retrieval / search-index crawlers: build the live index the assistant searches in real time. Blocking this is what actually removes you from citations.

User-triggered fetchers: activate only when a live user asks the assistant to open a specific URL.

CompanyTraining crawlerRetrieval/citation crawlerUser-triggered fetcher
OpenAIGPTBotOAI-SearchBotChatGPT-User
AnthropicClaudeBotClaude-SearchBotClaude-User
GoogleGoogle-Extended (opt-out token, not a crawler)Googlebot (also powers AI Overviews)
PerplexityPerplexityBotPerplexity-User
OthersApplebot-Extended, Meta-ExternalAgent, Bytespider, CCBotAmazonbot

Most common mistake

Blocking GPTBot in robots.txt thinking it stops AI scraping, without realizing OAI-SearchBot is a completely separate crawler that controls ChatGPT citation access. The other common error: a blanket "User-agent: *" disallow rule that quietly takes out Googlebot too.

Note: Bytespider (ByteDance) has a documented history of ignoring robots.txt entirely regardless of what's set.

Recommended stance

Decide separately and deliberately whether you want (a) training access and (b) citation/retrieval access. For example, for Specbee, the default should be: allow all retrieval/search crawlers (OAI-SearchBot, PerplexityBot, Claude-SearchBot, Googlebot) unconditionally; treat training-crawler blocking as a separate business decision rather than an SEO one.

What drives AI citations?

The original GEO research (Princeton / Georgia Tech / IIT Delhi, arXiv:2311.09735) rewrote the same content with different tactics and measured visibility in AI-generated answers. Ranked by effectiveness:

RankTacticWhy it works
1Add expert quotationsAttributed statements ("…," says [named expert]) were the single strongest lever tested.
2Add statisticsHard numbers beat vague claims, e.g., a specific ROI figure over "strong returns."
3Improve fluencyCleaner, better-structured writing gets selected more, independent of the facts.
4Cite sourcesNamed, dated attribution ("according to [Org]'s [Year] report") adds a verifiable anchor.
5Use technical termsDomain vocabulary helps specifically on specialized / niche queries.

Lower down the list: authoritative tone, simplification, and unique vocabulary had moderate effect. Keyword stuffing was the only tactic that actively hurt visibility: the opposite of classic SEO instinct. The common thread across everything that worked: it made the content read like evidence rather than marketing copy.

One more finding worth internalizing: AI systems measurably favor fresher content. One study found cited sources were roughly 25.7% fresher than average. Quarterly refreshes of cornerstone content are now a citation-eligibility factor, in addition to good housekeeping.

Myth to retire: llms.txt

llms.txt adoption sits around 10% of sites, and none of the major platforms, not Google, not OAI-SearchBot, not Perplexity, not Bing/Copilot, use it as a citation input. One analysis found that removing llms.txt as a variable actually improved a model predicting citation frequency, meaning it's pure noise. Only add it if your CMS generates it for free; don't spend real effort here.

The step-by-step action plan

Step 1: Technical foundation (do this first; it's free and structural)

Confirm indexability and crawlability using Search Console's URL Inspection tool. Audit robots.txt against the crawler table above. Add Article/BlogPosting + FAQPage + BreadcrumbList schema. Validate all of it in Google's Rich Results Test.

Step 2: Restructure content for extraction

Every H2 leads with a direct 40–60 word answer before expanding. Keep sections to 200–400 words as self-contained, independently citable units. Use comparison tables and numbered lists liberally. Define key terms/entities on first mention.

Step 3: Load content with what the research says gets cited

Add a specific, named-source statistic every 150–200 words. Add real expert quotes (interview an actual architect/developer rather than fabricating attribution). Cut marketing framing and reframe as evidence.

Step 4: Build topical authority rather than one-offs

Use a pillar-and-cluster structure with descriptive internal anchor text. Publish consistently across core service areas. Refresh cornerstone content quarterly given the freshness bias.

Step 5: Fix credibility signals the March 2026 update checks for

Use real author bylines with credentials in place of generic bios. Layer visible human editing and original examples into any AI-assisted drafts.

Step 6: Measure honestly

Check Search Console's Generative AI tab weekly for impressions by page (remember: impressions only, which falls short of proof of traffic). Use a third-party citation tracker, such as Profound, Otterly.ai, LLMrefs, or Semrush's AI toolkit, for cross-platform citation visibility. Manually spot-check target queries directly in ChatGPT/Perplexity monthly.

Work these steps in this order: technical access first, since no crawler can cite a page it can't reach; extractable structure and evidence next, since that's what earns the citation once the page is reachable; measurement last, to confirm it's working.

Final thoughts

AEO and GEO aren't separate campaigns from SEO. They're what SEO looks like now that AI systems stand between many buyers and the sites they used to click through to. The technical checklist, content structure, and citation-worthy evidence in this playbook reinforce each other.

The teams that move first get an advantage while it still exists. Quarterly content refreshes, real author credentials, and open access for retrieval crawlers are cheap to implement and hard for competitors to copy quickly. Start with the technical foundation, then build out expert quotes and FAQs, since those are the fastest wins. Our team of SEO experts already runs AEO and GEO audits for Drupal sites. Talk to us about turning this playbook into an execution plan for your site.

Frequently Asked Questions

What's the difference between AEO and GEO?

AEO optimizes for Google's AI Overviews and AI Mode, which sit on top of Google's existing search index. GEO optimizes for standalone LLMs like ChatGPT, Perplexity, and Claude, which crawl and retrieve content on their own. Both reward the same core pattern: answer-first structure, real evidence, and demonstrable authorship.

How much traffic am I actually losing to AI Overviews?

Referral traffic to publishers from Google Search has already dropped roughly 60% for small sites and 22% for large ones since AI Overviews launched. Google's AI Mode alone passed 1 billion monthly users in 2026, with query volume doubling every quarter. Being cited inside the answer is quickly becoming the new page-one ranking.

How do AI crawlers access my content, and can I control it?

Each AI company runs at least two separate crawlers: one that scrapes content to train future models, and one that retrieves live pages for citations. Blocking the training crawler doesn't affect your citation access at all. To stay visible in AI answers, allow the retrieval crawlers: OAI-SearchBot, PerplexityBot, Claude-SearchBot, and Googlebot.

Does blocking GPTBot in robots.txt keep me out of AI citations?

No, GPTBot only controls OpenAI's training crawler. ChatGPT's citation access runs through a separate crawler, OAI-SearchBot, so blocking GPTBot alone leaves your citation visibility untouched. Check each platform's retrieval crawler before you decide what to disallow in robots.txt.

What actually gets content cited by AI systems like ChatGPT and Perplexity?

Research from Princeton, Georgia Tech, and IIT Delhi ranked expert quotations and statistics as the two strongest levers for AI citation, tested across the same content rewritten in different ways. Clean, fluent writing and named source attribution came next. Keyword stuffing was the only tactic that actively hurt visibility, the opposite of standard SEO instinct.

Do I need an llms.txt file to get cited by AI search engines?

Not for most sites. llms.txt adoption sits around 10%, and none of the major platforms, including Google, OAI-SearchBot, and Perplexity, use it as a citation input yet. One analysis even found that removing it as a variable improved a model's citation predictions. Add it only if your CMS generates it for free.

Last updated on September 8, 2026

About the author

sid

Siddhant

Meet Siddhant Gyawali, a skilled SEO/SEM Specialist and ardent DC comics enthusiast. He’s passionate about staying updated on new AI developments. A former state and national-level basketball player, Sid was a finalist for the Australian Web Marketing Awards and a Global Search Awards winner. Sid dreams of exploring Bhutan with his family.

Read more about Siddhant