Blog

Optimise Content for AI Search

Sunny Patel

Sunny Patel

SEO Consultant & AI Strategist

Optimise Content for AI Search

AI search is the default experience for a growing share of queries: Google AI Overviews, ChatGPT Search, Perplexity, and Bing Copilot all generate answers directly from web content, often without sending a click. A portfolio-wide audit of my own sites found 120,409 Bing Copilot citations across 11 sites in one pull, a 57:1 ratio against organic clicks from the same pages. This post covers what actually earns those citations in August 2026, not what earned them a year ago. The mechanics keep changing underneath the same recycled advice.

GEO, AEO, and SEO Are Three Different Jobs Now

Generative Engine Optimisation (GEO) is the practice of getting your brand or content cited inside AI-generated answers from ChatGPT, Perplexity, and Gemini. Answer Engine Optimisation (AEO) is the narrower practice of getting a specific passage extracted as a direct answer: an AI Overview, a featured snippet, a voice result. Traditional SEO earns ranking positions in a list of ten blue links.

These overlap but they are not the same job. Treating them as one strategy is where most sites lose ground. A page can rank position 3 on Google and never appear in an AI Overview. A page can sit at position 11, invisible to anyone scrolling results, and still earn dozens of ChatGPT citations a day. It answers the question directly, which is what the model rewards. The skill that spans all three is structuring content so a machine can lift out a correct, self-contained answer. Everything below is about that skill.

A Top-10 Ranking No Longer Guarantees an AI Citation

One 2026 analysis of AI Overview citation patterns found that 76.1% of cited URLs still came from the traditional top 10 results. Ranking well remains the strongest single predictor for Google specifically. But a separate large-scale study by AI Mode Boost, covering over 15,000 AI Overview results across 63 industry verticals, found the correlation with traditional ranking position had fallen to r=0.18, close to no relationship. 47% of cited passages came from pages ranking below position 5.

Both can be true at once. Google mostly cites from its own top 10. Which page within that top 10 gets the citation has almost nothing to do with rank order. The same study found content scoring 8.5 out of 10 or higher on a semantic completeness measure was 4.2 times more likely to be cited than lower-scoring pages on the same topic. Rank gets you into the pool. Structure decides who gets pulled out of it.

How Google, ChatGPT, and Perplexity Actually Choose What to Cite

Each system selects sources differently. Conflating them is the fastest way to waste effort.

Google AI Overviews synthesise an answer in real time for each query: scanning indexed content, assessing structure and claim clarity, and assembling a response from multiple sources. Google has not published the exact selection logic. Its own Search Central guidance points back to standard helpful-content and E-E-A-T principles rather than a separate AI ranking system.

ChatGPT weighs content-answer fit heavily: how precisely a passage answers the exact question asked, not how many keywords or backlinks the domain has. Brand mentions across the open web (being named and described in other people's content, not just linked to) correlate roughly three times more strongly with ChatGPT citation than backlink counts do.

Perplexity runs a retrieval pipeline with several distinct stages: a broad initial retrieval using keyword and semantic matching, a cross-encoder that scores query-document pairs together, then a final reranker weighing entity signals, domain authority, freshness, and source diversity. Of roughly ten pages it retrieves for a query, only three or four typically get cited in the answer. Perplexity also leans on Reddit far more than Google does: Reddit accounts for close to half of Perplexity's citations by some measures, against a low single-digit share of Google AI Overview citations. Optimise for one platform's citation logic, assume it transfers to the others, and effort gets misallocated fast.

What the Reddit-ChatGPT drop in August 2026 shows about how fast this moves

Reddit's share of ChatGPT Search citations fell from an average of 3.83% (18 July to 7 August) to 0.52% (14 to 17 August), an 86% relative decline. The drop coincided with a change to how ChatGPT generates its background search queries on 8 August. Nobody, including Reddit and OpenAI, has given a confirmed reason. Whatever caused it, a source that supplied roughly one in twenty-five ChatGPT citations dropped to roughly one in two hundred within a week, with no warning and no announcement.

That is the honest state of GEO in 2026: you are optimising against a system whose selection logic can shift materially inside a single week, for reasons the platform itself has not explained. Build on documented, current mechanics. Expect to revisit them.

The Structural Unit AI Systems Extract: Self-Contained Answer Chunks

The content pattern that survives across all three platforms is the self-contained answer chunk: a passage of roughly 134 to 167 words that fully answers one specific question without requiring the reader to have read anything before or after it. Every model in this space extracts passages, not pages. A page built as one long argument that only makes sense read start to finish gives the model nothing clean to lift.

Two formatting rules follow directly from that:

  • Answer first, context second. State the direct answer in the first sentence of a section. Add the supporting detail, caveats, and method after it, not before.
  • Bold question, plain-sentence answer. A heading formatted as What is [term]? followed immediately by a one or two sentence definition is the clearest signal you can give a model about where an extractable answer starts and ends.

The pattern holds in my own data. When I compared high-citation and low-citation pages across the portfolio behind that 120,409-citation figure, the winners were technical troubleshooting guides, specification pages, and research compilations with named sources, structured as one precise answer per section. The losers were opinion-led reviews and listicles with no supporting data. Specificity, structure, and cited authority, in that order, is still the shortest description I have of what gets lifted.

Does Schema Markup Still Help You Get Cited

Yes, with one important 2026 correction. Google and Microsoft have both confirmed their AI systems use structured data to understand and verify page content. Content with correctly implemented schema is associated with roughly 2.5 times higher odds of appearing in AI-generated answers, per one recent study. Schema tells a model what type of content it is looking at, who wrote it, and what it claims, without the model having to infer that from prose.

The correction: Google removed FAQ rich results from standard search entirely on 7 May 2026. If you added FAQPage schema purely to win that rich snippet, that specific payoff is gone. FAQPage and QAPage schema still matter for AI citation. AI systems keep parsing the structured question-answer pairs for context even without a visible rich result attached. Add it for the AI layer, not for the SERP feature that no longer exists. Pair it with Organization and Person schema and consistent sameAs links across your other profiles. Entity credibility, the degree to which a system recognises your brand as a known, verifiable source, still measurably affects which of two similar pages gets the citation. My guide on adding schema markup to any platform covers implementation.

Does llms.txt Actually Help You Get Cited

Mostly no. llms.txt gets recommended reflexively, well ahead of what the data supports. llms.txt is a proposed text file, similar to robots.txt, that lists a site's key pages in a format meant to be easy for an LLM to read. Adoption reached roughly 10% of domains analysed in one 300,000-domain study by early 2026. The same study built a machine learning model to test whether having an llms.txt file predicted AI citation frequency. Removing the llms.txt variable from the model improved its prediction accuracy, meaning the file's presence had no positive correlation with citations. Google's John Mueller has said none of the AI providers have confirmed using it for retrieval.

Where it does earn its keep: AI-assisted coding tools such as Cursor and Continue do read llms.txt to navigate a site's documentation. If developers are a real audience for your site, add it for them. Do not add it expecting an AI search citation lift. The current data does not support one.

How I Measure Whether Content Is Actually Getting Cited

Three sources, none of them complete on their own:

  1. Bing Webmaster Tools still reports AI Copilot citation counts per page and per query, which is how I found the 120,409-citation figure above. It is the most direct citation-level data available from any platform.
  2. Google Search Console began rolling out a dedicated Generative AI performance report in June 2026, under Performance, showing impressions for AI Overviews and AI Mode separately from standard web results. It reached UK site owners first. Click data is not in it yet, only impressions. Treat it as visibility confirmation, not a full picture.
  3. GA4 referral sessions from chatgpt.com, perplexity.ai, and similar sources measure the traffic floor, the users who actually clicked through. Between roughly 35% and 70% of genuine AI referral sessions arrive with no referrer header and land in Direct traffic instead, depending on platform. GA4 will always undercount this channel as a result. I run this measurement quarterly across the portfolio with a documented method rather than quoting industry-wide averages. The full breakdown, click floor measured against citation volume for real sites rather than projected ones, is in my AI referral traffic study.

None of these three tools alone tells you whether AI search is working for a given page. Citations without clicks still put your brand in front of the person asking the question. Clicks without citation visibility mean you are flying blind on everything AI search does that never resolves into a session.

What To Do With a Page Like This One

Check whether your highest-value pages answer their target question inside the first sentence of the relevant section. Check whether each section stands alone as roughly 150 words of self-contained answer. Check whether Organization and Person schema is current and consistent across your other profiles. Then check Bing Webmaster Tools for citation data you are probably not looking at yet. I cover all of this as part of every technical SEO audit I run, plus dedicated AI search optimisation work for sites where citations, not rankings, are the target. Get in touch if you want a second pair of eyes on where your pages are losing AI visibility.

Related Articles

Free Resource

Free SEO Audit Checklist

The same 47-point checklist I use for client audits. Covers technical SEO, content gaps, and quick wins you can fix today.

One email with the checklist. No mailing list.

Get Started

Ready to grow your organic traffic?

Request a free 20-minute SEO diagnosis: focus on the biggest issue and the most useful next step.

Request Free Diagnosis
15+ years experienceFree 20-minute diagnosisNo contracts

Usually responds within a few hours