In one sentence: Content still wins, but the unit that wins is no longer a page that matches a keyword; it is a fact an AI engine can lift, check against other sources, and attribute to you.

Here is the short version. A classic search engine is a librarian. It walks every aisle of an infinite library, writes a card for each page, and when you ask a question it hands you the ten cards whose words match yours. So for twenty years the craft was matching words: put “water heater repair” in the title, in the heading, in the first paragraph, and repeat it until the librarian filed you under the right card.

An AI engine is a reader, not a librarian. When you ask ChatGPT, Gemini, Claude or Perplexity who to call, the engine fetches pages live, reads them, and writes an answer with two, three or four names in it. In a study of 200,085 local searches, ChatGPT named 4.1 businesses on average, Google’s AI Mode 3.5, and AI Overviews 2.5.[1] The reader does not want the page that says the words most. It wants the page that gives it a fact it can trust and reuse. Those are different objects.

What a reader rewards, measured

The first controlled study of this, from Princeton and the Allen Institute in 2023, tested nine ways of rewriting a page and measured how much of each version showed up in the generated answer across 10,000 questions. The winners were not keyword tricks. Adding quotations, adding statistics, and citing sources raised a page’s share of the answer by up to 41% on the objective measure and 28% on the subjective one.[2] Keyword stuffing scored below every other method tested. The authors put it plainly: “simple methods like Keyword Stuffing traditionally used in SEO don’t perform well.”[2]

A 2026 measurement of 602 prompts and 21,143 citations across ChatGPT, Google and Perplexity found the pages that actually shaped answers were “longer, more modular, more semantically aligned with the generated answer, and more likely to contain extractable evidence genres such as definitions, numerical facts, comparisons, and procedural steps.”[3] The same study found a thing many agencies still sell does nothing on its own: “Q&A formatting alone does not improve absorption.”[3] A question heading with a vague answer under it is decoration. A question heading with a number under it is evidence.

Readers also cite carelessly, which changes how a fact should be written. A 2023 audit of four generative search engines found only 51.5% of generated sentences fully supported by their citations, and only 74.5% of citations supporting the sentence they were attached to.[4] So a fact that lives in one self-contained sentence, with your name in that sentence, survives being half-quoted. A fact spread across a paragraph does not.

One more measured difference matters for a local business. A 2025 comparison of AI search against Google found the AI engines leaned overwhelmingly on third-party sources rather than a brand’s own pages.[5] The librarian was happy to hand out your own brochure. The reader wants to hear it from someone else first.

The two clocks: training and retrieval

Two different machines are involved, and most of the confusion in this industry comes from mixing them up.

The retrieval clock runs in days. Every major assistant fetches the web live for questions like “who should I call.” OpenAI documents a search crawler and a user-fetch agent; Anthropic documents the same pair; Google’s Gemini grounds its answers on Google Search and returns inline citations; Perplexity’s user fetcher goes out when you ask.[6][7][8][9] Change what those agents can read and the answers change within days to weeks.

The training clock runs in years. Models are built from filtered snapshots of the public web. The FineWeb dataset, for example, was cut from 96 Common Crawl snapshots into 15 trillion tokens after language filtering, quality filtering and deduplication.[10] The RefinedWeb team showed “properly filtered and deduplicated web data alone can lead to powerful models,” with roughly half the crawl removed as duplicate.[11] Nobody can place a business inside a model on demand. What a business can do is be present, in clean original prose, wherever the next snapshot is cut. Common Crawl’s own bot obeys robots.txt, reads the sitemap you declare, and executes no JavaScript.[12] Block it, or hide your facts behind scripts, and you are absent from the next generation of models.

Will the good content get you penalised for using AI?

No, and the platform that polices this hardest says so in writing. Google’s guidance is that “appropriate use of AI or automation is not against our guidelines.”[13] What Google penalises is a named spam pattern, scaled content abuse: “creating large amounts of unoriginal content that provides little to no value to users, no matter how it’s created.”[14] The offence is volume without value. The tool is not the offence.

That is worth reading twice, because it is the line between the content that gets blocked and the content that gets quoted. Thousands of pages spun from one template, city names swapped, no source, no author, nothing a competitor’s page doesn’t say: that is what the platforms are filtering out, and they are right to. An engine that trusts filler poisons its own product. The content a reader keeps is the opposite on every axis: original, first-hand, sourced, attributed to a person who can be held to it, and written so that a homeowner can follow it without a dictionary.

The comparison, then, looks like this. Old rules: match the words, repeat them, publish often, link from anywhere. New rules: state the fact once and clearly, give the number and its source, name who is speaking and why they know, keep one page per real thing, and get other people to say it about you. A page written to the old rules can still rank in a list. It will rarely be the page a reader quotes.

What we do about it, in order

This is the order of work, because each step is a gate the next one depends on.

1. Open the door. Every AI crawler and user-fetch agent allowed, in robots.txt and at the CDN. Google names the CDN explicitly as a place that must allow crawling.[15] We found two of our own sites blocking AI crawlers at the edge by default; the block never showed in robots.txt.

2. Get indexed and recognised. Google says there are no additional requirements to appear in AI Overviews or AI Mode; a page must be indexed and eligible, and standard practice applies.[15] Recognition means one consistent name, address and phone number on every profile the engines already cite.

3. Put the facts in the page itself. Who you are, what you do, where, for whom, at what price, at what hours, in plain sentences the reader can lift. Structured data that matches the visible words.[16]

4. Write evidence, not adjectives. A definition, a number with a source, a comparison, the steps of a job, a quotation from a real person. That is the recipe the measurements reward.[2][3]

5. Earn the third-party mentions. The directories, the review platforms, the local press, the supplier listings the engines cite for your trade and town.[5]

6. Measure it, per engine, every month. Engines differ from each other in freshness, in how many sources they cite and in how they react to a rephrased question.[5][3] One engine’s answer is never evidence about another’s.

Why now

Because the ad layer is arriving, the way it did in search. Perplexity began selling sponsored follow-up questions in November 2024, two years after launch, while stating that “the content of the answers you receive on Perplexity will not be influenced by advertisers.”[17] OpenAI’s own crawler documentation now lists an advertising bot alongside its search bot.[6] The pattern from search was an organic layer that persisted, an ad layer that arrived within a few years, and rising costs for whoever waited. The organic slot in an AI answer is earned by accumulated, verifiable presence, and that presence compounds. The businesses that are the recognised answer in their trade and town before the auction opens are the ones the engines keep naming after it does.

What we will not claim

No one can guarantee what an AI answers, and we don’t. No one outside the platforms knows the exact weighting inside them, and we don’t either; we measure the answers instead. And no one can sell you a place in a model’s training data; the money in that market flows the other way, from AI companies to publishers, as the OpenAI–News Corp agreement shows.[18] What we can show you is what the engines say about you today, who they name instead, and how that changes when the inputs change.