SMACK HAPPY BLOG
Google Still Writes the Rules for AI Search. Here’s the Checklist.
Search has had one standard-setter for two decades, and an AI-generated summary at the top of the results page hasn’t changed that. Every SEO practice in wide use today exists because Google decided, at some point, what mattered and what didn’t, and the rest of the industry built around it. AEO and GEO get pitched like an unmapped frontier, which happens to be exactly the story an agency needs to tell if it wants to sell something new. But the company that decides what gets crawled, indexed, and shown to billions of searches a day just published, in its own documentation, what matters for its AI features and what doesn’t. A vendor’s guess about where things are headed doesn’t carry the same weight.
Here for a quick reference instead of the full read? Jump straight to the checklist or the bot glossary.
For the bigger picture on why discovery now happens on platforms long before anyone reaches a website, see The New Path to Yes: A Guide for the Modern Website Buyer Journey.
That authority has limits, and they matter here. Google sets the terms for AI Overviews and AI Mode because it owns that index and that ranking system. Its reach thins out fast past its own product. ChatGPT, Perplexity, and other AI tools draw on their own training data and their own retrieval, and they answer to different signals. Mostly what gets written about a brand elsewhere on the web matters more to them than what a site does technically. Treat every AI tool like the same animal and that gap disappears exactly when it matters most.
This isn’t an endorsement. Google runs an advertising business built on people staying inside its own results pages, and it has every reason to frame its AI features as additive rather than as the thing quietly eating into classic search traffic. Whether the documentation is accurate is a separate question from whether the company deserves anyone’s trust in general. Three claims in this post rest on that documentation: what robots.txt controls, what the Performance report measures, and whether llms.txt does anything at all. Independent research and crawler logs confirm all three. The documentation gets cited here because it’s the most reliable source available on Google’s own index, not because the company running it has earned any goodwill.
One quick housekeeping note before going further. The industry already calls this AIaaS, AI as a Service. We call it AaaS instead, dropping the I, because that missing letter says something true about the category. These tools produce probable outputs, not informed ones. Ordering from Instacart doesn’t make you a better cook (credit to entrepreneur Max Hofert for that one, we’ve never found a cleaner way to say it). You’ll see AaaS standing in for the generic stuff throughout this post. The exceptions are Google’s own product names, AI Overviews and AI Mode, because that’s what Google calls them, and a few industry acronyms we’ll spell out the first time each one shows up.
One more ground rule. Whatever grander name the industry lands on next, Super Intelligence, Super-dee-duper-intelligence, whatever shows up in the next presser, we won’t be using that one either. Terminology that needs a promotion before it can do the job usually can’t do the job.
With that settled, here’s what Google put in writing. Two documents matter here, a short page on how AI Overviews (AIO) and AI Mode work with a site, and a longer guide on optimizing for generative AI search overall. We read through both. Here’s what to act on.
- There’s no separate playbook for AI in search
- The single biggest signal is content nobody else could have written
- Every site is included in AI features by default, unless it was deliberately excluded
- Google’s own documentation says no new AI text files or special markup are needed to appear in AI Overviews or AI Mode
- A blunt block-AI-training setting can quietly take out normal Google or Bing crawling too
- A lot of what AEO and GEO agencies are selling doesn’t line up with Google’s own guidance
Nothing here requires a new strategy
Google’s guide undercuts its own premise almost immediately. AIO and AI Mode pull from the same index and the same ranking systems as regular Google results, and the documentation says so without hedging. No additional requirements. No special optimizations. Two things happen behind the scenes that didn’t happen before. The systems retrieve and summarize real pages instead of generating an answer out of nothing, which is why sources get cited at all. They also quietly run a handful of related searches before answering. A query about a lawn full of weeds might also pull results for herbicides, or for pulling weeds by hand. Neither trick requires a separate playbook. A site solid on fundamentals is already in the running.
Say something only you could say
Google draws one hard line between commodity content and content built on lived experience. Commodity content is the kind of listicle advice anyone could write, or any chatbot could generate in a few seconds flat. Google’s own example makes the point almost too well. It sets “7 Tips for First-Time Homebuyers” against “Why We Waived the Inspection and Saved Money: A Look Inside the Sewer Line.” Anyone who’s never bought a house could have written the first one. Only someone who lived through waiving an inspection could have written the second. That distinction, according to Google, does more for visibility than anything else in the guide, and it happens to be the same advice we’ve given clients for years, long before anyone called it AIO or GEO.
The technical basics haven’t changed
There’s nothing exotic here, just fundamentals that mattered before anyone slapped a trendy acronym on them. A site still needs to be crawlable and properly indexed. Page speed and mobile behavior matter exactly as much as they always did. Duplicate content still costs crawl budget and still confuses Google about which version of a page deserves to rank, so cleaning it up still pays off.
Feed Google the details if you’re local or selling something
For local businesses and ecommerce clients, an AaaS answer can pull straight from Google Business Profile and Merchant Center feeds. Keeping those current shapes what the answer says about a business or a product, not just how the listing looks on the page.
Check the one report that tells you if any of this is working
Every property in Search Console defaults to “Include my site’s links and content in Search generative AI features.” Google finished rolling that control, and the reporting behind it, out to every site worldwide on August 31. Nothing needs switching on. The thing to check, once, is whether a property inherited an exclusion from a parent property, or got switched off manually at some point. That setting lives under Settings, then Search generative AI.
Once that’s confirmed, pull the Generative AI performance report. It sits as its own report in the interface, though the numbers underneath it come from the same Web search type as the regular Performance report. An empty report almost always means low impressions, not a broken setting, unless the exclusion check above turned something up.
Skip the llms.txt theater
This is the one that stings. Google’s documentation says it outright. No new machine-readable files, no AI text files, no special schema.org markup needed to appear in either feature. Gary Illyes, from Google’s Search Relations team, has confirmed the same thing more specifically: Google doesn’t support llms.txt and isn’t planning to. John Mueller has separately compared the file to the old meta keywords tag, a field where a site owner could claim anything with nothing forcing it to be true. Crawler behavior backs all of that up.
| What SE Ranking found | Number |
|---|---|
| AI bot requests tracked over a 90-day window | 500+ million |
| Of those, requests that targeted llms.txt directly | 408 |
| Domains analyzed for a link between llms.txt and AI citations | ~300,000 |
| Of the 50 most AI-cited domains, how many had an llms.txt file | 1 |
llms.txt has been a standard part of our AaaS-visibility work for clients. The file can still make sense for clients chasing visibility on other AaaS platforms that reportedly do reference it. As a play aimed at Google specifically, it does nothing.
A few other circulating tactics fare no better. Breaking content into tiny chunks for AaaS to parse, and rewriting copy to match how people phrase AaaS queries, both assume Google’s systems need help understanding a page. They don’t. Synonyms and phrasing variations get handled without any of that. Chasing mentions across blogs and forums as a visibility hack belongs in the same bin. Google’s ranking systems weigh the quality of a source, not how many times a brand gets namedropped across the web. Structured data survives the cut, though not because it helps AIO or AI Mode specifically. It still earns rich results in regular Search, and that alone justifies keeping it.
Robots.txt decides who gets in
Google’s own list of best practices for AI features says it plainly. Make sure crawling is allowed in robots.txt, and by any CDN or hosting setup sitting in front of it. llms.txt can’t enforce anything. It can only suggest. Robots.txt can say no, and every major AI crawler except one well-documented holdout honors it. The real decision was never whether to let AI in or keep it out. It’s choosing, piece by piece, whether content should train a model, get cited in a live AI answer, do both, or do neither. Robots.txt is how that choice gets made, crawler by crawler, because different bots are built to do different jobs.
Most sites land on the same pattern OpenAI’s own documentation points toward: allow the crawlers that power live citations, and make a deliberate call on the ones that only train models. One detail worth knowing before touching any of it. Googlebot, Applebot, and Bingbot are multi-purpose, so a blunt block on AI training can quietly take out ordinary Google or Bing crawling too, and nobody notices until traffic drops.
| Crawler | What it’s for | Usual call |
|---|---|---|
| OAI-SearchBot, Claude-SearchBot, PerplexityBot | Fetches pages to build citations inside a live AI answer | Usually left open |
| GPTBot, ClaudeBot, CCBot | Trains the underlying model rather than answering a specific query | Business decision, not a technical default |
| Googlebot, Applebot, Bingbot | Multi-purpose: the same crawler often feeds both regular search and AI training | Allow for search; check before blocking AI training |
Full definitions for every bot named in this post are in the glossary at the end.
Cloudflare shows the infrastructure catching up in real time. About a year after launching a single toggle to block AI bots from scraping sites without permission, the company split it into three separate controls: Search, Training, and Agent. Accounts that already had the old “Block AI Bots” or “Managed Robots.txt” setting turned on were migrated automatically into the new categories, with any manual changes preserved. Domains signing up for Cloudflare for the first time after September 15, 2026 start from a different default:
| Control | Default for new sign-ups |
|---|---|
| Search | Allowed |
| Training | Disallowed on pages that carry ads |
| Agent | Disallowed on pages that carry ads |
A CDN isn’t a requirement here. Robots.txt alone gets a site most of the way to a deliberate choice instead of a default nobody picked. A CDN like Cloudflare sits on top of that and needs its own check regardless of what robots.txt already says. A CDN-level block can override robots.txt entirely, and a site owner can go a long time without noticing.
Watch for agencies selling AEO and GEO like it’s a brand-new category
A wave of agencies has rebranded around AEO (Answer Engine Optimization) and GEO (Generative Engine Optimization) over the past year, and a good chunk of what they’re selling doesn’t survive contact with what Google’s guide says. One common pitch frames the service as the thing that will separate the brands that thrive from the ones that fade, a race only their particular package can win. Google’s own guidance cuts that framing off at the knees. There’s no separate playbook for AI in search to sell.
Individual line items don’t fare much better under scrutiny. Rewriting content specifically for AaaS answer quality shows up in nearly every pitch deck, even though Google’s guide says content needs no special-case rewriting at all. Question-based keyword mapping follows the same logic. It rests on the idea that phrasing for a chatbot demands its own keyword strategy, when Google already handles synonyms and intent without one. And structured data gets marketed as an AI-citation feature more often than it should be, when what it mostly earns is a regular rich result.
Independent research backs this up too. See the numbers under Skip the llms.txt theater above. Whatever a pitch implies, having the file in place isn’t an insider’s advantage.
That doesn’t make AEO and GEO meaningless as concepts, or every agency using those words a con artist. It means reading the pitch against Google’s own guidance before signing anything, especially the parts arriving with a deadline attached.
We’ve made this same point in that guide too. No agency, freelancer, or platform can promise a guaranteed spot inside an AI answer nobody controls, whether the pitch is dressed up as AEO, GEO, or whatever comes next.
Tools for checking the work
Nothing above needs to be taken on faith.
Semrush has a free AI Visibility Checker that shows whether a domain is showing up in AI answers at all, and its Site Audit tool includes an AI Search Health score that flags when robots.txt or meta tags are blocking the crawlers that matter. Semrush’s own analysis of Google’s guide makes a point Google’s own documentation leaves out entirely. Platforms like ChatGPT and Perplexity draw on their own training data and retrieval rather than Google’s index, so a brand’s reputation on review sites, Reddit, and Wikipedia can end up mattering as much as anything on the site itself.
SparkToro answers a different question. It shows which sources an audience trusts, and where it spends its attention, broken down by website, podcast, subreddit, even which AI tools people use. It has nothing to do with technical SEO. What it tells you is something neither tool above does, where an audience really goes, and whether the content shows up there.
The checklist
- Check that the site hasn’t been excluded from Search generative AI features under Settings, especially on properties that inherit the setting from a parent
- Pull the Generative AI performance report and see where the site currently stands
- Flag any page that reads generic or replaceable, and rewrite it around real, specific experience
- Check that headings and structure read cleanly for an actual person, not just for keywords
- Verify Google Business Profile and Merchant Center feeds are current, for local and ecommerce clients
- Confirm the site is crawlable, mobile-friendly, and free of obvious duplicate content
- Keep schema markup. It still earns rich results, just not AaaS visibility specifically
- Reframe or drop llms.txt as a Google-specific deliverable, depending on which platforms a client cares about
- Set robots.txt to allow the crawlers that power live citations (OAI-SearchBot, Claude-SearchBot, PerplexityBot) and make a deliberate call on the ones that only train models (GPTBot, ClaudeBot, CCBot)
- If the site sits behind Cloudflare or another CDN, check its AI-crawler settings directly, since a CDN-level block can override robots.txt without anyone noticing
- Weigh any AEO or GEO vendor pitch against what Google’s own guidance says before signing anything
- Run a free AI visibility check (Semrush’s AI Visibility Checker or similar) to confirm whether the site is showing up in AI answers today
- Pull an audience snapshot (SparkToro or similar) to see which sites, communities, and AI tools the audience uses
That’s the whole list. Not one item on it requires a department of its own, whatever acronym eventually wins.
Bot glossary
Every crawler named in this post.
| Bot | Company | What it does |
|---|---|---|
| Googlebot | Multi-purpose: powers regular Search and feeds Google’s AI features | |
| Applebot | Apple | Multi-purpose: powers Siri, Spotlight, and Apple’s AI features |
| Bingbot | Microsoft | Multi-purpose: powers Bing Search and Copilot |
| OAI-SearchBot | OpenAI | Fetches pages to build live citations inside ChatGPT answers |
| GPTBot | OpenAI | Crawls to train OpenAI’s underlying models |
| Claude-SearchBot | Anthropic | Fetches pages to build live citations inside Claude answers |
| ClaudeBot | Anthropic | Crawls to train Anthropic’s underlying models |
| PerplexityBot | Perplexity | Fetches pages to build live citations inside Perplexity answers |
| CCBot | Common Crawl | Builds the open web dataset a number of AI labs train on |