How AI engines decide which websites to check
Short answer: Google's AI features pull from Google's own Search index — a page must be indexed and eligible to show with a snippet to be cited. ChatGPT builds its pool from its own search crawler (OAI-SearchBot) plus third-party search providers. Both rewrite your question into several narrower queries before choosing sources. Neither publishes ranking weights, and Google says its Search ignores llms.txt.
Claims about what Google and OpenAI do are sourced to their own documentation, linked at the bottom. Where something is inference rather than documentation, it's labelled as such.
Ask ChatGPT or Google for a recommendation and you get a short answer with a handful of links. Somebody's website got pulled into that answer. Most of the time it isn't the business that spent the most on advertising — it's the one whose pages were reachable, indexed, and clearly useful on the exact question being asked.
There's a lot of guesswork circulating about how this works. Below is what OpenAI and Google actually document, and nothing beyond it. Where the answer is "they don't say," we say that too.
How does Google decide which websites to use in AI Overviews and AI Mode?
Google's AI features run on Google's own Search index. Google states that its generative AI features are rooted in its core Search ranking and quality systems, and it names two techniques:
- Retrieval-augmented generation (RAG), which Google also calls grounding — the core ranking systems retrieve relevant, up-to-date pages from the Search index, and the model builds its response from what those pages say.
- Query fan-out — a set of concurrent related queries the model generates alongside the original one. Google's own example: a question about fixing a lawn full of weeds may fan out into "best herbicides for lawns," "remove weeds without chemicals," and "how to prevent weeds in lawn."
Google doesn't say so directly, but fan-out is the mechanism that would let a page get cited without ranking on page one for the obvious keyword — it's answering one of the side questions instead.
AI Overviews don't appear on every search. Google says they show only when its systems determine an overview is additive to classic Search, so they often don't trigger at all.
One user-side lever worth knowing about: since May 2026, Google's preferred sources feature — where a searcher picks sites they want to see more of — also applies to AI Overviews and AI Mode.
How does ChatGPT decide which websites to check?
ChatGPT assembles its pool from more than one place. OpenAI documents a dedicated search crawler, OAI-SearchBot, whose stated job is to surface sites in ChatGPT's search features. OpenAI also states that ChatGPT search sometimes partners with third-party search providers; the ones it names in its ChatGPT Search help article are Bing and Shopify. ChatGPT decides on its own whether a question warrants a web search, and the user can also force one.
ChatGPT rewrites your question before searching. OpenAI's published example: a researcher asking about CCR8 cancer drugs might trigger an initial query of "CCR8 immunotherapy drug development 2025," followed by narrower queries after the model reviews what came back. Same principle as Google's fan-out, different plumbing.
For local questions, OpenAI says ChatGPT estimates general location from IP address and may share that general location with third-party search providers to improve result accuracy. It says it does not share the IP address itself, or your account information, with those providers.
What makes a page eligible to be cited?
Eligibility is where most local businesses lose before the ranking question is ever asked.
To be eligible as a supporting link in AI Overviews or AI Mode, a page must be indexed and eligible to appear in Google Search with a snippet. Google states there are no additional technical requirements — and also that meeting every requirement still doesn't guarantee crawling, indexing, or serving.
Worth knowing that Google's own two pages on this don't quite agree, and the newer one is stricter. The AI features page (last updated December 2025) says there are no additional technical requirements. The generative AI optimization guide (last updated July 2026) adds one: beyond the Search technical requirements, a site must be included in Search generative AI features in Search Console to be eligible for display. If you read only the first page, you'd miss the second gate entirely.
That second gate is real, but it isn't live for most sites yet. Search Console has a Search generative AI control under Settings, covering AI Overviews, AI Mode, and generative AI features in Discover. Sites are included by default. Excluding your site means no links, no grounding, no impressions, and no traffic from those features. Google is rolling the control out to a subset of website owners in the UK first, with a global rollout stated but not dated — so if you're in the US, you most likely won't see the setting yet. Worth knowing about before it reaches you, and worth checking the day it does, especially if a previous developer or agency has touched your Search Console.
Worth knowing how fast this has moved, too. Google originally gave a hard date — a note on the Search Console help page said it would begin taking the control into account on June 17, 2026. Archived snapshots show that sentence on the page on June 3 and again on June 21; by July 9 it was gone, and the live page now mentions only the limited rollout. Google retired the FAQ structured data documentation the same quiet way in June: the old URL still returns a 200, but it redirects to a changelog entry instead of the page it used to serve. We checked both against the Internet Archive rather than trusting our own earlier notes. That's the real lesson underneath all of this — the documentation moves without announcement, so any GEO advice quoting a date is worth re-checking at the source before you act on it.
OpenAI
OpenAI documents four separate agents, each controlled independently in robots.txt:
| User agent | What it does | Matters for AI visibility? |
|---|---|---|
| OAI-SearchBot | Surfaces sites in ChatGPT search features | Yes — this is the one |
| GPTBot | Crawls content that may train foundation models | Independent of search; see below |
| ChatGPT-User | Fetches a page because a user action asked for it | Not used for search inclusion |
| OAI-AdsBot | Validates landing pages submitted as ChatGPT ads | Only if you run ChatGPT ads |
Sites opted out of OAI-SearchBot will not be shown in ChatGPT search answers, though OpenAI notes they can still appear as navigational links.
Robots.txt is only half the check. OpenAI says that to be included you should allow OAI-SearchBot and make sure your site host and content delivery network allow traffic from its published IP addresses, which it lists at openai.com/searchbot.json. This is the quiet failure mode: a site can have a perfectly correct robots.txt and still be invisible because a managed host or CDN is dropping the requests before they arrive.
There's one more wrinkle if you do opt out. OpenAI says that if it learns a disallowed page's URL from a third-party search provider or by crawling other pages, and has signals the page is relevant, it may surface just the link and page title in ChatGPT Atlas. To prevent that, it points to the noindex meta tag — and notes the crawler has to be allowed to crawl the page in order to read the tag. Blocking everything is not the same as opting out cleanly.
Does blocking GPTBot hurt my ChatGPT visibility?
No. OpenAI states the settings are independent: you can allow OAI-SearchBot so ChatGPT can cite you while disallowing GPTBot to signal that your content shouldn't be used for training. If both are allowed, OpenAI may reuse a single crawl for both purposes to avoid crawling your site twice.
Two related details people get wrong. ChatGPT-User is user-initiated, so OpenAI says robots.txt rules may not apply to it — and it is explicitly not used to decide whether content appears in search. And after you change robots.txt, OpenAI says its systems take roughly 24 hours to adjust for search.
How do the engines pick winners among eligible pages?
This is where public documentation thins out. Neither company publishes ranking weights, and anyone who tells you otherwise is guessing. OpenAI says plainly that there is no way to guarantee top placement.
Here's what OpenAI does say about how ChatGPT search results are determined: language models evaluate content by meaning, intent, and relevance; automated systems decide which results to present, weighing user intent, relevance, and recency; cited sources appear inline; and the sidebar sources are ordered partly by the third-party provider's own ranking systems. OpenAI also says it may decline to surface sites containing illegal, harmful, or sensitive content.
Google's position is that generative AI features run on the same ranking and quality systems as classic Search — so the usual quality signals apply, with one emphasis worth noting. Google's guidance says content with a unique point of view tends to do better than commodity content. Its own comparison: a generic "7 Tips for First-Time Homebuyers" post is commodity; a post about why the writer waived the inspection, saved money, and what the sewer line inspection actually turned up is not. First-hand experience beats a restatement of what's already online.
What can I skip? (Google's own mythbusting)
Google published a list of things that don't help in its generative AI search guidance:
- llms.txt and other AI-specific files — Google Search ignores them. Keeping one is neither helpful nor harmful for Google. OpenAI has not documented whether its search crawler uses llms.txt in either direction, so treat any claim that it does — or doesn't — as unsourced.
- "Chunking" content into small pieces — not required. There's no ideal page length.
- Rewriting content specifically for AI — not required; the systems handle synonyms and intent.
- Chasing inauthentic mentions — spam systems and quality ranking both feed the AI features. Google clarified in May 2026 that its spam policies apply to generative AI responses too.
- Overfocusing on structured data — not required for generative AI search, and there's no special schema to add. It's still worth using where it drives rich results in regular Search, and any markup you use should match the visible text on the page.
- Spinning up pages to chase fan-out queries — Google warns that creating separate content for every variation of how people might search, including fan-out queries, primarily to manipulate rankings or AI responses, violates its scaled content abuse spam policy.
One more, related: FAQ rich results were retired from Google Search on May 7, 2026. FAQ reporting left Search Console in June 2026, with Search Console API support ending in August 2026, and Google removed the FAQ structured data documentation altogether in June 2026. FAQPage is still a valid schema.org type, and Google's position on markup it no longer uses is that leaving it in place causes no problems for Search — but it has no visible effect there either. Write FAQs because customers ask those questions, not for a SERP feature that no longer exists.
Google also warns against third-party tools claiming access to internal Google ranking or AI metrics. No such access exists.
What actually moves the needle for a local practice
Strip away the noise and the checklist is short:
- Confirm you're crawlable and indexed. Robots.txt, CDN rules, and hosting-level bot blocking all count. Verify your site in Search Console.
- Allow OAI-SearchBot if you want ChatGPT to cite you. Decide about GPTBot separately.
- Check your host and CDN aren't blocking OpenAI's IP ranges. A correct robots.txt doesn't help if your firewall drops the request. OpenAI publishes the ranges at
openai.com/searchbot.json. - Put important information in text. Not in an image, not behind JavaScript that fails to render.
- Keep your Google Business Profile current. Google names Business Profile and Merchant Center directly in its generative AI guidance — for a local practice, that's the record AI answers lean on for hours, location, and services.
- Write from experience. The specific case, the local ordinance, the thing you fix every week. That's what a model has no other source for.
- Measure it. Search Console's Performance report counts AI Overviews and AI Mode under the Web search type. There's also a separate Generative AI performance report, launched June 2026 and currently rolling out to a subset of sites, so you may not have it yet. ChatGPT referrals arrive tagged
utm_source=chatgpt.com, which OpenAI adds automatically. - Know about the Search generative AI control. It's in Search Console under Settings, it defaults to include, and it's currently rolling out to a subset of UK site owners. If your property has it, confirm it's set to include.
Nothing on that list is exotic. That's the actual finding: the gate is technical eligibility plus content worth quoting, and most local professional service sites fail the first test before they're ever judged on the second.
Common questions
Do I need llms.txt to appear in AI answers?
Not for Google — Google says its Search ignores the file entirely. OpenAI hasn't said either way. If another service you use reads it, keeping one does no harm.
Is GEO different from SEO?
Google's official answer is no. Its guidance addresses "AEO" and "GEO" directly and says that from Google Search's perspective, optimizing for generative AI search is optimizing for the search experience — still SEO. ChatGPT is the part that genuinely requires separate technical attention, because OAI-SearchBot is a different crawler with its own robots.txt permission and its own IP ranges to allow.
How fast do changes show up?
OpenAI says about 24 hours after a robots.txt update for its search systems to adjust. Google's crawling can take days to months depending on how often its systems decide a page needs refreshing; you can request a recrawl in Search Console. For sites that have the Search generative AI control, Google says changes generally take a few days.
Can I be in Google's AI answers but not ChatGPT's, or the other way around?
Yes. Neither company documents this explicitly, but they're separate systems with separate crawlers and separate permissions, so it follows. It's common for a site to be eligible in one and blocked in the other without anyone realizing.
Sources
- Google Search Central — AI features and your website
- Google Search Central — Optimizing your website for generative AI features on Google Search
- Search Console Help — Search generative AI control
- Google — New opportunities, control and insights for website owners (UK rollout)
- Google Search Central Blog — Introducing Search Generative AI performance reports in Search Console
- Google Search Central — Documentation updates: removing the FAQ rich result feature
- Search Console Help — Data anomalies in Search Console (confirms May 7, 2026)
- Google Search Central Blog — Changes to HowTo and FAQ rich results (Google on markup it no longer uses)
- OpenAI — Overview of OpenAI Crawlers
- OpenAI — Transparency & content moderation
- OpenAI Help Center — ChatGPT search
- OpenAI Help Center — Publishers and Developers FAQ
Claivo is a digital marketing business in Burlington, Iowa, working with local professional service practices across southeast Iowa — attorneys, dentists, physicians, chiropractors, and contractors.