AI crawler access checker
Enter a site to see which AI crawlers its robots.txt allows: GPTBot, OAI-SearchBot, ClaudeBot, PerplexityBot, Googlebot, Google-Extended and more. A site that blocks search crawlers cannot be cited by them.
How to read the result
| Result | What it means |
|---|---|
| Allowed | The crawler may read the site. If some paths are closed to it, such as an admin area or redirect links, the result says how many. That is normal. |
| Partly | The site is closed to this crawler, with some paths left open. |
| Blocked | A rule closes the whole site to this crawler. |
| Rule used | Whether the result comes from a rule written for that crawler by name, or from the general rule for all crawlers. |
Three kinds of AI crawler
- Search crawlers build the index an assistant searches when it answers. Block these and the site cannot be found by that assistant.
- Training crawlers collect pages to train future models. Blocking them does not affect live search.
- User-requested fetchers open a page only when a person asks the assistant to read it.
Which index each assistant searches is explained in AI engine indexes, and how they pick what to cite in how AI search chooses sources.
Why this matters for off-page work
Mentions and links on other sites help AI visibility only if the assistants can read those sites, and yours. Before paying for a placement meant to be cited by AI search, check that the publisher is open to the search crawlers. The measuring side is covered in how to measure AI visibility.
What AI crawlers are
An AI crawler is a program that an AI company runs to read web pages. Each one announces itself with a name, and a site’s robots.txt file can give that name its own rules. The names matter because one company often runs several crawlers for different jobs, and blocking the wrong one has the opposite effect to the one you wanted.
The jobs fall into the three kinds described above: building a search index, collecting training data, and fetching a page because a user asked.
The main crawler names by company
Everything in this table comes from each company’s own documentation, read on 4 October 2026.
| Name | Company | Kind | What the company says it does |
|---|---|---|---|
| OAI-SearchBot | OpenAI | Search | Used “to surface websites in search results in ChatGPT’s search features” |
| GPTBot | OpenAI | Training | Crawls “content that may be used in training our generative AI foundation models” |
| ChatGPT-User | OpenAI | User-requested | Used “for certain user actions in ChatGPT and Custom GPTs” |
| Claude-SearchBot | Anthropic | Search | Crawls “to improve search result quality for users” |
| ClaudeBot | Anthropic | Training | Collects web content “that could potentially contribute to their training” |
| Claude-User | Anthropic | User-requested | May access websites when people ask Claude questions |
| PerplexityBot | Perplexity | Search | Designed “to surface and link websites in search results on Perplexity”. Not used for foundation model training |
| Perplexity-User | Perplexity | User-requested | Supports “user actions within Perplexity” |
| Googlebot | Search | The crawler for Google Search and all its features | |
| Google-Extended | Training control | A robots.txt name only. Controls use for Gemini training and grounding | |
| Applebot-Extended | Apple | Training control | A robots.txt name only. Apple says it “does not crawl webpages” |
| CCBot | Common Crawl | Open archive | Builds “an open repository of web crawl data” |
OpenAI. Its documentation says sites opted out of OAI-SearchBot “will not be shown in ChatGPT search answers, though can still appear as navigational links”. For ChatGPT-User it says that because the actions are started by a user, “robots.txt rules may not apply”.
Anthropic. Its help page says its bots honour “industry standard directives in robots.txt”, and that blocking Claude-SearchBot or Claude-User “may reduce your site’s visibility” in search results and in answers to user queries.
Perplexity. Its documentation says Perplexity-User “generally ignores robots.txt rules”, because a user starts the request.
Google. Google’s crawler documentation says Google-Extended “does not impact a site’s inclusion in Google Search nor is it used as a ranking signal in Google Search”. For AI Overviews and AI Mode, Google’s page on AI features says a page must be indexed and eligible to show with a snippet, with no additional technical requirements.
Apple. Apple’s support page says pages that disallow Applebot-Extended “can still be included in search results”. One quirk: if a robots.txt file has no rules for Applebot but has rules for Googlebot, Apple says Applebot follows the Googlebot rules.
Microsoft. Copilot draws on Bing’s index, as AI engine indexes explains, which makes Bingbot the name to allow. A Bing Webmaster blog post from 22 September 2023 describes page-level tags for its chat answers: content tagged NOARCHIVE “will not be included in Bing Chat answers”.
The tool also checks crawlers from Meta, Amazon, ByteDance, DuckDuckGo and Mistral, which this page does not describe.
How robots.txt rules are matched
The tool applies the matching rules in Google’s robots.txt specification, last updated on 31 August 2026 when we read it.
1. A crawler follows one group: the most specific one that names it. If no group names it, it follows the general group, User-agent: *. Google’s specification says “user agent specific groups and global groups (*) are not combined”. The name is not case sensitive. Paths are.
2. Within the group, the longest matching path wins. In Google’s words, crawlers “use the most specific rule based on the length of the rule path”.
3. On a tie, the less restrictive rule wins. “In case of conflicting rules, including those with wildcards, Google uses the least restrictive rule.”
The same specification says a robots.txt that returns a 4xx status, such as 404, is treated as if no file exists, so nothing is blocked. A 5xx status stops crawling while Google retries, and the tool shows “Unknown”. Rules apply only to the host they sit on, so check a subdomain separately.
The verdict in the tool is about the home page. “Blocked” means a rule closes the whole site from the root. “Partly” means the root is closed but some paths are opened with Allow. A site that closes one important folder still shows “Allowed”, with the number of closed paths.
How to allow or block a crawler
Block training, keep search
This file refuses the training crawlers and controls by name. Every other crawler, including OAI-SearchBot, Claude-SearchBot, PerplexityBot, Googlebot and Bingbot, falls to the general group and may read everything except the admin area.
User-agent: GPTBot
User-agent: ClaudeBot
User-agent: CCBot
User-agent: Google-Extended
User-agent: Applebot-Extended
Disallow: /
User-agent: *
Disallow: /admin/
The trap: a named group replaces the general one
User-agent: *
Disallow: /checkout/
Disallow: /internal-search/
User-agent: OAI-SearchBot
Disallow: /drafts/
Here OAI-SearchBot follows only its own group. The two general rules do not apply to it, so /checkout/ and /internal-search/ are open to it. If you write a group for a crawler, repeat in it every rule you want that crawler to keep.
Close the site but open one section
User-agent: PerplexityBot
Disallow: /
Allow: /blog/
For /blog/a-post/, the Allow rule is the longer match, so the page is open. Everything else is closed. The tool reports this as “Partly”.
The accidental block
User-agent: *
Disallow: /
This closes the site to every crawler that has no group of its own, search crawlers included. In the tool it shows as most crawlers blocked, with “the general rule for all crawlers” as the rule used.
Changes are not instant. OpenAI’s documentation says it can take about 24 hours for its systems to adjust after a robots.txt update, and Perplexity’s says up to 24 hours.
What blocking each kind does to visibility in AI answers
| You block | What it does | What it does not do |
|---|---|---|
| A search crawler | Removes the site from that assistant’s search answers, or makes it less visible there | Remove you from other assistants, or from models already trained |
| A training crawler or control | Tells the company not to use your pages for future training | Change whether you are found by live search |
| A user-requested fetcher | May stop the assistant opening your page when a user asks | Reliably stop it. OpenAI and Perplexity both say robots.txt may not apply to these |
| Googlebot or Bingbot | Removes you from the search index and from the AI features built on it | Anything selective. robots.txt has no AI-only setting for these two names |
Two limits apply. Google’s introduction to robots.txt says the file “cannot enforce crawler behavior”: it is up to the crawler to obey. And robots.txt controls reading, not mentioning. An assistant can still name your brand because other sites wrote about you, as brand mentions and citations explains.
To limit what Google shows from a page without blocking the crawler, Google’s AI features page points to the nosnippet, data-nosnippet, max-snippet and noindex controls.
Does llms.txt help?
llms.txt is a proposal, published at llmstxt.org in September 2024, for a Markdown file at the root of a site that tells AI agents where its key content is. No search or AI company has confirmed using it to choose which pages to cite. Google’s AI features page says: “You don’t need to create new machine readable files, AI text files, or markup to appear in these features.” SE Ranking’s November 2025 study of 300,000 domains found no relationship between having the file and how often a domain was cited.
The tool tells you whether the file exists and nothing more. What is GEO, AEO and LLMO sets this beside the other claims made about AI visibility.
Firewalls and bot protection that robots.txt does not show
A robots.txt file is a notice. A firewall, a CDN or a bot protection service is a locked door. A bot-blocking setting at the CDN can stop a search crawler that robots.txt allows. The reverse also happens: a crawler blocked only at the firewall will look “Allowed” in this tool.
The tool cannot test this. It gives one hint: if the home page answers with an error status, the result says the site may be blocking automated visitors.
What the companies say about firewalls:
- Perplexity advises site owners to allow its crawlers in the firewall by matching the user agent together with its published IP ranges.
- Anthropic warns that blocking its IP addresses “may not work correctly or persistently guarantee an opt-out”, because it stops the bot reading your robots.txt.
- OpenAI, Perplexity, Anthropic and Common Crawl each publish their IP ranges as a file.
To check your own site, look in your firewall or CDN settings for any AI bot rule, then search your server logs for the crawler names.
A checklist for a site that wants to be cited by AI search
- robots.txt returns status 200 or 404, not a server error.
- Googlebot, Bingbot, OAI-SearchBot, PerplexityBot and Claude-SearchBot all show “Allowed”.
- No leftover
Disallow: /sits in the general group. - The firewall or CDN is not blocking the same crawlers.
- The home page and key pages carry no
noindexornosnippetinstruction you did not intend. - Key pages are indexed in Bing as well as Google.
- After a change, wait a day and run the check again.
What gets a readable page cited is covered in link building for AI visibility and query fan-out, and what to measure afterwards in how to measure AI visibility and AI referral traffic.
The same check works on publishers: run the domain through this tool, then the article through the publisher page checker.
Common questions
What does the AI crawler access checker test?
It reads the robots.txt file of the site you enter and works out, for each AI crawler, whether the file lets it read the site or blocks it from the whole site, and how many paths are closed to it. It also notes whether the home page carries a robots instruction and whether an llms.txt file exists.
If a crawler is blocked, will the site disappear from AI answers?
Blocking a search crawler such as OAI-SearchBot, PerplexityBot or Googlebot removes the site from the index that the assistant searches, so it is unlikely to be cited from a live search. Blocking a training crawler such as GPTBot or ClaudeBot affects future model training, not live search. The two are separate decisions.
Does robots.txt actually stop a crawler?
It is an instruction, not a lock. The companies listed here state that their crawlers follow it. A site can also block crawlers at its firewall, which this tool cannot see.
Do I need an llms.txt file?
No search or AI company has confirmed that it uses llms.txt to choose which pages to cite. The tool reports whether the file exists, for information only.
What is Google-Extended?
It is a robots.txt name that controls whether Google may use a site for Gemini training and grounding. It is not a separate crawler, and blocking it does not remove a site from Google Search or from AI Overviews, which rely on Googlebot.


