Guides

Why ChatGPT, Claude or Perplexity can't see your website

When an assistant can't read your site, the cause is almost always one of five layers: robots.txt, your CDN, JavaScript, discovery or the assistant's own index. A test for each, in the vendors' own words.

By The Laarpi teamUpdated 6 min read

When ChatGPT, Claude or Perplexity can't read your website, the cause is almost always one of five layers, and you can test each one yourself: robots.txt refuses the bot, your CDN or firewall blocks it, the text only appears after JavaScript runs, the assistant can't find your pages, or your pages aren't in that assistant's search index yet. Work through them in that order, because each layer hides the ones after it. The free AI-ready check runs most of these tests at once. They are the same basics that make up an AI-ready website.

First, name the failure

"Can't see my site" covers two different problems, and they have different causes.

What you seeWho is fetchingLayers to check first
You paste your URL and the assistant says it can't open itA user-triggered fetcher: ChatGPT-User, Claude-User or Perplexity-User2, 3, then 1
The assistant opens it, but describes it wrongly or thinlyThe same fetcher, reading what your server sends3
Your site never appears in the assistant's search answersA search crawler: OAI-SearchBot, Claude-SearchBot or PerplexityBot1, 2, 4, then 5

The training crawlers (GPTBot, ClaudeBot) don't decide either outcome. They affect what future models learn, not what an assistant can open or cite today.

Layer 1: robots.txt says no

Each vendor uses separate tokens for separate jobs, so a rule written for one bot doesn't cover the others. OpenAI says a site that blocks OAI-SearchBot "will not be shown in ChatGPT search answers", and that blocking GPTBot only opts out of training. Anthropic says blocking Claude-User "prevents our system from retrieving your content in response to a user query", and its web fetch tool lists robots.txt among the reasons a URL comes back as url_not_allowed. For user-triggered fetchers the rules differ by vendor: OpenAI says robots.txt "may not apply" to ChatGPT-User, and Perplexity says Perplexity-User "generally ignores" it.

Three mistakes cause most layer-1 failures:

  • A forgotten Disallow: / from a staging site, still in production.
  • A named group that refuses more than you meant. A crawler that has its own group ignores the * group completely.
  • A robots.txt that returns a server error. RFC 9309 says a crawler that gets a 5xx for robots.txt "MUST assume complete disallow".
curl -s -o /dev/null -w "%{http_code}\n" https://example.com/robots.txt
curl -s https://example.com/robots.txt | grep -i -A3 "user-agent: \(\*\|OAI-SearchBot\|Claude\|Perplexity\)"

robots.txt for AI crawlers has the full table of tokens and three files to copy.

Layer 2: your CDN or firewall blocks it

A robots.txt that says yes doesn't help if the edge says no first. Bot protection, a web application firewall or a "checking your browser" page can answer a crawler with a 403, a 429 or a challenge page, while people see the site normally. Google gives the same advice for its own crawler: make sure Googlebot "isn't blocked" and that the page returns HTTP 200.

Cloudflare changed its defaults on 15 September 2026. Since then, new domains get one of two presets. Sites that earn money from ads get "Disallow AI Training", plus agents blocked on pages that show an ad, while sites without ads get "Allow" for both. Existing domains keep their settings, but check them anyway: the controls for Search, Training and Agent bots sit in your zone's Security settings, on every plan.

# Status code your edge returns to a search crawler's user agent
curl -s -o /dev/null -w "%{http_code}\n" \
  -A "Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/131.0.0.0 Safari/537.36; compatible; OAI-SearchBot/1.4; +https://openai.com/searchbot" \
  https://example.com/

# Does the body look like a challenge page?
curl -s -A "PerplexityBot/1.0" https://example.com/ | grep -i -o "<title>[^<]*"

A forged user agent only tests rules that match on the user agent. An edge that verifies crawlers by IP address treats your request differently from the real bot's, so a 200 here is a good sign, not proof. To let real crawlers through without opening the door to imitators, allow them by their published address ranges (OpenAI's openai.com/searchbot.json and openai.com/chatgpt-user.json, Perplexity's perplexity.com/perplexitybot.json), or by signature. OpenAI's allowlisting help page says ChatGPT's cloud browser signs its requests with HTTP Message Signatures (RFC 9421) and a Signature-Agent header of https://chatgpt.com. When we fetched https://chatgpt.com/.well-known/http-message-signatures-directory on 6 October 2026, it returned an Ed25519 public key that your edge can use to verify that a request really came from ChatGPT.

Layer 3: there's no text without JavaScript

If your page is an empty shell that a script fills in, most AI fetchers see the shell. Vercel and MERJ measured AI crawler traffic across Vercel's network and found that "none of the major AI crawlers currently render JavaScript", the exceptions being Gemini, which uses Google's infrastructure, and Applebot. ChatGPT and Claude did download JavaScript files (11.50% and 23.84% of their fetches), but didn't run them. That study dates from December 2024, so retest your own pages rather than relying on it. Anthropic's current documentation agrees for its own tool: web fetch "does not support websites dynamically rendered with JavaScript".

# Is your headline in the HTML the server sends?
curl -s https://example.com/ | grep -c "Your headline here"

A count of 0 means the text isn't in the server HTML. The fix is to render the page on the server or at build time, which any modern framework can do. Keep prices, opening hours and product names as HTML text, not drawn into a canvas or an image. How AI agents read websites shows the same page as a crawler, a browser agent and a screenshot each see it.

Layer 4: it can't find your pages

Search crawlers find pages through links and sitemaps, like any crawler. Make sure robots.txt names your sitemap, every important page is reachable through a plain <a href> link, and old URLs redirect to new ones. That last point matters more than it sounds: in the same Vercel study, 34.82% of ChatGPT's fetches and 34.16% of Claude's hit 404 pages, against 8.22% for Googlebot. AI crawlers ask for a lot of URLs that don't exist, and a redirect turns some of those misses into pages.

For agents that read your site on someone's behalf, a short llms.txt linked from your pages gives them a map. What is llms.txt explains the evidence on who reads it, and why linking it matters more than publishing it.

curl -s https://example.com/robots.txt | grep -i "^sitemap:"
curl -s -o /dev/null -w "%{http_code}\n" https://example.com/sitemap.xml

Layer 5: it isn't in the assistant's index yet

If layers 1 to 4 pass, the honest answer is that some of what happens next isn't documented. What is documented:

  • Robots changes take about a day. OpenAI says about 24 hours, and Perplexity says up to 24 hours.
  • Google's AI features use Google's index. A page appears in them if it is "indexed and eligible to be shown in Google Search with a snippet", with "no additional technical requirements". Search Console shows whether a page is indexed.
  • Model knowledge isn't search. What a model learned in training and what it finds through live search are separate. A new site can be missing from the first and still be found by the second.

What isn't documented: how often each assistant's search index refreshes, how it ranks sources, and which outside indexes it draws on. Claims like "the training cycle runs 6 to 12 months behind" don't come from any vendor. If a page passes every layer and still doesn't show up, keep it crawlable and linked, and test again in a few weeks.

A ten-minute diagnosis

  1. Run the AI-ready check on your homepage. It makes a plain request the way a crawler does, reports bot walls and challenge pages, reads robots.txt for GPTBot, ClaudeBot, PerplexityBot and Google-Extended, and checks for readable text, a sitemap and llms.txt.
  2. Fix any robots.txt group that refuses the search crawler or user fetcher you want.
  3. Look at your CDN's bot settings, especially if the domain was added to Cloudflare after 15 September 2026.
  4. Check your headline with curl. If it's missing, render it on the server.
  5. Confirm your sitemap returns 200 and is named in robots.txt.
  6. Paste the URL into the assistant again the next day, once robots caches have refreshed.

A site that passes all five layers is open to every assistant that respects them. The custom domain setup matters here too: a proxy or firewall in front of your domain is layer 2, so check it whenever you change it.

Sources

Checked 6 October 2026. Standards and vendor pages change; the linked pages are the authority.

  1. OpenAI: overview of OpenAI crawlers
  2. Anthropic: does Anthropic crawl data from the web? (updated 7 April 2026)
  3. Anthropic: web fetch tool documentation
  4. Perplexity: Perplexity crawlers
  5. Google Search Central blog: top ways to ensure your content performs well in Google's AI experiences (John Mueller, 21 May 2025)
  6. Google Search Central: AI features and your website (updated 10 December 2025)
  7. Vercel and MERJ: the rise of the AI crawler (17 December 2024)
  8. Cloudflare: stay discoverable in search while disallowing AI training (15 September 2026)
  9. Cloudflare: your site, your rules, new AI traffic options (1 July 2026)
  10. OpenAI Help: ChatGPT cloud browser allowlisting
  11. RFC 9421: HTTP Message Signatures (IETF, February 2024)
  12. RFC 9309: Robots Exclusion Protocol (IETF, September 2022)
Questions

Fair questions

Why does ChatGPT say it can't access my website?

When you paste a link, ChatGPT fetches it on your behalf. The usual causes are a firewall or challenge page that answers bots instead of the page, a server error, or a page whose text only appears after JavaScript runs. OpenAI says robots.txt may not apply to these user-triggered visits, so check your CDN and your server HTML first.

I blocked GPTBot. Is that why ChatGPT search doesn't show my site?

No. OpenAI documents GPTBot as its training crawler. ChatGPT's search features use OAI-SearchBot, and a site that blocks OAI-SearchBot 'will not be shown in ChatGPT search answers'. Check both groups in your robots.txt.

Does ChatGPT search use Google's or Bing's index?

OpenAI's crawler documentation describes OAI-SearchBot as the bot that surfaces sites in ChatGPT's search features and doesn't say which other indexes feed its answers. Treat confident claims either way with care, and make sure OAI-SearchBot can reach your pages.

How long after I fix it will assistants see my site?

For robots.txt changes, OpenAI says about 24 hours and Perplexity says up to 24 hours. No vendor documents how long it takes for a new or changed page to appear in its search answers, so anyone who quotes a fixed number is guessing.

My robots.txt allows Claude, but Claude still can't fetch my page. Why?

Anthropic's web fetch tool lists several reasons a URL is refused or fails: robots.txt, private addresses, HTTP errors and content types other than text, HTML and PDF. It also doesn't support pages rendered with JavaScript. A challenge page from your CDN shows up as an error too.

Is there one test that checks all of this?

The free AI-ready check at /tools/ai-ready-check makes a plain request like a crawler, reads your robots.txt for GPTBot, ClaudeBot, PerplexityBot and Google-Extended, checks whether the main text is there without JavaScript, and looks for a sitemap and llms.txt. It reports each result with a fix.

Related

Start building

Describe the site in a sentence. It asks what matters, then designs and builds it from scratch.

One sentence is enough.