Why ChatGPT, Claude or Perplexity can't see your website
When an assistant can't read your site, the cause is almost always one of five layers: robots.txt, your CDN, JavaScript, discovery or the assistant's own index. A test for each, in the vendors' own words.
When ChatGPT, Claude or Perplexity can't read your website, the cause is almost always one of five layers, and you can test each one yourself: robots.txt refuses the bot, your CDN or firewall blocks it, the text only appears after JavaScript runs, the assistant can't find your pages, or your pages aren't in that assistant's search index yet. Work through them in that order, because each layer hides the ones after it. The free AI-ready check runs most of these tests at once. They are the same basics that make up an AI-ready website.
First, name the failure
"Can't see my site" covers two different problems, and they have different causes.
| What you see | Who is fetching | Layers to check first |
|---|---|---|
| You paste your URL and the assistant says it can't open it | A user-triggered fetcher: ChatGPT-User, Claude-User or Perplexity-User | 2, 3, then 1 |
| The assistant opens it, but describes it wrongly or thinly | The same fetcher, reading what your server sends | 3 |
| Your site never appears in the assistant's search answers | A search crawler: OAI-SearchBot, Claude-SearchBot or PerplexityBot | 1, 2, 4, then 5 |
The training crawlers (GPTBot, ClaudeBot) don't decide either outcome. They affect what future models learn, not what an assistant can open or cite today.
Layer 1: robots.txt says no
Each vendor uses separate tokens for separate jobs, so a rule written for one bot doesn't cover the others. OpenAI says a site that blocks OAI-SearchBot "will not be shown in ChatGPT search answers", and that blocking GPTBot only opts out of training. Anthropic says blocking Claude-User "prevents our system from retrieving your content in response to a user query", and its web fetch tool lists robots.txt among the reasons a URL comes back as url_not_allowed. For user-triggered fetchers the rules differ by vendor: OpenAI says robots.txt "may not apply" to ChatGPT-User, and Perplexity says Perplexity-User "generally ignores" it.
Three mistakes cause most layer-1 failures:
- A forgotten
Disallow: /from a staging site, still in production. - A named group that refuses more than you meant. A crawler that has its own group ignores the
*group completely. - A robots.txt that returns a server error. RFC 9309 says a crawler that gets a 5xx for robots.txt "MUST assume complete disallow".
curl -s -o /dev/null -w "%{http_code}\n" https://example.com/robots.txt
curl -s https://example.com/robots.txt | grep -i -A3 "user-agent: \(\*\|OAI-SearchBot\|Claude\|Perplexity\)"
robots.txt for AI crawlers has the full table of tokens and three files to copy.
Layer 2: your CDN or firewall blocks it
A robots.txt that says yes doesn't help if the edge says no first. Bot protection, a web application firewall or a "checking your browser" page can answer a crawler with a 403, a 429 or a challenge page, while people see the site normally. Google gives the same advice for its own crawler: make sure Googlebot "isn't blocked" and that the page returns HTTP 200.
Cloudflare changed its defaults on 15 September 2026. Since then, new domains get one of two presets. Sites that earn money from ads get "Disallow AI Training", plus agents blocked on pages that show an ad, while sites without ads get "Allow" for both. Existing domains keep their settings, but check them anyway: the controls for Search, Training and Agent bots sit in your zone's Security settings, on every plan.
# Status code your edge returns to a search crawler's user agent
curl -s -o /dev/null -w "%{http_code}\n" \
-A "Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/131.0.0.0 Safari/537.36; compatible; OAI-SearchBot/1.4; +https://openai.com/searchbot" \
https://example.com/
# Does the body look like a challenge page?
curl -s -A "PerplexityBot/1.0" https://example.com/ | grep -i -o "<title>[^<]*"
A forged user agent only tests rules that match on the user agent. An edge that verifies crawlers by IP address treats your request differently from the real bot's, so a 200 here is a good sign, not proof. To let real crawlers through without opening the door to imitators, allow them by their published address ranges (OpenAI's openai.com/searchbot.json and openai.com/chatgpt-user.json, Perplexity's perplexity.com/perplexitybot.json), or by signature. OpenAI's allowlisting help page says ChatGPT's cloud browser signs its requests with HTTP Message Signatures (RFC 9421) and a Signature-Agent header of https://chatgpt.com. When we fetched https://chatgpt.com/.well-known/http-message-signatures-directory on 6 October 2026, it returned an Ed25519 public key that your edge can use to verify that a request really came from ChatGPT.
Layer 3: there's no text without JavaScript
If your page is an empty shell that a script fills in, most AI fetchers see the shell. Vercel and MERJ measured AI crawler traffic across Vercel's network and found that "none of the major AI crawlers currently render JavaScript", the exceptions being Gemini, which uses Google's infrastructure, and Applebot. ChatGPT and Claude did download JavaScript files (11.50% and 23.84% of their fetches), but didn't run them. That study dates from December 2024, so retest your own pages rather than relying on it. Anthropic's current documentation agrees for its own tool: web fetch "does not support websites dynamically rendered with JavaScript".
# Is your headline in the HTML the server sends?
curl -s https://example.com/ | grep -c "Your headline here"
A count of 0 means the text isn't in the server HTML. The fix is to render the page on the server or at build time, which any modern framework can do. Keep prices, opening hours and product names as HTML text, not drawn into a canvas or an image. How AI agents read websites shows the same page as a crawler, a browser agent and a screenshot each see it.
Layer 4: it can't find your pages
Search crawlers find pages through links and sitemaps, like any crawler. Make sure robots.txt names your sitemap, every important page is reachable through a plain <a href> link, and old URLs redirect to new ones. That last point matters more than it sounds: in the same Vercel study, 34.82% of ChatGPT's fetches and 34.16% of Claude's hit 404 pages, against 8.22% for Googlebot. AI crawlers ask for a lot of URLs that don't exist, and a redirect turns some of those misses into pages.
For agents that read your site on someone's behalf, a short llms.txt linked from your pages gives them a map. What is llms.txt explains the evidence on who reads it, and why linking it matters more than publishing it.
curl -s https://example.com/robots.txt | grep -i "^sitemap:"
curl -s -o /dev/null -w "%{http_code}\n" https://example.com/sitemap.xml
Layer 5: it isn't in the assistant's index yet
If layers 1 to 4 pass, the honest answer is that some of what happens next isn't documented. What is documented:
- Robots changes take about a day. OpenAI says about 24 hours, and Perplexity says up to 24 hours.
- Google's AI features use Google's index. A page appears in them if it is "indexed and eligible to be shown in Google Search with a snippet", with "no additional technical requirements". Search Console shows whether a page is indexed.
- Model knowledge isn't search. What a model learned in training and what it finds through live search are separate. A new site can be missing from the first and still be found by the second.
What isn't documented: how often each assistant's search index refreshes, how it ranks sources, and which outside indexes it draws on. Claims like "the training cycle runs 6 to 12 months behind" don't come from any vendor. If a page passes every layer and still doesn't show up, keep it crawlable and linked, and test again in a few weeks.
A ten-minute diagnosis
- Run the AI-ready check on your homepage. It makes a plain request the way a crawler does, reports bot walls and challenge pages, reads robots.txt for GPTBot, ClaudeBot, PerplexityBot and Google-Extended, and checks for readable text, a sitemap and llms.txt.
- Fix any robots.txt group that refuses the search crawler or user fetcher you want.
- Look at your CDN's bot settings, especially if the domain was added to Cloudflare after 15 September 2026.
- Check your headline with curl. If it's missing, render it on the server.
- Confirm your sitemap returns 200 and is named in robots.txt.
- Paste the URL into the assistant again the next day, once robots caches have refreshed.
A site that passes all five layers is open to every assistant that respects them. The custom domain setup matters here too: a proxy or firewall in front of your domain is layer 2, so check it whenever you change it.
Sources
Checked 6 October 2026. Standards and vendor pages change; the linked pages are the authority.
- OpenAI: overview of OpenAI crawlers
- Anthropic: does Anthropic crawl data from the web? (updated 7 April 2026)
- Anthropic: web fetch tool documentation
- Perplexity: Perplexity crawlers
- Google Search Central blog: top ways to ensure your content performs well in Google's AI experiences (John Mueller, 21 May 2025)
- Google Search Central: AI features and your website (updated 10 December 2025)
- Vercel and MERJ: the rise of the AI crawler (17 December 2024)
- Cloudflare: stay discoverable in search while disallowing AI training (15 September 2026)
- Cloudflare: your site, your rules, new AI traffic options (1 July 2026)
- OpenAI Help: ChatGPT cloud browser allowlisting
- RFC 9421: HTTP Message Signatures (IETF, February 2024)
- RFC 9309: Robots Exclusion Protocol (IETF, September 2022)
