Free tool
AI crawler checker
Can ChatGPT, Claude, Perplexity and Gemini actually read your site? Enter a URL to check your robots.txt rules for 13 AI crawlers, your llms.txt, and whether the page's content is in the HTML they receive. It's free and needs no sign-up.
robots.txt, per AI crawler
GPTBot, OAI-SearchBot, ChatGPT-User, ClaudeBot, Claude-SearchBot, Claude-User, PerplexityBot, Perplexity-User, Google-Extended, Applebot-Extended, meta-externalagent, CCBot and Bytespider. We evaluate them against your page's path, with Allow/Disallow precedence and wildcards.
llms.txt
Whether it exists, whether it's your app shell served for a missing file, and whether it follows the llmstxt.org format. We also check for llms-full.txt.
Server-rendered content
We fetch the page as GPTBot and measure the text in the raw HTML. An empty app shell fails. We also compare against what a browser gets, and flag a firewall that blocks AI crawlers, missing title, description and H1 tags, and missing JSON-LD.
Questions
- What does the AI crawler checker test?
- Three things an AI engine needs before it can cite a page. First, whether your robots.txt lets the AI crawlers in (OpenAI, Anthropic, Perplexity, Google, Apple, Meta, Common Crawl and ByteDance). Second, whether you publish a valid llms.txt. Third, whether the page's content is in the raw HTML a crawler receives, or only appears after JavaScript runs.
- Should I block GPTBot and the other training crawlers?
- It's your call. Training crawlers such as GPTBot, ClaudeBot and Google-Extended collect content for future models. Search and browsing crawlers such as OAI-SearchBot, ChatGPT-User, Claude-SearchBot and PerplexityBot fetch pages for live answers. Blocking a training crawler keeps you out of future training data but doesn't stop live answers citing you. Blocking a search crawler does, so the checker fails that and only warns about the other.
- What is Google-Extended?
- Google-Extended is a robots.txt token, not a separate crawler. It controls whether Google may use your content for its Gemini models. Blocking it doesn't remove you from Google Search or AI Overviews.
- Why does server rendering matter for AI search?
- Most AI crawlers download the HTML and don't run JavaScript. If your page is a client-rendered app shell, they see an empty page with nothing to quote, however good the content is in a browser.
- What is llms.txt?
- llms.txt is a markdown file at your domain root that gives AI agents a short summary of your site and links to its key pages (see llmstxt.org). The checker confirms it exists, that it isn't your HTML app shell served for a missing file, and that it has a title, a summary and a link list.
- How often are results updated?
- Results are cached for 24 hours per URL, and the badge shows the same cached result. If you've just fixed something, the badge catches up within a day.
Crawlable is step one. Being cited is step two.
Terradium tracks how often ChatGPT, Perplexity, Google AI Overviews and Gemini mention and cite you for the questions your buyers ask, and writes the content that closes the gaps.
Start free for 7 days