Foundlab.

/ Free tool

Can AI crawlers read your page?

Enter a page and see which search and AI crawlers your robots.txt lets in, whether the page can be indexed, what structured data it carries and how much of it is readable without JavaScript. No email, nothing stored.

Any public page. A homepage is a good start. Nothing you enter is stored.

What it checks

Six things, for one page.

  1. AI and search crawlers.

    Reads your robots.txt and works out, for this exact page, whether each of 9 named crawlers is allowed. Search and answer crawlers are kept apart from training crawlers, because blocking training is a choice, not a fault.

  2. Whether it can be indexed.

    Looks for a noindex instruction in the page and in its headers, and checks which address the page names as its main version.

  3. Whether the content can be read.

    Counts the words in the page as it is first delivered, before any JavaScript runs. Most AI crawlers do not run JavaScript, so a page that builds itself in the browser can look empty to them.

  4. Structured data.

    Lists the JSON-LD types on the page, such as Organization, LocalBusiness or FAQPage, and flags any block that is broken and would be ignored.

  5. What search results show.

    The title and meta description with their lengths, and whether links shared in messages get a proper preview.

  6. Site files.

    Whether robots.txt can be read, whether a sitemap is published and declared, and whether an llms.txt file exists.

What it cannot tell you

Four limits worth knowing.

  • It checks one page. Other pages on the site can have different rules and different problems.
  • Allowed in robots.txt is not the same as reachable. A firewall or bot protection can still turn a crawler away, and robots.txt cannot override that.
  • It confirms structured data is present and readable, not that it qualifies for rich results. Google's Rich Results Test does that.
  • The request comes from Foundlab's server, which is not in Australia. A site that blocks or treats overseas visitors differently may show a different result here.

To see whether assistants actually name you, work through the findability checklist. For the whole site checked, fixed in order and written up, see the Findability Audit.

FAQ

Questions about the checker.

Should I block AI crawlers?

It depends which ones. Crawlers such as OAI-SearchBot, Claude-SearchBot and PerplexityBot fetch pages for search results and answers, so blocking them keeps you out of those answers. Training crawlers such as GPTBot, ClaudeBot and CCBot collect content for model training. Blocking those is a reasonable choice and does not stop you appearing in search.

Do I need an llms.txt file?

No. It is an optional summary file that some AI tools read. It is not a standard, and Google says no special AI file is needed to appear in its AI features. It can be useful, but its absence is not a problem.

My page shows only a few words. Is that bad?

If the page is meant to have content, yes. It usually means the content is built in the browser with JavaScript. Google can render JavaScript, but most AI crawlers do not, so they see a nearly empty page. Rendering the main content on the server fixes it.

Why does my site say it is blocking automated visitors?

The site answered this checker with an error or a challenge page, which usually means a firewall or bot protection. When that happens the page-level checks are skipped rather than reported on a block page. Ask your host whether verified search and AI crawlers are allowed through.

Do you store the pages I check?

No. The address, the result and your IP address are not saved anywhere. Analytics records only that a check ran and how many issues it found.

Not sure where to start?

Book a fifteen minute call. If it is not the right thing, you hear that on the call, before you pay.