AI readiness checks explained
What each readiness check tests, which AI crawlers Signal evaluates in robots.txt, and how scores are calculated.
Updated
The readiness audit fetches a page from Signal's servers and evaluates whether answer engines can access, understand, and attribute it. It runs on demand in the app and in the free AI readiness checker.
Crawler access
Signal parses robots.txt according to RFC 9309 and evaluates each crawler against the audited path. Crawlers are split into two groups because they have different consequences.
| Crawler | Operator | Purpose | Group |
|---|---|---|---|
| OAI-SearchBot | OpenAI | ChatGPT search results and citations | Retrieval |
| ChatGPT-User | OpenAI | Pages fetched when a user asks | Retrieval |
| PerplexityBot | Perplexity | Perplexity answer index | Retrieval |
| Claude-SearchBot | Anthropic | Claude search results | Retrieval |
| Googlebot | Google Search, including AI Overviews | Retrieval | |
| Bingbot | Microsoft | Bing index, which grounds Copilot | Retrieval |
| GPTBot | OpenAI | Model training | Training |
| ClaudeBot | Anthropic | Model training | Training |
| Google-Extended | Gemini training opt-out token | Training | |
| Applebot-Extended | Apple | Apple Intelligence training opt-out token | Training |
| CCBot | Common Crawl | Open web corpus | Training |
Blocking a retrieval crawler removes the page from that engine's live answers and is reported as a failure. Blocking a training crawler is a legitimate policy choice and is reported for information only.
Discovery, content, and schema
The audit also checks for a sitemap declared in robots.txt, an optional llms.txt file, noindex directives, a canonical URL, a descriptive title and meta description, a single H1, the amount of readable text present without JavaScript, Open Graph metadata, valid JSON-LD, and an Organization, Product, or SoftwareApplication entity.
Scoring
Each check is pass (1), warning (0.5), or fail (0). Informational checks are not scored. The score is the average of scored checks, from 0 to 100. A high score means the page is technically accessible and well described; it does not guarantee that any engine will index, cite, or recommend it.
Network and privacy
The audit fetches only public http and https URLs on standard ports, refuses private and reserved network addresses, follows at most four redirects, and reads at most 1.5 MB of HTML. Requests identify as ManifestSignalBot.