Skip to content

AI readiness checks explained

What each readiness check tests, which AI crawlers Signal evaluates in robots.txt, and how scores are calculated.

Updated

The readiness audit fetches a page from Signal's servers and evaluates whether answer engines can access, understand, and attribute it. It runs on demand in the app and in the free AI readiness checker.

Crawler access

Signal parses robots.txt according to RFC 9309 and evaluates each crawler against the audited path. Crawlers are split into two groups because they have different consequences.

CrawlerOperatorPurposeGroup
OAI-SearchBotOpenAIChatGPT search results and citationsRetrieval
ChatGPT-UserOpenAIPages fetched when a user asksRetrieval
PerplexityBotPerplexityPerplexity answer indexRetrieval
Claude-SearchBotAnthropicClaude search resultsRetrieval
GooglebotGoogleGoogle Search, including AI OverviewsRetrieval
BingbotMicrosoftBing index, which grounds CopilotRetrieval
GPTBotOpenAIModel trainingTraining
ClaudeBotAnthropicModel trainingTraining
Google-ExtendedGoogleGemini training opt-out tokenTraining
Applebot-ExtendedAppleApple Intelligence training opt-out tokenTraining
CCBotCommon CrawlOpen web corpusTraining

Blocking a retrieval crawler removes the page from that engine's live answers and is reported as a failure. Blocking a training crawler is a legitimate policy choice and is reported for information only.

Discovery, content, and schema

The audit also checks for a sitemap declared in robots.txt, an optional llms.txt file, noindex directives, a canonical URL, a descriptive title and meta description, a single H1, the amount of readable text present without JavaScript, Open Graph metadata, valid JSON-LD, and an Organization, Product, or SoftwareApplication entity.

Scoring

Each check is pass (1), warning (0.5), or fail (0). Informational checks are not scored. The score is the average of scored checks, from 0 to 100. A high score means the page is technically accessible and well described; it does not guarantee that any engine will index, cite, or recommend it.

Network and privacy

The audit fetches only public http and https URLs on standard ports, refuses private and reserved network addresses, follows at most four redirects, and reads at most 1.5 MB of HTML. Requests identify as ManifestSignalBot.