robots.txt · AI crawlers

Audit robots.txt GPTBot and ClaudeBot disallow rules across domains

robots.txt Parser & AI Crawler Block Checker fetches public robots.txt files and parses crawl rules for AI crawlers and LLM-training bots (examples: GPTBot, ClaudeBot, Google-Extended). It extracts per-user-agent allow/disallow rules, restricted paths, and crawl-delay directives, then reports per-bot dispositions (blocked, partially blocked, allowed). This is a robots.txt policy audit — not HTML head-tag extraction, and not sitemap.xml URL discovery.

$11.00 per 1,000 robots.txt AI policy rows ($0.011 per delivered policy or missing-file row)

Open robots.txt Parser & AI Crawler Block Checker on Apify

robots.txt AI crawl policies, not page head tags and not sitemap XML

Use this page when the job is bulk parsing of robots.txt to see whether named AI crawlers are allowed, partially blocked, or disallowed. Use Meta Tag & OpenGraph Scraper when the job is title, canonical, Open Graph, Twitter Card, hreflang, JSON-LD, and the HTML robots meta tag on known page URLs. Use Sitemap Scraper & Analyzer when the job is nested sitemap.xml URL inventories.

This Actor Meta Tag & OpenGraph Scraper Sitemap Scraper & Analyzer
Intent Audit robots.txt crawl policies for GPTBot, ClaudeBot, Google-Extended, and other AI crawlers Technical SEO audit of HTML head metadata, social previews, JSON-LD, and the robots meta tag Discover URLs from nested XML sitemaps; optional HEAD 200 check
Primary input Required domains (maximum 500) Known page URLs in urls (maximum 500) Required sitemapUrls (sitemap.xml or a domain that auto-discovers /sitemap.xml)
What it reads Public HTTP(S) robots.txt bodies and response headers HTML head and embedded JSON-LD from the initial HTML (no JavaScript render) Sitemap XML and sitemap index files; optional HEAD on discovered URLs
Primary output Per-domain policy snapshots: aiPolicies, dispositions, rowType, source robots.txt URL, change diffs Title, description, canonical, robots meta, Open Graph, Twitter Card, hreflang, JSON-LD, issue flags URL inventory, directory structure, lastmod/changefreq analysis, optional HTTP 200
GPTBot / ClaudeBot disallow Yes — parsed from robots.txt user-agent directives No — HTML robots meta is not a robots.txt AI-bot policy No — sitemap XML does not encode GPTBot/ClaudeBot allow/disallow

The live Store description is: Audit robots.txt files across thousands of websites to detect specific crawl policies, disallowed paths, and user-agents for GPTBot and ClaudeBot. The live schema description is: Analyze robots.txt files for AI crawler blocking policies. Track which sites block GPTBot, ClaudeBot, Google-Extended, and 13 more. Published README copy also names allow/disallow rules, restricted directory paths, crawl-delay directives, policy snapshots, change diffs, severity-scored findings with remediation guidance, and raw header/body evidence.

Source requests allow only public HTTP(S) hosts. Local, private, link-local, metadata, credential-bearing, and non-HTTP(S) URLs are rejected. Redirects are rechecked and limited to 3 hops; responses are capped at 1 MiB and requests time out after 10 seconds. The published README says this is passive scanning of public signals, not a penetration test: no exploitation, fuzzing, or auth bypass. Only scan sites you have authorization for.

Use cases

Published README Use Cases table: Developers automate recurring fetches; data teams pipe structured output into analytics warehouses; ops teams monitor changes via webhook alerts; product managers track competitor/market signals. Published Key Features: compliance-first reports mapping findings to WCAG, GDPR, and SOC2; non-invasive scanning of observable public signals; severity-scored output with remediation guidance; delta-alerting of new findings since the last run via webhook; evidence export of raw headers/responses.

Store Quickstart starts with store-input.example.json at demoMode=true. Published templates are Demo Quickstart (trial run), Production Monitor (recurring dataset snapshots), and Webhook Alert (policy-change notifications). README tips: schedule weekly runs against production domains to catch config drift; use webhook delivery to pipe findings into a SIEM (Splunk, Datadog, Elastic); for CI, block releases on critical severity findings using exit codes. Combine with SSL/TLS Certificate Scraper when you also need certificate expiry coverage.

How is robots.txt Parser & AI Crawler Block Checker different from Meta Tag & OpenGraph Scraper and Sitemap Scraper & Analyzer?

This Actor fetches and parses robots.txt files. It extracts per-user-agent allow/disallow rules, restricted paths, and crawl-delay directives, then reports per-bot dispositions (blocked, partially blocked, allowed) for AI crawlers such as GPTBot, ClaudeBot, and Google-Extended. Meta Tag & OpenGraph Scraper audits HTML head tags on known page URLs: title, canonical, the HTML robots meta tag, Open Graph, Twitter Card, hreflang, and JSON-LD. That robots field is page-head metadata, not a robots.txt file. Sitemap Scraper & Analyzer reads sitemap.xml and nested sitemap indexes to discover URL inventories; it does not parse robots.txt AI crawler policies. Use this Actor to audit crawl policies in robots.txt, including GPTBot/ClaudeBot disallow. Use Meta Tag & OpenGraph Scraper for page head tags. Use Sitemap Scraper & Analyzer for XML sitemap URL discovery.

What input is required?

Open the Actor on the Apify Store and supply domains (required, maximum 500). Live input schema fields:

Field Type Default Notes
domains string[] required Domains to analyze robots.txt for AI crawler policies (max 500). Schema prefill is google.com, github.com, nytimes.com, openai.com.
delivery string dataset dataset or webhook. In demoMode, delivery is always dataset.
webhookUrl string — POST target when delivery is webhook. Slack, Discord, or any HTTP endpoint.
snapshotKey string robotstxt-snapshots Record key for change snapshots inside the fixed robotstxt-ai-checker-state Key-Value Store.
datasetMode string changes_only (live schema) all or changes_only. all emits every result. changes_only emits only changed results and emits zero rows and zero charges when nothing changed. The README input table still lists default all.
concurrency integer 5 Parallel requests; minimum 1, maximum 10.
dryRun boolean false Run without saving results or sending webhooks (for testing).
demoMode boolean false Checks only 1 domain, returns compact policy fields, and disables webhook/snapshot writes.

The schema sets additionalProperties to false. Required field is domains only.

Published Store Quickstart input that matches the live schema (prefill domains; this example sets datasetMode to all rather than the live default changes_only):

{
  "domains": [
    "google.com",
    "github.com",
    "nytimes.com",
    "openai.com"
  ],
  "delivery": "dataset",
  "snapshotKey": "robotstxt-snapshots",
  "datasetMode": "all",
  "concurrency": 5,
  "dryRun": false,
  "demoMode": false
}

Published Store example run input (demoMode true, schema-valid):

{
  "domains": [
    "openai.com",
    "github.com",
    "nytimes.com"
  ],
  "demoMode": true,
  "concurrency": 3,
  "dryRun": false
}

Run a robots.txt AI crawler policy check on Apify

Which README input examples are not live schema fields?

README examples labeled Single domain AI bot audit, Bulk competitor sites, and All-AI-bot policy snapshot send bots, emitPerBotDisposition, and detectAllAiBots. Those names are not in the live input schema. additionalProperties is false, so do not send them. Schema-valid published inputs use domains plus optional delivery, webhookUrl, snapshotKey, datasetMode, concurrency, dryRun, and demoMode.

What does a result contain?

The published README output table lists meta, results, and per-result domain, status, summary, aiPolicies, changes, checkedAt, actorName (stable actor identifier), rowType (robots_txt_policy_snapshot, robots_txt_missing, robots_txt_error, or robots_txt_invalid_domain), sourceUrl (validated final robots.txt URL, or null), fetchedAt, warnings, chargedEvent (apify-default-dataset-item for billed Dataset rows; otherwise null), billingEventName (apify-default-dataset-item), demoApplied, detailsMasked, and error.

The published Store README sample is truncated mid-object after the first GPTBot aiPolicies fields. It is a README illustration, not a live coverage guarantee. Published fragment:

{
  "meta": {
    "generatedAt": "2026-02-22T17:50:20.909Z",
    "totals": {
      "total": 1,
      "requestedDomains": 2,
      "processedDomains": 1,
      "withRobotsTxt": 1,
      "noRobotsTxt": 0,
      "invalidDomains": 0,
      "blockingAi": 0,
      "errors": 0
    },
    "demoApplied": true,
    "limits": {
      "maxDomains": 1,
      "compactPolicies": true,
      "webhookEnabled": false,
      "snapshotWriteEnabled": false
    },
    "upgradeHint": "Demo mode checks 1 domain, disables webhook delivery, and returns a compact policy view. Set demoMode=false to unlock bulk checks and full policy details."
  },
  "results": [
    {
      "domain": "openai.com",
      "status": "ok",
      "summary": {
        "totalCrawlers": 16,
        "blocked": 0,
        "partialBlock": 16,
        "allowed": 0,
        "changed": 0
      },
      "aiPolicies": [
        {
          "crawler": "GPTBot",
          "company": "OpenAI",
          "blocked": false,
          "partialBlock": true,
          "allowed": false

The live Store README cuts off at "allowed": false on that GPTBot object. The upgradeHint in the fragment restates that demo mode checks 1 domain, disables webhook delivery, and returns a compact policy view.

How do datasetMode, demoMode, snapshots, and delivery work?

Live schema default datasetMode is changes_only: it emits only changed results and emits zero rows and zero charges when nothing changed. all emits every result. The README Monitoring Contract still says datasetMode: "all" is the default and that changes_only emits the first observed baseline, then only changed domains. Snapshots persist in the fixed named Key-Value Store robotstxt-ai-checker-state; snapshotKey selects only the record inside that store (default robotstxt-snapshots). A pending state reservation is saved before delivery and committed only after delivery succeeds. Failed delivery stays blocked instead of being charged again automatically. Every emitted Dataset row is pushed individually through the Apify Actor SDK and must receive chargedCount >= 1. The run writes PHASE89_DELIVERY_AUDIT after delivery. The fixed snapshot write is not a Dataset event.

demoMode true checks only 1 domain, returns compact policy fields, and disables webhook/snapshot writes. delivery defaults to dataset; set webhook and provide webhookUrl for Slack, Discord, or any HTTP endpoint. In demoMode, delivery is always dataset. dryRun true runs without saving results or sending webhooks. Store Quickstart: start with store-input.example.json at demoMode=true, then store-input.templates.json with Demo Quickstart, Production Monitor, or Webhook Alert.

How is robots.txt Parser & AI Crawler Block Checker priced?

Billing is pay per event. The published Store card is $11.00 / 1,000 robots.txt AI policy rows. The billed event is robots.txt AI policy row at $0.011, charged for one delivered source-linked robots.txt AI crawler policy or missing-file row. The README Cost section names apify-default-dataset-item at $0.011 per delivered result and gives 1,000 delivered items = $11.00 as the example after scheduled start-fee removal. The same README notes that current transition pricing may include a temporary legacy start fee until that update activates. changes_only runs with no changed domains emit zero rows and incur zero result charges. The fixed snapshot write is not a Dataset event. No subscription is required — you only pay for what you use.

$11.00 per 1,000 robots.txt AI policy rows ($0.011 per delivered policy or missing-file row)

See robots.txt Parser & AI Crawler Block Checker pricing on Apify

Limits to keep in mind

Open robots.txt Parser & AI Crawler Block Checker on Apify

Related pages