robots.txt · AI crawlers
Audit robots.txt GPTBot and ClaudeBot disallow rules across domains
robots.txt Parser & AI Crawler Block Checker fetches public robots.txt files and parses crawl rules for AI crawlers and LLM-training bots (examples: GPTBot, ClaudeBot, Google-Extended). It extracts per-user-agent allow/disallow rules, restricted paths, and crawl-delay directives, then reports per-bot dispositions (blocked, partially blocked, allowed). This is a robots.txt policy audit — not HTML head-tag extraction, and not sitemap.xml URL discovery.
$11.00 per 1,000 robots.txt AI policy rows ($0.011 per delivered policy or missing-file row)
Open robots.txt Parser & AI Crawler Block Checker on Apify
robots.txt AI crawl policies, not page head tags and not sitemap XML
Use this page when the job is bulk parsing of robots.txt to see whether named AI crawlers are allowed, partially blocked, or disallowed. Use Meta Tag & OpenGraph Scraper when the job is title, canonical, Open Graph, Twitter Card, hreflang, JSON-LD, and the HTML robots meta tag on known page URLs. Use Sitemap Scraper & Analyzer when the job is nested sitemap.xml URL inventories.
| This Actor | Meta Tag & OpenGraph Scraper | Sitemap Scraper & Analyzer | |
|---|---|---|---|
| Intent | Audit robots.txt crawl policies for GPTBot, ClaudeBot, Google-Extended, and other AI crawlers | Technical SEO audit of HTML head metadata, social previews, JSON-LD, and the robots meta tag | Discover URLs from nested XML sitemaps; optional HEAD 200 check |
| Primary input | Required domains (maximum 500) |
Known page URLs in urls (maximum 500) |
Required sitemapUrls (sitemap.xml or a domain that auto-discovers /sitemap.xml) |
| What it reads | Public HTTP(S) robots.txt bodies and response headers |
HTML head and embedded JSON-LD from the initial HTML (no JavaScript render) | Sitemap XML and sitemap index files; optional HEAD on discovered URLs |
| Primary output | Per-domain policy snapshots: aiPolicies, dispositions, rowType, source robots.txt URL, change diffs |
Title, description, canonical, robots meta, Open Graph, Twitter Card, hreflang, JSON-LD, issue flags | URL inventory, directory structure, lastmod/changefreq analysis, optional HTTP 200 |
| GPTBot / ClaudeBot disallow | Yes — parsed from robots.txt user-agent directives | No — HTML robots meta is not a robots.txt AI-bot policy | No — sitemap XML does not encode GPTBot/ClaudeBot allow/disallow |
The live Store description is: Audit robots.txt files across thousands of websites to detect specific crawl policies, disallowed paths, and user-agents for GPTBot and ClaudeBot. The live schema description is: Analyze robots.txt files for AI crawler blocking policies. Track which sites block GPTBot, ClaudeBot, Google-Extended, and 13 more. Published README copy also names allow/disallow rules, restricted directory paths, crawl-delay directives, policy snapshots, change diffs, severity-scored findings with remediation guidance, and raw header/body evidence.
Source requests allow only public HTTP(S) hosts. Local, private, link-local, metadata, credential-bearing, and non-HTTP(S) URLs are rejected. Redirects are rechecked and limited to 3 hops; responses are capped at 1 MiB and requests time out after 10 seconds. The published README says this is passive scanning of public signals, not a penetration test: no exploitation, fuzzing, or auth bypass. Only scan sites you have authorization for.
Use cases
- Audit robots.txt policies for named AI crawlers (for example GPTBot and ClaudeBot) across large domain lists.
- Monitor domains on a schedule (daily or weekly) to detect policy changes and configuration drift.
- Bulk-verify competitor and market sites for AI-crawler dispositions and indexability signals.
- Produce audit-ready policy snapshots and evidence for compliance and security reviews, mapping findings to standards (WCAG, GDPR, SOC2).
- Generate severity-scored findings with remediation guidance for engineering and ops triage.
- Capture raw headers and responses to support compliance documentation.
- Integrate policy-change detection into CI/ops workflows to gate releases or trigger alerts.
Published README Use Cases table: Developers automate recurring fetches; data teams pipe structured output into analytics warehouses; ops teams monitor changes via webhook alerts; product managers track competitor/market signals. Published Key Features: compliance-first reports mapping findings to WCAG, GDPR, and SOC2; non-invasive scanning of observable public signals; severity-scored output with remediation guidance; delta-alerting of new findings since the last run via webhook; evidence export of raw headers/responses.
Store Quickstart starts with store-input.example.json at demoMode=true. Published templates are Demo Quickstart (trial run), Production Monitor (recurring dataset snapshots), and Webhook Alert (policy-change notifications). README tips: schedule weekly runs against production domains to catch config drift; use webhook delivery to pipe findings into a SIEM (Splunk, Datadog, Elastic); for CI, block releases on critical severity findings using exit codes. Combine with SSL/TLS Certificate Scraper when you also need certificate expiry coverage.
How is robots.txt Parser & AI Crawler Block Checker different from Meta Tag & OpenGraph Scraper and Sitemap Scraper & Analyzer?
This Actor fetches and parses robots.txt files. It extracts per-user-agent allow/disallow rules, restricted paths, and crawl-delay directives, then reports per-bot dispositions (blocked, partially blocked, allowed) for AI crawlers such as GPTBot, ClaudeBot, and Google-Extended. Meta Tag & OpenGraph Scraper audits HTML head tags on known page URLs: title, canonical, the HTML robots meta tag, Open Graph, Twitter Card, hreflang, and JSON-LD. That robots field is page-head metadata, not a robots.txt file. Sitemap Scraper & Analyzer reads sitemap.xml and nested sitemap indexes to discover URL inventories; it does not parse robots.txt AI crawler policies. Use this Actor to audit crawl policies in robots.txt, including GPTBot/ClaudeBot disallow. Use Meta Tag & OpenGraph Scraper for page head tags. Use Sitemap Scraper & Analyzer for XML sitemap URL discovery.
What input is required?
Open the Actor on the Apify Store and supply domains (required, maximum 500). Live input schema fields:
| Field | Type | Default | Notes |
|---|---|---|---|
domains |
string[] | required | Domains to analyze robots.txt for AI crawler policies (max 500). Schema prefill is google.com, github.com, nytimes.com, openai.com. |
delivery |
string | dataset |
dataset or webhook. In demoMode, delivery is always dataset. |
webhookUrl |
string | — | POST target when delivery is webhook. Slack, Discord, or any HTTP endpoint. |
snapshotKey |
string | robotstxt-snapshots |
Record key for change snapshots inside the fixed robotstxt-ai-checker-state Key-Value Store. |
datasetMode |
string | changes_only (live schema) |
all or changes_only. all emits every result. changes_only emits only changed results and emits zero rows and zero charges when nothing changed. The README input table still lists default all. |
concurrency |
integer | 5 | Parallel requests; minimum 1, maximum 10. |
dryRun |
boolean | false | Run without saving results or sending webhooks (for testing). |
demoMode |
boolean | false | Checks only 1 domain, returns compact policy fields, and disables webhook/snapshot writes. |
The schema sets additionalProperties to false. Required field is domains only.
Published Store Quickstart input that matches the live schema (prefill domains; this example sets datasetMode to all rather than the live default changes_only):
{
"domains": [
"google.com",
"github.com",
"nytimes.com",
"openai.com"
],
"delivery": "dataset",
"snapshotKey": "robotstxt-snapshots",
"datasetMode": "all",
"concurrency": 5,
"dryRun": false,
"demoMode": false
}
Published Store example run input (demoMode true, schema-valid):
{
"domains": [
"openai.com",
"github.com",
"nytimes.com"
],
"demoMode": true,
"concurrency": 3,
"dryRun": false
}
Run a robots.txt AI crawler policy check on Apify
Which README input examples are not live schema fields?
README examples labeled Single domain AI bot audit, Bulk competitor sites, and All-AI-bot policy snapshot send bots, emitPerBotDisposition, and detectAllAiBots. Those names are not in the live input schema. additionalProperties is false, so do not send them. Schema-valid published inputs use domains plus optional delivery, webhookUrl, snapshotKey, datasetMode, concurrency, dryRun, and demoMode.
What does a result contain?
The published README output table lists meta, results, and per-result domain, status, summary, aiPolicies, changes, checkedAt, actorName (stable actor identifier), rowType (robots_txt_policy_snapshot, robots_txt_missing, robots_txt_error, or robots_txt_invalid_domain), sourceUrl (validated final robots.txt URL, or null), fetchedAt, warnings, chargedEvent (apify-default-dataset-item for billed Dataset rows; otherwise null), billingEventName (apify-default-dataset-item), demoApplied, detailsMasked, and error.
The published Store README sample is truncated mid-object after the first GPTBot aiPolicies fields. It is a README illustration, not a live coverage guarantee. Published fragment:
{
"meta": {
"generatedAt": "2026-02-22T17:50:20.909Z",
"totals": {
"total": 1,
"requestedDomains": 2,
"processedDomains": 1,
"withRobotsTxt": 1,
"noRobotsTxt": 0,
"invalidDomains": 0,
"blockingAi": 0,
"errors": 0
},
"demoApplied": true,
"limits": {
"maxDomains": 1,
"compactPolicies": true,
"webhookEnabled": false,
"snapshotWriteEnabled": false
},
"upgradeHint": "Demo mode checks 1 domain, disables webhook delivery, and returns a compact policy view. Set demoMode=false to unlock bulk checks and full policy details."
},
"results": [
{
"domain": "openai.com",
"status": "ok",
"summary": {
"totalCrawlers": 16,
"blocked": 0,
"partialBlock": 16,
"allowed": 0,
"changed": 0
},
"aiPolicies": [
{
"crawler": "GPTBot",
"company": "OpenAI",
"blocked": false,
"partialBlock": true,
"allowed": false
The live Store README cuts off at "allowed": false on that GPTBot object. The upgradeHint in the fragment restates that demo mode checks 1 domain, disables webhook delivery, and returns a compact policy view.
How do datasetMode, demoMode, snapshots, and delivery work?
Live schema default datasetMode is changes_only: it emits only changed results and emits zero rows and zero charges when nothing changed. all emits every result. The README Monitoring Contract still says datasetMode: "all" is the default and that changes_only emits the first observed baseline, then only changed domains. Snapshots persist in the fixed named Key-Value Store robotstxt-ai-checker-state; snapshotKey selects only the record inside that store (default robotstxt-snapshots). A pending state reservation is saved before delivery and committed only after delivery succeeds. Failed delivery stays blocked instead of being charged again automatically. Every emitted Dataset row is pushed individually through the Apify Actor SDK and must receive chargedCount >= 1. The run writes PHASE89_DELIVERY_AUDIT after delivery. The fixed snapshot write is not a Dataset event.
demoMode true checks only 1 domain, returns compact policy fields, and disables webhook/snapshot writes. delivery defaults to dataset; set webhook and provide webhookUrl for Slack, Discord, or any HTTP endpoint. In demoMode, delivery is always dataset. dryRun true runs without saving results or sending webhooks. Store Quickstart: start with store-input.example.json at demoMode=true, then store-input.templates.json with Demo Quickstart, Production Monitor, or Webhook Alert.
How is robots.txt Parser & AI Crawler Block Checker priced?
Billing is pay per event. The published Store card is $11.00 / 1,000 robots.txt AI policy rows. The billed event is robots.txt AI policy row at $0.011, charged for one delivered source-linked robots.txt AI crawler policy or missing-file row. The README Cost section names apify-default-dataset-item at $0.011 per delivered result and gives 1,000 delivered items = $11.00 as the example after scheduled start-fee removal. The same README notes that current transition pricing may include a temporary legacy start fee until that update activates. changes_only runs with no changed domains emit zero rows and incur zero result charges. The fixed snapshot write is not a Dataset event. No subscription is required — you only pay for what you use.
$11.00 per 1,000 robots.txt AI policy rows ($0.011 per delivered policy or missing-file row)
See robots.txt Parser & AI Crawler Block Checker pricing on Apify
Limits to keep in mind
- Maximum 500 domains per run. Required field is
domains. concurrencyis 1–10 (default 5). Requests time out after 10 seconds. Responses are capped at 1 MiB. Redirects are limited to 3 hops.- Public HTTP(S) hosts only. Local, private, link-local, metadata, credential-bearing, and non-HTTP(S) URLs are rejected.
demoModechecks 1 domain, returns compact fields, and disables webhook/snapshot writes.- Live schema names are
domains,delivery,webhookUrl,snapshotKey,datasetMode,concurrency,dryRun, anddemoMode. Unpublished README aliasesbots,emitPerBotDisposition, anddetectAllAiBotsare not working inputs. - This is a robots.txt policy audit, not HTML head-tag extraction and not sitemap XML discovery. It is not a penetration test.
- Published README FAQ: only scan sites you have authorization for; weekly scans for production domains, daily if config-change velocity is high; webhook or Dataset API for compliance-tool export; evidence artifacts may support SOC2 CC7.1 continuous monitoring, but the Actor is not itself a SOC2 certification.
Open robots.txt Parser & AI Crawler Block Checker on Apify
Related pages
- Broken Link Checker — HTML crawl for 404s, not robots.txt AI crawler policies
- Meta Tag & OpenGraph Scraper — HTML head tags, Open Graph, JSON-LD, and the robots meta tag, not robots.txt AI crawler policies
- Sitemap Scraper & Analyzer — nested sitemap.xml URL inventories, not robots.txt allow/disallow
- SSL/TLS Certificate Scraper — certificate expiry and fingerprints; the README pairs it with this Actor for layered coverage
- AI Brand Visibility Scraper | ChatGPT Search Rank — published README see-also for measuring public AI brand visibility after crawler policy changes; not a robots.txt parser
- Structured Data Scraper & Validator — schema.org JSON-LD on page HTML, not robots.txt
- Bulk URL Status Checker — HTTP status on a known URL list, not robots.txt parsing
- Website Content Extractor — cleaned page body, not crawl-policy files
- DMARC & Email Security Checker — SPF/DKIM/DMARC DNS, not robots.txt
- DNS Propagation Checker — public DNS records, not robots.txt
- HHS Healthcare Data Breach Change Scraper — HHS OCR disclosure changes, not robots.txt
- Site Governance Monitor — robots + sitemap + schema drift in one watch
- GDPR & CCPA Cookie Compliance Scraper — cookie banners/policies, not robots.txt
- Site QA Indexability AI Crawler Report — robots.txt and llms.txt signals.
- Technical SEO & AI Crawler Audit — canonical/noindex/sitemap regressions.
- Tools