Site QA · indexability
Public pages, robots.txt, and llms.txt indexability issues
This Actor unofficially audits user-supplied public pages, robots.txt, and llms.txt signals for AI crawler indexability issues and source-linked report rows.
Not a robots.txt parser-only check (use robots.txt Parser & AI Crawler Block Checker). Not a sitemap enumerator. Not a broken-link crawler. You supply the URLs.
from $30.00 / 1,000 AI crawler policy checked ($0.03 per delivered ai-crawler-policy-checked); $0.12 per Indexability issue detected; $6.00 per AI crawler indexability report; $8.00 per Indexability export generated. No Actor Start.
Open Site QA Indexability AI Crawler Report on Apify
Site QA Indexability AI Crawler Report, not a neighboring Actor
Use this page for this Actor’s job. Use robots.txt Parser & AI Crawler Block Checker for GPTBot/ClaudeBot disallow audits; Use Sitemap Scraper & Analyzer for Nested XML sitemap URLs; Use Broken Link Checker for Crawl start URLs for 404s.
| This Actor | robots.txt Parser & AI Crawler Block Checker | Sitemap Scraper & Analyzer | Broken Link Checker | |
|---|---|---|---|---|
| Intent | Site QA Indexability AI Crawler Report | GPTBot/ClaudeBot disallow audits | Nested XML sitemap URLs | Crawl start URLs for 404s |
| Primary input | urls | origin URLs | sitemap URLs | start URLs |
| What it reads | User-supplied public pages, robots.txt, llms.txt | robots.txt | sitemap.xml | Linked HTTP statuses |
| Primary output | Policy-check, issue, report, and export events | Disallow rows | Sitemap URL rows | Broken-link rows |
| Not this job | Not a robots.txt parser-only check (use robots.txt Parser & AI Crawler Block Checker). Not a sitemap enumerator. Not a broken-link crawler. You supply the URLs. | Not an llms.txt + indexability report pack. | Not robots/llms.txt indexability issues. | Not AI crawler policy checks. |
Store ID: taroyamada/site-qa-indexability-ai-crawler-report-scraper. Respect source terms, robots.txt, and rate limits.
Use cases
- Known URLs with checkRobotsTxt / checkLlmsTxt
- Keep generateReport/emitExport off on first run
- authorizedUseConfirmed for live audits
How is Site QA Indexability AI Crawler Report different from robots.txt Parser & AI Crawler Block Checker and Sitemap Scraper & Analyzer?
Site QA Indexability AI Crawler Report (taroyamada/site-qa-indexability-ai-crawler-report-scraper): This Actor unofficially audits user-supplied public pages, robots.txt, and llms.txt signals for AI crawler indexability issues and source-linked report rows. Not a robots.txt parser-only check (use robots.txt Parser & AI Crawler Block Checker). Not a sitemap enumerator. Not a broken-link crawler. You supply the URLs. robots.txt Parser & AI Crawler Block Checker is for GPTBot/ClaudeBot disallow audits (input origin URLs; robots.txt; Disallow rows). Not an llms.txt + indexability report pack. Sitemap Scraper & Analyzer is for Nested XML sitemap URLs (input sitemap URLs; sitemap.xml; Sitemap URL rows). Not robots/llms.txt indexability issues. Broken Link Checker is for Crawl start URLs for 404s (input start URLs; Linked HTTP statuses; Broken-link rows). Not AI crawler policy checks.
What input is required?
Live required fields: urls. Published exampleRunInput is shown below.
| Field | Type | Default | Notes |
|---|---|---|---|
urls |
any[] | ["https://example.com/?siteQaCanary=indexability-ai-crawler-v1"] |
Public page URLs. Public URLs that you are allowed to audit. Required. |
aiCrawlerUserAgents |
any[] | ["GPTBot", "Google-Extended", "PerplexityBot", "ClaudeBot"] |
AI crawler user agents. Crawler user-agent names to inspect in robots.txt. |
maxPages |
integer | 10 |
Max pages. Maximum number of input pages to check. minimum=1 maximum=100 |
checkRobotsTxt |
boolean | true |
Check robots.txt. Fetch and inspect robots.txt at each site origin. |
checkLlmsTxt |
boolean | true |
Check llms.txt. Fetch and inspect llms.txt at each site origin. |
authorizedUseConfirmed |
boolean | false |
Authorized use confirmed. Required for non-dry runs. Confirms each URL is owned by you, your client, or otherwise authorized for this audit. |
emitPageRows |
boolean | false |
Emit page snapshot rows. Emit optional public page indexability snapshot rows. |
generateReport |
boolean | true |
Generate report row. Generate a site-level AI crawler indexability report row. |
emitExport |
boolean | false |
Generate export row. Generate an export row for handoff workflows. |
emitUnchanged |
boolean | false |
Emit unchanged rows. When false, repeated unchanged runs emit zero rows and zero charges. |
dryRun |
boolean | true |
Dry run. Emit local sample rows without charging. |
initialRunMode |
string enum | emit_backfill |
Initial run mode. Deployment canary control used by automation; emit_backfill allows first-run proof rows, baseline_only is used for no-change proof. enum: emit_backfill, baseline_only |
snapshotKey |
string | empty |
Snapshot key. Optional state namespace for canary and recurring no-change proof runs. |
Published Store example run input (omitted fields take schema defaults):
{
"urls": [
"https://example.com/?siteQaCanary=indexability-ai-crawler-v1"
],
"maxPages": 1,
"dryRun": true
}
Run Site QA Indexability AI Crawler Report on Apify
How do dataset, webhook, and dry-run delivery work?
Dataset output is the billable surface when rows are written. See live PPE for which events bill. dryRun true validates or samples without the usual dataset/webhook side effects described on the Store schema. Unchanged runs that write zero default-dataset rows typically charge $0.00 on current PPE.
What does a result contain?
README output fence is truncated or invalid JSON on Store. Illustration head: { "rowType": "ai_crawler_indexability_report", "billingEventName": "ai-crawler-indexability-report", "sourceUrl": "https://example.com/robots.txt", "checkedUrlCount": 3, "issueCount": 1, "reportStatus": "action_needed" } Treat README samples as illustrations, not a live coverage guarantee. There is no published output JSON schema on the Store page.
How is Site QA Indexability AI Crawler Report priced?
Billing is pay per event. The live Store card is from $30.00 / 1,000 AI crawler policy checked ($0.03 per delivered ai-crawler-policy-checked); $0.12 per Indexability issue detected; $6.00 per AI crawler indexability report; $8.00 per Indexability export generated. No Actor Start. Current PPE:
| Event | Price | Emitted when |
|---|---|---|
ai-crawler-policy-checked (AI crawler policy checked) |
$0.03 | Charged for each source-linked robots.txt or llms.txt policy observation row. |
indexability-issue-detected (Indexability issue detected) |
$0.12 | Charged for each source-linked AI crawler or page indexability issue row. |
ai-crawler-indexability-report (AI crawler indexability report) |
$6.00 | Charged for each site-level AI crawler indexability report row. |
indexability-export-generated (Indexability export generated) |
$8.00 | Charged for each generated indexability handoff export row. |
from $30.00 / 1,000 AI crawler policy checked ($0.03 per delivered ai-crawler-policy-checked); $0.12 per Indexability issue detected; $6.00 per AI crawler indexability report; $8.00 per Indexability export generated. No Actor Start.
See Site QA Indexability AI Crawler Report pricing on Apify
Limits to keep in mind
- You supply urls
- urls required
- Report and export are separate billed events
- Not a site-wide crawler beyond maxPages
- Respect source terms, robots.txt, and rate limits.
Open Site QA Indexability AI Crawler Report on Apify
Related pages
- robots.txt Parser & AI Crawler Block Checker — Not an llms.txt + indexability report pack.
- Sitemap Scraper & Analyzer — Not robots/llms.txt indexability issues.
- Broken Link Checker — Not AI crawler policy checks.
- Technical SEO & AI Crawler Audit — canonical/noindex/sitemap regressions.
- Site QA Broken Link Report — URL health issues, not llms.txt.
- Tools