Technical SEO
robots, noindex, sitemap, canonical, and AI crawler access changes
Audit public pages you are authorized to test for robots, noindex, sitemap, canonical, and AI crawler access changes. Deliver source-linked regressions, portfolio reports, and optional exports.
Not an llms.txt + indexability report pack (use Site QA Indexability AI Crawler Report). Not a robots.txt parser-only check. Not a sitemap enumerator. You supply the URLs and must confirm authorized use for live runs.
from $12.00 / 1,000 SEO page audited ($0.01 per delivered seo-page-audited); $0.08 per SEO regression detected; $3.00 per SEO portfolio report; $2.00 per SEO agency export. No Actor Start.
Open Technical SEO & AI Crawler Audit on Apify
Technical SEO & AI Crawler Audit, not a neighboring Actor
Use this page for this Actor’s job. Use Site QA Indexability AI Crawler Report for robots.txt and llms.txt signals; Use robots.txt Parser & AI Crawler Block Checker for GPTBot/ClaudeBot disallow audits; Use Sitemap Scraper & Analyzer for Nested XML sitemap URLs.
| This Actor | Site QA Indexability AI Crawler Report | robots.txt Parser & AI Crawler Block Checker | Sitemap Scraper & Analyzer | |
|---|---|---|---|---|
| Intent | Technical SEO / AI crawler regression on supplied URLs | robots.txt and llms.txt signals | GPTBot/ClaudeBot disallow audits | Nested XML sitemap URLs |
| Primary input | urls | urls | origin URLs | sitemap URLs |
| What it reads | User-supplied public pages, robots.txt, optional llms.txt | robots.txt / llms.txt | robots.txt | sitemap.xml |
| Primary output | Page-audited, regression, report, and export events | Policy-check and indexability events | Disallow rows | Sitemap URL rows |
| Not this job | Not an llms.txt + indexability report pack (use Site QA Indexability AI Crawler Report). Not a robots.txt parser-only check. Not a sitemap enumerator. You supply the URLs and must confirm authorized use for live runs. | Not a technical SEO regression pack covering canonical/noindex. | Not a canonical/noindex/sitemap regression report. | Not robots/canonical/noindex regressions. |
Store ID: taroyamada/technical-seo-portfolio-regression-report. Respect source terms, robots.txt, and rate limits.
Use cases
- A few authorized URLs with dryRun true first
- authorizedUseConfirmed required for non-dry live audits
- Keep generateReport/emitExport off until the URL list is stable
How is Technical SEO & AI Crawler Audit different from Site QA Indexability AI Crawler Report and robots.txt Parser & AI Crawler Block Checker?
Technical SEO & AI Crawler Audit (taroyamada/technical-seo-portfolio-regression-report): Audit public pages you are authorized to test for robots, noindex, sitemap, canonical, and AI crawler access changes. Deliver source-linked regressions, portfolio reports, and optional exports. Not an llms.txt + indexability report pack (use Site QA Indexability AI Crawler Report). Not a robots.txt parser-only check. Not a sitemap enumerator. You supply the URLs and must confirm authorized use for live runs. Site QA Indexability AI Crawler Report is for robots.txt and llms.txt signals (input urls; robots.txt / llms.txt; Policy-check and indexability events). Not a technical SEO regression pack covering canonical/noindex. robots.txt Parser & AI Crawler Block Checker is for GPTBot/ClaudeBot disallow audits (input origin URLs; robots.txt; Disallow rows). Not a canonical/noindex/sitemap regression report. Sitemap Scraper & Analyzer is for Nested XML sitemap URLs (input sitemap URLs; sitemap.xml; Sitemap URL rows). Not robots/canonical/noindex regressions.
What input is required?
Live required fields: urls. Published exampleRunInput is shown below.
| Field | Type | Default | Notes |
|---|---|---|---|
urls |
any[] | ["https://example.com/"] |
Public page URLs. Public URLs that you are allowed to audit. Required. |
aiCrawlerUserAgents |
any[] | ["GPTBot", "ClaudeBot"] |
AI crawler user agents. Crawler user-agent names to inspect in robots.txt. |
maxPages |
integer | 3 |
Max pages. Maximum number of input pages to check. minimum=1 maximum=100 |
checkRobotsTxt |
boolean | true |
Check robots.txt. Fetch and inspect robots.txt at each site origin. |
checkLlmsTxt |
boolean | true |
Check llms.txt. Fetch and inspect llms.txt at each site origin. |
authorizedUseConfirmed |
boolean | false |
Authorized use confirmed. Required for non-dry runs. Confirms each URL is owned by you, your client, or otherwise authorized for this audit. |
emitPageRows |
boolean | false |
Emit page snapshot rows. Emit optional public page indexability snapshot rows. |
generateReport |
boolean | false |
Generate report. Generate a site-level SEO portfolio report row. |
emitExport |
boolean | false |
Generate export row. Generate an export row for handoff workflows. |
emitUnchanged |
boolean | false |
Emit unchanged. When false, repeated unchanged runs emit zero rows and zero charges. |
dryRun |
boolean | true |
Dry run. Emit local sample rows without charging. |
initialRunMode |
string enum | "baseline_only" |
Initial run mode. Initial run mode for this run. enum: baseline_only, emit_backfill |
snapshotKey |
string | empty |
Snapshot key. Optional state namespace for canary and recurring no-change proof runs. |
watchlists |
any[] | [] |
Portfolio watchlists. Named portfolio watch definitions. |
monitorKey |
string | empty |
Monitor key. Monitor key for this run. |
emitRawRows |
boolean | false |
Emit raw rows. Emit optional public page indexability snapshot rows. |
maxChargeUsd |
number | 0 |
Maximum charge (USD). Maximum charge (USD) for this run. minimum=0 |
Published Store example run input (omitted fields take schema defaults):
{
"urls": [
"https://example.com/"
],
"maxPages": 3,
"dryRun": true,
"maxChargeUsd": 0
}
Run Technical SEO & AI Crawler Audit on Apify
How do dataset, webhook, and dry-run delivery work?
Dataset output is the billable surface when rows are written. See live PPE for which events bill. dryRun true validates or samples without the usual dataset/webhook side effects described on the Store schema. Unchanged runs that write zero default-dataset rows typically charge $0.00 on current PPE.
What does a result contain?
Published README Output Example / Sample Output JSON. Illustration: { "actorName": "technical-seo-portfolio-regression-report", "rowType": "indexability_issue", "billingEventName": "seo-regression-detected", "issueType": "ai_crawler_disallowed_by_robots", "severity": "high", "sourceUrl": "https://example.com/?siteQaCanary=indexability-ai-crawler-v1" } Treat README samples as illustrations, not a live coverage guarantee. There is no published output JSON schema on the Store page.
How is Technical SEO & AI Crawler Audit priced?
Billing is pay per event. The live Store card is from $12.00 / 1,000 SEO page audited ($0.01 per delivered seo-page-audited); $0.08 per SEO regression detected; $3.00 per SEO portfolio report; $2.00 per SEO agency export. No Actor Start. Current PPE:
| Event | Price | Emitted when |
|---|---|---|
seo-page-audited (SEO page audited) |
$0.01 | Charged only when a seo page audited row is delivered. |
seo-regression-detected (SEO regression detected) |
$0.08 | Charged only when a seo regression detected row is delivered. |
seo-portfolio-report (SEO portfolio report) |
$3.00 | Charged only when a seo portfolio report row is delivered. |
seo-agency-export (SEO agency export) |
$2.00 | Charged only when a seo agency export row is delivered. |
from $12.00 / 1,000 SEO page audited ($0.01 per delivered seo-page-audited); $0.08 per SEO regression detected; $3.00 per SEO portfolio report; $2.00 per SEO agency export. No Actor Start.
See Technical SEO & AI Crawler Audit pricing on Apify
Limits to keep in mind
- urls required
- authorizedUseConfirmed for live audits
- Schema default dryRun is true
- Not a site-wide crawler beyond maxPages
- Respect source terms, robots.txt, and rate limits.
Open Technical SEO & AI Crawler Audit on Apify
Related pages
- Site QA Indexability AI Crawler Report — Not a technical SEO regression pack covering canonical/noindex.
- robots.txt Parser & AI Crawler Block Checker — Not a canonical/noindex/sitemap regression report.
- Sitemap Scraper & Analyzer — Not robots/canonical/noindex regressions.
- Website Accessibility Checker — WCAG 2.1 URL grades, not canonical/noindex.
- WCAG Accessibility Checker — Lighthouse Regression Report — Lighthouse + axe-core, not technical SEO.
- Tools