robots · sitemap · schema
Post-deploy robots.txt, sitemap, and schema governance diffs
This Actor checks domains you supply for robots.txt AI-crawler rules, sitemap.xml health, and JSON-LD/Microdata on sampled paths (homepage-first by default). One summary-first payload per domain with alerts.
Not a sitewide HTML crawl, not a known-URL status checker, and not a standalone robots.txt-only AI policy audit. Homepage-first samplePaths default to [/].
from $9.00 / 1,000 delivered monitoring result rows ($0.009 per new or changed row). No Actor Start. Unchanged runs with emitUnchanged false are free.
Open Site Governance Monitor on Apify
Site Governance Monitor, not a neighboring Actor
Use this page for this Actor’s job. Use robots.txt Parser & AI Crawler Block Checker for robots.txt AI policy only; Use Sitemap Scraper & Analyzer for sitemap.xml URL inventory; Use Broken Link Checker for HTML crawl for 404s.
| This Actor | robots.txt Parser & AI Crawler Block Checker | Sitemap Scraper & Analyzer | Broken Link Checker | |
|---|---|---|---|---|
| Intent | Site Governance Monitor | robots.txt AI policy only | sitemap.xml URL inventory | HTML crawl for 404s |
| Primary input | domains |
domains max 500 | sitemapUrls | startUrls max 10 |
| What it reads | Public sources listed on the Store page | robots.txt | sitemap XML | crawled HTML anchors |
| Primary output | Dataset rows billed per live PPE | per-bot allow/disallow rows | URL inventory | broken-link rows |
| Not this job | Not a sitewide HTML crawl, not a known-URL status checker, and not a standalone robots.txt-only AI policy audit. Homepage-first samplePaths default to [/]. | Not sitemap+schema bundle | Not robots+schema governance score | Not robots/sitemap/schema bundle |
Store ID: taroyamada/site-governance-monitor. Respect source terms, robots.txt, and rate limits.
Use cases
- Release QA for robots/sitemap/schema drift
- AI crawler allow/block watches
- Portfolio governance summaries
How is Site Governance Monitor different from robots.txt Parser & AI Crawler Block Checker and Sitemap Scraper & Analyzer?
Site Governance Monitor — Robots, Sitemap & Schema (taroyamada/site-governance-monitor): This Actor checks domains you supply for robots.txt AI-crawler rules, sitemap.xml health, and JSON-LD/Microdata on sampled paths (homepage-first by default). One summary-first payload per domain with alerts. Not a sitewide HTML crawl, not a known-URL status checker, and not a standalone robots.txt-only AI policy audit. Homepage-first samplePaths default to [/]. robots.txt Parser & AI Crawler Block Checker is for robots.txt AI policy only (input domains max 500; robots.txt; per-bot allow/disallow rows). Not sitemap+schema bundle. Sitemap Scraper & Analyzer is for sitemap.xml URL inventory (input sitemapUrls; sitemap XML; URL inventory). Not robots+schema governance score. Broken Link Checker is for HTML crawl for 404s (input startUrls max 10; crawled HTML anchors; broken-link rows). Not robots/sitemap/schema bundle.
What input is required?
Live required fields: domains. exampleRunInput is vercel.com, samplePaths [/], all three checks true, concurrency 1, maxSitemapUrls 5000. Schema concurrency default 3. Keep snapshotKey stable across deploys.
| Field | Type | Default | Notes |
|---|---|---|---|
domains |
array required | empty |
Sites or domains to monitor. Starter quickstart: begin with 1-3 sites for a lightweight first success. Homepage-first runs stay intentionally small, while larger recurring watches can expand to broader portfolios with… |
samplePaths |
array | empty |
Sample paths for release QA. Path-only routes to validate on every domain. Keep the starter quickstart homepage-first with ["/"], then add /pricing and /docs when you want broader release-QA template coverage. Default… |
delivery |
string | dataset |
Delivery mode. Starter path: dataset keeps the first run low-friction and still writes the full summary-first payload to OUTPUT. Advanced path: webhook sends the same meta/alerts/results payload to your endpoint for r… |
webhookUrl |
string | empty |
Webhook URL. Advanced delivery only: required when delivery is webhook. Must be a valid http(s) URL. The payload includes the executive summary, flattened alerts, workflow metadata, and per-domain governance results. |
snapshotKey |
string | site-governance-monitor-snapshots |
Snapshot key for recurring checks. Keep this stable when you move from the homepage-first quickstart to recurring release-QA, portfolio, or webhook workflows so governance drift stays comparable run to run. |
checkAiBots |
boolean | true |
Robots.txt monitor. Monitor robots.txt for missing files, AI crawler allow/block rules, and drift after releases. |
checkSchema |
boolean | true |
Schema validator monitor. Validate JSON-LD and Microdata on homepage, pricing, docs, and other release-sensitive templates. |
checkSitemap |
boolean | true |
Sitemap monitor. Monitor sitemap.xml reachability, freshness, robots.txt declarations, and URL inventory drift. |
followRedirects |
boolean | true |
Follow redirects. Follow redirects before grading sitemap and sampled page responses so canonical properties are scored correctly. |
concurrency |
integer | 3 |
Concurrency. Starter runs usually only need 1-3 parallel checks. Increase it for broader recurring portfolios when you want more properties covered in one run. |
batchDelayMs |
integer | 250 |
Batch delay (ms). Pause between batches to keep recurring sweeps polite and stable across portfolios and platform estates. |
requestTimeoutSecs |
integer | 15 |
Request timeout (seconds). Timeout used for robots.txt, sitemap, and sampled page fetches during each recurring governance check. |
maxSitemapUrls |
integer | 5000 |
Max sitemap URLs. Maximum number of sitemap URLs to parse when expanding sitemap indexes for larger sites or portfolios. |
emitUnchanged |
boolean | false |
Emit unchanged site rows. Keep disabled for recurring monitoring so unchanged runs write zero dataset rows and incur zero result charges. |
dryRun |
boolean | false |
Dry run. Preview the governance bundle without saving snapshots, dataset rows, or sending webhooks. Useful for validation, not for capturing a recurring baseline. |
Published Store example run input (omitted fields take schema defaults):
{
"domains": [
"vercel.com"
],
"samplePaths": [
"/"
],
"delivery": "dataset",
"snapshotKey": "site-governance-homepage-quickstart",
"checkAiBots": true,
"checkSchema": true,
"checkSitemap": true,
"concurrency": 1,
"batchDelayMs": 250,
"requestTimeoutSecs": 15,
"maxSitemapUrls": 5000,
"dryRun": false
}
Run Site Governance Monitor on Apify
How do dataset, webhook, and dry-run delivery work?
delivery defaults to dataset on the live schema. Dataset output is the billable surface when rows are written. webhookUrl is used when delivery is webhook (and typically not during dryRun). dryRun true validates or samples without the usual dataset/webhook side effects described on the Store schema. emitUnchanged controls whether unchanged watches write rows. Unchanged runs that write zero default-dataset rows typically charge $0.00 on current PPE.
What does a result contain?
Published README Output Example / Sample Output JSON.
{
"meta": {
"executiveSummary": {
"overallStatus": "attention_needed",
"recommendedCadence": "daily"
},
"runProfile": {
"tier": "starter",
"label": "Starter first-success path"
},
"upgradeSuggestions": [
{
"type": "webhook",
"templateId": "action_needed_webhook",
"title": "Route action-needed domains to your endpoint"
}
],
"nextWorkflow": {
"type": "same_actor_template",
"id": "action_needed_webhook",
"title": "Next best step: Action-Needed Webhook Handoff"
}
},
"alerts": [
{
"domain": "client-release.example",
"severity": "high",
"component": "sitemapHealth",
"type": "sitemap_missing",
"message": "No reachable XML sitemap was found for this domain."
}
],
"results": [
{
"domain": "client-release.example",
"status": "changed",
"severity": "high",
"brief": "3 alert(s): No reachable XML sitemap was found for this domain.",
"recommendedActions": [
"Publish a reachable XML sitemap for the domain and keep it updated.",
"Publish a robots.txt file so the robots.txt monitor can confirm which AI crawlers you allow or block."
]
}
]
}
There is no published output JSON schema on the Store page.
How is Site Governance Monitor priced?
Billing is pay per event. The live Store card is from $9.00 / 1,000 delivered monitoring result rows ($0.009 per new or changed row). No Actor Start. Unchanged runs with emitUnchanged false are free. Current PPE:
| Event | Price | Emitted when |
|---|---|---|
apify-default-dataset-item (Delivered monitoring result row) |
$0.009 | Charged only when a new or changed monitoring result row is delivered. Unchanged runs are free. |
The published README Cost block is stale versus the live Store pricing tab. README Cost quotes actor-start $0.01 + dataset-item $0.003. Live: Delivered monitoring result row $0.009, no start. README Input Examples that send urls are not live schema (use domains + samplePaths). This page quotes live PPE only.
from $9.00 / 1,000 delivered monitoring result rows ($0.009 per new or changed row). No Actor Start. Unchanged runs with emitUnchanged false are free.
See Site Governance Monitor pricing on Apify
Limits to keep in mind
- Not a full-site crawler
- samplePaths are path-only
- maxSitemapUrls 100–50000
- README urls examples are stale
- Respect source terms, robots.txt, and rate limits.
Open Site Governance Monitor on Apify
Related pages
- robots.txt Parser & AI Crawler Block Checker — robots.txt-only AI bot audit.
- Sitemap Scraper & Analyzer — Sitemap URL discovery without schema/robots bundle.
- Structured Data Scraper & Validator — JSON-LD/Microdata score on known URLs, not domain governance diffs.
- GDPR & CCPA Cookie Compliance Scraper — Cookie banners/policies, not robots/sitemap/schema
- Tools