Broken links

Crawl start URLs for broken links, 404s, and redirect chains

Broken Link Checker starts from public pages you supply, follows internal links up to a bounded depth, and reports broken links with HTTP status codes, redirect chains, anchor text, and internal/external class. It is crawl-based link QA — not a known URL-list status checker, not sitemap XML discovery, and not a robots.txt AI-crawler audit.

from $8.00 / 1,000 results ($0.001 Actor Start + $0.008 per default-dataset result)

Open Broken Link Checker on Apify

HTML crawl for broken links, not a known URL list

Use this page when you have a start URL and need to discover broken links from source HTML. Use Bulk URL Status Checker when you already have the URL list. Use Sitemap Scraper & Analyzer to inventory sitemap.xml URLs. Use robots.txt Parser & AI Crawler Block Checker for GPTBot/ClaudeBot allow/disallow rules.

This Actor Bulk URL Status Checker Sitemap Scraper & Analyzer
Intent Discover broken links from crawled HTML HTTP status QA on a known URL list Discover URLs from nested sitemap.xml
Input startUrls (max 10) urls (max 1000) sitemapUrls
What it reads Start pages plus internal hops up to maxDepth Live HTTP(S) for URLs you already have; no crawl sitemap.xml and sitemap indexes
Primary output Broken-link findings: status, redirects, anchor, severity Status code, final URL, redirect chain, timing URL inventory, lastmod, optional HEAD 200
Not this job Status-only pass over a fixed list HTML crawl to find links Anchor-text crawl of HTML pages

Store ID: taroyamada/broken-link-checker. Crawl only public sites you control or are authorized to check, with reasonable limits.

Use cases

How is Broken Link Checker different from Bulk URL Status Checker and Sitemap Scraper?

This Actor crawls from start URLs you supply, follows internal links up to maxDepth, and reports broken links, HTTP status codes, and redirect chains. Required input is startUrls (maximum 10). Store ID taroyamada/broken-link-checker. Bulk URL Status Checker checks a known URL list (urls, maximum 1000) for HTTP status, redirect chains, and timing; it does not crawl pages. Sitemap Scraper & Analyzer parses sitemap.xml into URL inventories and optionally HEADs discovered URLs; it does not crawl HTML for anchors. robots.txt Parser & AI Crawler Block Checker audits AI crawler allow/disallow rules, not page links. Use this Actor to discover broken links from source HTML. Use Bulk URL Status Checker when you already have the URL list.

What input is required?

startUrls is required: public start pages, maximum 10. Schema prefill is https://example.com. additionalProperties is false. README examples that send includeExternalLinks, pages, maxLinksPerPage, or externalOnly are not live schema fields. Do not send them.

Field Type Default Notes
startUrls string[] required Max 10. Prefill https://example.com. For a fixed list, use Bulk URL Status Checker
maxDepth integer 2 Internal-link hops; maximum 5
maxPages integer 50 Max source pages / page-level dataset rows
concurrency integer 5 Parallel requests. Use reasonable values on public sites
checkExternal boolean true Also check external-domain links
timeoutMs integer 10000 Request timeout in milliseconds
delivery string dataset dataset or webhook
webhookUrl string — POST when delivery is webhook; not during dryRun
dryRun boolean false Local crawl only; no dataset/webhook. Store example run sets true

Published Store example run input (dry run — schema default dryRun is false):

{
  "startUrls": ["https://example.com"],
  "maxDepth": 1,
  "maxPages": 5,
  "concurrency": 3,
  "checkExternal": true,
  "timeoutMs": 10000,
  "delivery": "dataset",
  "dryRun": true
}

Published Store Quickstart (live crawl; keep the first run small):

{
  "startUrls": ["https://example.com"],
  "maxDepth": 1,
  "maxPages": 100,
  "delivery": "dataset",
  "dryRun": false
}

Run Broken Link Checker on Apify

What does a result contain?

The published README field list is rowType, sourcePage, targetUrl, statusCode, anchorText, linkType, redirectChain, severity, and fetchedAt. The published sample wraps meta plus rows[] with rowType broken_link. The live schema description says each dataset row is a crawled source page including broken-link details; that may differ from the per-link sample shape. Treat the sample as a README illustration, not a live coverage guarantee.

How do dataset, webhook, and dry-run delivery work?

delivery defaults to dataset. Dataset output is written first; webhook POSTs after PPE output succeeds when delivery is webhook, webhookUrl is set, and dryRun is false. dryRun true runs a local crawl only and skips dataset writes and webhook. The default dataset is the billable surface. README: healthy/no-issue checks should avoid default dataset charges; dry runs, missing-key warnings, and unchanged polls should not write payable default-dataset rows.

How is Broken Link Checker priced?

Billing is pay per event. The live Store card is from $8.00 / 1,000 results. Live Store events are Actor Start at $0.001 (charged when the Actor starts; number of events depends on Actor memory, one event per GB, minimum one event) and result at $0.008 (single result in the default dataset). The README pricing section matches those figures: $0.001 actor start and $0.008 per broken-link finding row. Older README figures of actor-start $0.01 / dataset-item $0.005 are stale. Platform usage is listed as included.

from $8.00 per 1,000 results ($0.001 Actor Start + $0.008 per default-dataset result)

See Broken Link Checker pricing on Apify

Limits to keep in mind

Published README next-report Store listings (no local landings): Site QA Broken Link Report Scraper (taroyamada/site-qa-broken-link-report-scraper), plus indexability/AI-crawler report Actors named in that README.

Open Broken Link Checker on Apify

Related pages