Broken links
Crawl start URLs for broken links, 404s, and redirect chains
Broken Link Checker starts from public pages you supply, follows internal links up to a bounded depth, and reports broken links with HTTP status codes, redirect chains, anchor text, and internal/external class. It is crawl-based link QA — not a known URL-list status checker, not sitemap XML discovery, and not a robots.txt AI-crawler audit.
from $8.00 / 1,000 results ($0.001 Actor Start + $0.008 per default-dataset result)
Open Broken Link Checker on Apify
HTML crawl for broken links, not a known URL list
Use this page when you have a start URL and need to discover broken links from source HTML. Use Bulk URL Status Checker when you already have the URL list. Use Sitemap Scraper & Analyzer to inventory sitemap.xml URLs. Use robots.txt Parser & AI Crawler Block Checker for GPTBot/ClaudeBot allow/disallow rules.
| This Actor | Bulk URL Status Checker | Sitemap Scraper & Analyzer | |
|---|---|---|---|
| Intent | Discover broken links from crawled HTML | HTTP status QA on a known URL list | Discover URLs from nested sitemap.xml |
| Input | startUrls (max 10) |
urls (max 1000) |
sitemapUrls |
| What it reads | Start pages plus internal hops up to maxDepth |
Live HTTP(S) for URLs you already have; no crawl | sitemap.xml and sitemap indexes |
| Primary output | Broken-link findings: status, redirects, anchor, severity | Status code, final URL, redirect chain, timing | URL inventory, lastmod, optional HEAD 200 |
| Not this job | Status-only pass over a fixed list | HTML crawl to find links | Anchor-text crawl of HTML pages |
Store ID: taroyamada/broken-link-checker. Crawl only public sites you control or are authorized to check, with reasonable limits.
Use cases
- Single-page or small-site audits from a start URL.
- Outbound-link checks with
checkExternaltrue (schema default). - Technical SEO and site-maintenance QA of 404s and redirect chains.
- Webhook handoff after dataset/PPE output succeeds.
- Correlate later with robots.txt, sitemap, and known-list status Actors on this site.
How is Broken Link Checker different from Bulk URL Status Checker and Sitemap Scraper?
This Actor crawls from start URLs you supply, follows internal links up to maxDepth, and reports broken links, HTTP status codes, and redirect chains. Required input is startUrls (maximum 10). Store ID taroyamada/broken-link-checker. Bulk URL Status Checker checks a known URL list (urls, maximum 1000) for HTTP status, redirect chains, and timing; it does not crawl pages. Sitemap Scraper & Analyzer parses sitemap.xml into URL inventories and optionally HEADs discovered URLs; it does not crawl HTML for anchors. robots.txt Parser & AI Crawler Block Checker audits AI crawler allow/disallow rules, not page links. Use this Actor to discover broken links from source HTML. Use Bulk URL Status Checker when you already have the URL list.
What input is required?
startUrls is required: public start pages, maximum 10. Schema prefill is https://example.com. additionalProperties is false. README examples that send includeExternalLinks, pages, maxLinksPerPage, or externalOnly are not live schema fields. Do not send them.
| Field | Type | Default | Notes |
|---|---|---|---|
startUrls |
string[] | required | Max 10. Prefill https://example.com. For a fixed list, use Bulk URL Status Checker |
maxDepth |
integer | 2 | Internal-link hops; maximum 5 |
maxPages |
integer | 50 | Max source pages / page-level dataset rows |
concurrency |
integer | 5 | Parallel requests. Use reasonable values on public sites |
checkExternal |
boolean | true | Also check external-domain links |
timeoutMs |
integer | 10000 | Request timeout in milliseconds |
delivery |
string | dataset |
dataset or webhook |
webhookUrl |
string | — | POST when delivery is webhook; not during dryRun |
dryRun |
boolean | false | Local crawl only; no dataset/webhook. Store example run sets true |
Published Store example run input (dry run — schema default dryRun is false):
{
"startUrls": ["https://example.com"],
"maxDepth": 1,
"maxPages": 5,
"concurrency": 3,
"checkExternal": true,
"timeoutMs": 10000,
"delivery": "dataset",
"dryRun": true
}
Published Store Quickstart (live crawl; keep the first run small):
{
"startUrls": ["https://example.com"],
"maxDepth": 1,
"maxPages": 100,
"delivery": "dataset",
"dryRun": false
}
Run Broken Link Checker on Apify
What does a result contain?
The published README field list is rowType, sourcePage, targetUrl, statusCode, anchorText, linkType, redirectChain, severity, and fetchedAt. The published sample wraps meta plus rows[] with rowType broken_link. The live schema description says each dataset row is a crawled source page including broken-link details; that may differ from the per-link sample shape. Treat the sample as a README illustration, not a live coverage guarantee.
How do dataset, webhook, and dry-run delivery work?
delivery defaults to dataset. Dataset output is written first; webhook POSTs after PPE output succeeds when delivery is webhook, webhookUrl is set, and dryRun is false. dryRun true runs a local crawl only and skips dataset writes and webhook. The default dataset is the billable surface. README: healthy/no-issue checks should avoid default dataset charges; dry runs, missing-key warnings, and unchanged polls should not write payable default-dataset rows.
How is Broken Link Checker priced?
Billing is pay per event. The live Store card is from $8.00 / 1,000 results. Live Store events are Actor Start at $0.001 (charged when the Actor starts; number of events depends on Actor memory, one event per GB, minimum one event) and result at $0.008 (single result in the default dataset). The README pricing section matches those figures: $0.001 actor start and $0.008 per broken-link finding row. Older README figures of actor-start $0.01 / dataset-item $0.005 are stale. Platform usage is listed as included.
from $8.00 per 1,000 results ($0.001 Actor Start + $0.008 per default-dataset result)
See Broken Link Checker pricing on Apify
Limits to keep in mind
startUrlsmaximum 10.maxDepthmaximum 5 (default 2). DefaultmaxPages50.- Live schema names do not include
includeExternalLinks,pages,maxLinksPerPage, orexternalOnly. - Do not aggressive-crawl sites you do not control. Technical SEO/QA/maintenance only.
- For a known URL list, use Bulk URL Status Checker.
- Store example run is a dry run; schema default
dryRunis false.
Published README next-report Store listings (no local landings): Site QA Broken Link Report Scraper (taroyamada/site-qa-broken-link-report-scraper), plus indexability/AI-crawler report Actors named in that README.
Open Broken Link Checker on Apify
Related pages
- Bulk URL Status Checker — known URL list HTTP status, not an HTML crawl
- Sitemap Scraper & Analyzer — sitemap.xml inventories, not anchor crawls
- robots.txt Parser & AI Crawler Block Checker — robots.txt AI policies, not 404 crawls
- RDAP Domain Expiry & Whois Scraper — RDAP registration records, not link crawls
- Wayback Machine Bulk Checker — Internet Archive availability, not live 404 crawls
- Short URL Resolver & Scraper — expand short links, not site crawls
- Meta Tag & OpenGraph Scraper — head metadata on known URLs
- Website Accessibility Checker — WCAG 2.1 violations, not HTTP 404s.
- Website Content Extractor — cleaned page body
- Security Headers Checker — OWASP header grades on a URL list, not HTML crawl for 404s
- Site Governance Monitor — robots/sitemap/schema drift, not broken-link crawl
- Site QA Indexability AI Crawler Report — robots.txt and llms.txt signals.
- Site QA Broken Link Report — billed URL-health report pack on supplied pages.
- Tools