Look up Wayback snapshots closest to a site migration date
Use this page when you already have legacy URLs and need the Internet Archive snapshot nearest a cutover date. Wayback Machine Bulk Checker is the existing Store Actor. It is not a new Actor, not a sitemap crawler, and not a CDX full-history dump. Pass known URLs; set closest in YYYYMMDD form.
Live Store price: $0.008 per successfully delivered Wayback Machine lookup row ($8.00 / 1,000 wayback archive results). Cheapest first paid path: 1 URL = $0.008 if one lookup row is delivered. There is no Actor start event on the current Store pricing tab. Platform usage is listed as Included. Do not use older README Cost lines that listed actor-start $0.01 and dataset-item $0.003.
Try the cheapest first paid Wayback lookup on Apify Input
Open taroyamada/wayback-machine-checker on the Apify Store · Wayback Machine Bulk Checker landing
What a closest-snapshot row reports
The live schema description for closest is: Find snapshot closest to this date (YYYYMMDD format). Leave empty for latest. This page uses one README output-table set: url, archived, closestSnapshotUrl, and closestSnapshotDate (YYYYMMDDhhmmss). The same table also lists totalSnapshots, firstSnapshotDate, and lastSnapshotDate; treat those and original HTTP status as README-claimed columns. There is no published output_schema.json. The product landing documents the README sample, which uses different names.
The published README positions the Actor for mapping legacy site structures, auditing historical page versions during website migrations or domain changes, and auditing domain history before acquisitions. For broken-URL recovery without a target date, see Wayback 404 recovery. Discover URL inventories from nested XML with Sitemap Scraper & Analyzer, then look up those URLs here.
Tiny first-paid input
Start with the Store Quickstart, then add closest. Schema-valid closest is YYYYMMDD (for example 20200101), not YYYY-MM-DD. Keep concurrency low. Use dryRun: true before larger portfolio runs. Store Portfolio Archive Check covers bulk verification up to 500 URLs.
{
"urls": [
"https://example.com/old-article"
],
"closest": "20200101",
"concurrency": 3,
"maxChargeUsd": 1,
"delivery": "dataset",
"dryRun": false
}
No Internet Archive API key is required. The Actor uses archive.org/wayback/available. It cannot Save Page Now. Do not send checkAvailability, includeCdx, maxSnapshotsPerUrl, or compareToLatestLive — those names are not in the live schema, and additionalProperties is false. If one lookup row is delivered, documented PPE is $0.008.
Run this closest-snapshot input on Apify
Can Wayback Machine Bulk Checker map URLs from a site migration?
Yes, as a closest-snapshot lookup on URLs you already have (maximum 500). Set closest in YYYYMMDD form to find the snapshot nearest that date; leave it empty for the latest. The published README positions this for mapping legacy site structures and auditing historical page versions during migrations or domain changes. It does not crawl the old site, parse sitemap.xml, or dump CDX full history.
What live input finds the snapshot closest to a migration date?
urls is required. closest is YYYYMMDD (for example 20200101), not YYYY-MM-DD. concurrency defaults to 3 (maximum 5). maxChargeUsd defaults to 1. Live PPE is $0.008 per successfully delivered lookup row. The cheapest first paid path is one URL. Do not send README aliases such as checkAvailability, includeCdx, maxSnapshotsPerUrl, or compareToLatestLive. See Wayback Machine Bulk Checker pricing.
Open Wayback Machine Bulk Checker Input on Apify
Does a migration lookup return archived HTML for every URL?
No. The Actor returns structured availability rows using the README output table names url, archived, closestSnapshotUrl, and closestSnapshotDate. It does not fetch cached HTML. Snapshots exist only where the Archive crawled the page. A URL may be unavailable because it was never archived, blocked by robots.txt, or removed by a removal request. Use Website Content Extractor for live docs and product pages, and Article Content Extractor for live news and blog pages.
Open Wayback Machine Bulk Checker Input on Apify
Related pages
- Wayback Machine Bulk Checker — product landing and field list
- Wayback 404 recovery
- Sitemap Scraper & Analyzer — discover URLs from nested XML sitemaps
- Bulk URL Status Checker — live HTTP status on a known list
- Website Content Extractor
- Article Content Extractor
- Tools