Find Wayback snapshots for broken URLs

Use this page when you already have dead, 404, or deleted URLs and need to know whether the Internet Archive cached them. Wayback Machine Bulk Checker is the existing Store Actor for that read-only lookup. It is not a new Actor, not live HTTP status QA, and not a tool that saves new snapshots.

Live Store price: $0.008 per successfully delivered Wayback Machine lookup row ($8.00 / 1,000 wayback archive results). Cheapest first paid path: 1 URL = $0.008 if one lookup row is delivered. There is no Actor start event on the current Store pricing tab. Platform usage is listed as Included. Do not use older README Cost lines that listed actor-start $0.01 and dataset-item $0.003.

Try the cheapest first paid Wayback lookup on Apify Input

Open taroyamada/wayback-machine-checker on the Apify Store · Wayback Machine Bulk Checker landing

What a 404-recovery row reports

This page uses one README output-table set: url, archived, closestSnapshotUrl, and closestSnapshotDate. The same table also lists totalSnapshots, firstSnapshotDate, and lastSnapshotDate; treat those and original HTTP status as README-claimed columns, not as a published output_schema.json. Availability is not guaranteed per URL. The product landing documents the README sample, which uses different names.

The published README use cases include SEO recovery of deleted pages and cross-referencing dead URLs with historical cache data. For mapping a pre-migration date, see Wayback site migration snapshots.

Tiny first-paid input

The published Store Quickstart verifies a few archived URLs. Live required field is urls (maximum 500). Leave closest empty for the latest snapshot. Keep concurrency low (schema default 3, maximum 5). maxChargeUsd defaults to 1; results beyond the cap are kept as no-charge limit_reached rows. Do not send README aliases such as checkAvailability, includeCdx, maxSnapshotsPerUrl, or compareToLatestLive.

{
  "urls": [
    "https://example.com/old-page"
  ],
  "concurrency": 3,
  "maxChargeUsd": 1,
  "delivery": "dataset",
  "dryRun": false
}

Store templates named in the README are Quickstart (verify 3 archived URLs), Portfolio Archive Check (bulk verification, up to 500 URLs), and 404 Recovery after a broken-link crawl. If one lookup row is delivered, documented PPE is $0.008. If the 3-URL Quickstart delivers three rows, 3 × $0.008 = $0.024 at the same rate. Find live broken URLs with Bulk URL Status Checker, then look up archive availability here. This Actor cannot Save Page Now.

Run this Wayback 404 lookup on Apify

Can Wayback Machine Bulk Checker recover content from broken URLs?

It can report whether the Internet Archive cached those URLs and, when present, return closestSnapshotUrl and closestSnapshotDate. Those are the published README output table names (with url and archived). Required input is urls (maximum 500). Leave closest empty for the latest snapshot. The Actor does not fetch cached HTML, does not extract article bodies, and cannot create new snapshots. Coverage is not guaranteed per URL. Snapshots exist back to 1996 only where the Archive crawled the page. A URL may be unavailable because it was never archived, blocked by robots.txt, or removed by a removal request.

What is the cheapest first paid input for 404 archive lookup?

The cheapest first paid path is one urls item. If one lookup row is delivered, live PPE is $0.008. There is no Actor start event. The published Store Quickstart verifies 3 archived URLs (3 × $0.008 = $0.024 if all three are delivered). Live schema names are urls (required), optional closest in YYYYMMDD form, concurrency (default 3), maxChargeUsd (default 1), delivery dataset, and dryRun false. See Wayback Machine Bulk Checker pricing.

Open Wayback Machine Bulk Checker Input on Apify

Is this the same as live HTTP status checking?

No. Bulk URL Status Checker checks live HTTP status, redirects, and timing on a known list. This Actor queries archive.org/wayback/available. The README 404 Recovery template is for after a broken-link crawl: you already have dead URLs, then you look up archive availability. This Actor cannot Save Page Now. To save new pages, the README points to archive.org’s /save/ endpoint.

Open Wayback Machine Bulk Checker Input on Apify

Related pages