RAG readiness

Public website URLs into a RAG readiness audit

This Actor turns public website URLs into a decision-ready RAG readiness audit with coverage, chunking risk, retrieval cleanup actions, and source URLs. No user API key required.

Not a general website-to-markdown dump (use Website Content Extractor). Not article boilerplate cleanup (use Article Content Extractor). Not a JSON-LD validator.

$9.00 per Website RAG snapshot report (website_rag_snapshot_report); $29 per Website RAG readiness report. $0.001 Actor Start.

Open Website RAG Readiness Audit on Apify

Website RAG Readiness Audit, not a neighboring Actor

Use this page for this Actor’s job. Use Website Content Extractor for One-shot page-to-markdown; Use Article Content Extractor for News/blog HTML cleanup; Use Structured Data Scraper & Validator for JSON-LD / Microdata audits.

This ActorWebsite Content ExtractorArticle Content ExtractorStructured Data Scraper & Validator
IntentWebsite RAG Readiness AuditOne-shot page-to-markdownNews/blog HTML cleanupJSON-LD / Microdata audits
Primary inputurlsurlsurlsurls
What it readsUser-supplied public website URLsPage HTMLArticle HTMLPage head/body schema
Primary outputSnapshot and readiness report eventsMarkdown/textHeadline/byline/bodySchema validation rows
Not this jobNot a general website-to-markdown dump (use Website Content Extractor). Not article boilerplate cleanup (use Article Content Extractor). Not a JSON-LD validator.Not recurring pricing/legal diffs or RAG audits.Not landing-page CRO reports.Not RAG readiness reports.

Store ID: taroyamada/website-rag-readiness-audit. Respect source terms, robots.txt, and rate limits.

Use cases

How is Website RAG Readiness Audit different from Website Content Extractor and Article Content Extractor?

Website RAG Readiness Audit (taroyamada/website-rag-readiness-audit): This Actor turns public website URLs into a decision-ready RAG readiness audit with coverage, chunking risk, retrieval cleanup actions, and source URLs. No user API key required. Not a general website-to-markdown dump (use Website Content Extractor). Not article boilerplate cleanup (use Article Content Extractor). Not a JSON-LD validator. Website Content Extractor is for One-shot page-to-markdown (input urls; Page HTML; Markdown/text). Not recurring pricing/legal diffs or RAG audits. Article Content Extractor is for News/blog HTML cleanup (input urls; Article HTML; Headline/byline/body). Not landing-page CRO reports. Structured Data Scraper & Validator is for JSON-LD / Microdata audits (input urls; Page head/body schema; Schema validation rows). Not RAG readiness reports.

What input is required?

Live required fields: none listed (see live fields). Published exampleRunInput is shown below.

Field Type Default Notes
urls any[] empty Public URLs. Public website pages to audit. Use docs, pricing, help, blog, policy, or product pages. Login/paywall/private dashboard URLs are skipped as no-charge.
domain string empty Domain. Optional public domain. If provided without URLs, the Actor audits the homepage only in v1.
reportTier string enum snapshot Report tier. Choose the public launch tier. Watch summaries are proof-gated and not selectable publicly. enum: snapshot, readiness
seedQuestions any[] empty Seed questions. Optional buyer questions the RAG corpus should answer. Used for action recommendations only.
maxPages integer 3 Max pages. Maximum public pages to fetch in this run. maxChargeUsd is still checked before charging. minimum=1 maximum=25
maxReports integer 1 Max reports. Maximum report groups to charge. Usually 1 for this Actor. minimum=1 maximum=5
maxChargeUsd number 9 Max charge USD. Hard safety cap. If the selected report would exceed this cap, the Actor returns a no-charge limit_reached summary. minimum=0 maximum=100
demoMode boolean false Demo mode. Return a no-charge sample preview without fetching external pages.
dryRun boolean false Dry run. Return a no-charge previewReport and nextRunInput.
sourceDatasetId string empty Advanced source dataset ID. Advanced only. Optional future hook for prepared URL rows; not needed for first run.

Published Store example run input (omitted fields take schema defaults):

{
  "reportTier": "snapshot",
  "maxPages": 3,
  "maxReports": 1,
  "maxChargeUsd": 9,
  "demoMode": false,
  "dryRun": false
}

Run Website RAG Readiness Audit on Apify

How do dataset, webhook, and dry-run delivery work?

Dataset output is the billable surface when rows are written. See live PPE for which events bill. dryRun true validates or samples without the usual dataset/webhook side effects described on the Store schema. Unchanged runs that write zero default-dataset rows typically charge $0.00 on current PPE.

What does a result contain?

Published README Output Example / Sample Output JSON. Treat README samples as illustrations, not a live coverage guarantee. There is no published output JSON schema on the Store page.

How is Website RAG Readiness Audit priced?

Billing is pay per event. The live Store card is $9.00 per Website RAG snapshot report (website_rag_snapshot_report); $29 per Website RAG readiness report. $0.001 Actor Start. Current PPE:

Event Price Emitted when
apify-actor-start (Actor Start) $0.001 Charged when the Actor starts running. Number of events charged depends on Actor memory (one event per GB, minimum one event).
website_rag_snapshot_report (Website RAG snapshot report) $9.00 One public website RAG snapshot with coverage score, warnings, source URLs, and prioritized cleanup actions.
website_rag_readiness_report (Website RAG readiness report) $29 One deeper website RAG readiness report with chunking risks, retrieval QA actions, and implementation priorities.

$9.00 per Website RAG snapshot report (website_rag_snapshot_report); $29 per Website RAG readiness report. $0.001 Actor Start.

See Website RAG Readiness Audit pricing on Apify

Limits to keep in mind

Open Website RAG Readiness Audit on Apify

Related pages