Website content extractor for docs and help-center RAG
Use this page when the job is ingesting public docs, pricing, policy, and help-center page bodies into RAG. Website Content Extractor is the existing Store Actor for that cleanup step. It is not a new Actor, not a full-site crawler, and not a buyer-facing Site QA or RAG report. Pass HTTP(S) URLs you own or are authorized to fetch.
Live Store price: from $9.00 / 1,000 results (result at $0.009 per useful content row). Actor Start $0.001. Cheapest first paid path: 1 docs URL ≈ $0.010. Leave report-only fields off this extractor.
Try the cheapest first paid Website Content Extractor run on Apify Input
Open taroyamada/website-content-extractor on the Apify Store · Website Content Extractor landing
What RAG gets from each web_content row
The published README field list is rowType, url, title, markdown, text, wordCount, metadata, sourceUrl, and fetchedAt. Sample rowType is web_content. Markdown is the default outputFormat and the live schema recommends it for first-run proof and downstream reuse. Raw content rows remain source context for content QA and RAG ingestion — they are not a Site QA or RAG report.
For news, blog, newsroom, or press URLs, use Article Content Extractor and the news and blog RAG page. For head Open Graph, robots, and JSON-LD audits, use Meta Tag & OpenGraph Scraper. The four existing extractors sit on the Content Intelligence pack hub.
Tiny first-paid input
The published Store Quickstart is one public docs URL, outputFormat markdown, delivery dataset, and dryRun false. urls is required (maximum 200). Omit generateReport and emitExport.
{
"urls": [
"https://example.com/docs"
],
"outputFormat": "markdown",
"delivery": "dataset",
"dryRun": false
}
Replace the placeholder with one public docs or help-center URL you are authorized to fetch. If extraction succeeds, live PPE is Actor Start $0.001 plus one useful content row at $0.009 = $0.010. Failed or no-content rows stay out of the default dataset. The live Store card is from $9.00 / 1,000 results.
Run this website extractor input on Apify
Can this website extractor feed a docs and help-center RAG corpus?
Yes, as a bounded source step. It fetches public docs, pricing, product, policy, and help-center URLs you already have and returns structured web_content rows: title, markdown or text, wordCount, metadata, sourceUrl, and fetchedAt. outputFormat defaults to markdown. It is not a search surface, a full-site crawler, or a buyer-facing Site QA or RAG report Actor. You supply the URL list (maximum 200).
Why not send news or blog URLs to Website Content Extractor?
Those URLs are article-shaped. Article Content Extractor is for news, blog, newsroom, and press pages when you need headline, byline, publish date, article body, excerpt, and hero image (maximum 300). This Actor is for broader site pages that are not article-shaped (maximum 200). For news and blog RAG, see article extractor for news and blog RAG.
What is the cheapest first paid website extractor input for RAG?
The published Store Quickstart is one urls item, outputFormat markdown, delivery dataset, and dryRun false. Leave generateReport and emitExport off. Live PPE is $0.009 per useful content row plus Actor Start $0.001. If those two events charge, documented PPE is $0.010.
Open Website Content Extractor Input on Apify
Related pages
- Website Content Extractor — product landing and field list
- Content Intelligence pack — RSS, Google News, article, and website extractors
- Article extractor for news and blog RAG — news, blog, and press bodies
- Article Content Extractor
- Meta Tag & OpenGraph Scraper — head metadata, not page-body markdown
- YouTube transcript extractor for RAG — captions, not website HTML
- Tools