Article Content Extractor vs Website Content Extractor for RAG

Both Actors clean public HTML you already have into text you can ingest. They are different jobs, not a ranking of quality. Article Content Extractor Apify is for article-shaped news, blog, newsroom, and press URLs. Website Content Extractor is for broader docs, product, pricing, policy, and help-center pages. Neither is a search surface, a full-site crawler, or a buyer-facing RAG report. SAM.gov Scraper Apify is a bid-alert queue, not page HTML cleanup.

Open Article Content Extractor Apify on the Store · Article Content Extractor Apify landing · Open Website Content Extractor Input on Apify

Article-shaped URLs versus site-page markdown

Use this page when the job is choosing which extractor feeds RAG. Discover article URLs first with Google News Scraper or RSS & Atom Feed Extractor. Use Meta Tag & OpenGraph Scraper when the job is head metadata, not page body.

Article Content Extractor Apify Website Content Extractor
Intent Parse news, blog, newsroom, and press article pages Clean docs, product, pricing, policy, and help-center HTML
You have Article-shaped page URLs Broader site-page URLs
Max URLs 300 per run 200 per run
Output format markdown (default) or text — no HTML emit markdown (default), text, or html
Sample row article web_content
RAG-facing fields Headline, byline, publishedAt, articleText, excerpt, heroImage Title, markdown/text, wordCount, optional metadata
Optional extras includeImages, generateReport, emitExport includeMetadata, maxChargeUsd
Published price (sister LPs) Useful article row $0.008 ($8.00 / 1,000) From $9.00 / 1,000 results ($0.009 per row)

None of these Actors is a ranking tool, a full-site crawler, or a claim of publisher endorsement. Fetch public HTTP(S) URLs you own or are authorized to audit. Do not imply content ownership transfer.

When Article Content Extractor is the RAG source

Use Article Content Extractor when the URLs are article-shaped and you need a clean headline, byline, and body for RAG:

Leave generateReport and emitExport false for the first paid run. Those value events are $2.50 and $5.00 on the Article LP when you request them. The Store Quickstart also sets includeImages false for a low-cost first run.

{
  "urls": [
    "https://example.com/news/example"
  ],
  "outputFormat": "markdown",
  "generateReport": false,
  "emitExport": false
}

If one Actor Start event and one useful article row charge, documented PPE on the Article LP is $0.00005 + $0.008 = $0.00805. Failed or no-content rows stay out of the default dataset. See the cheapest first paid article run and news and blog RAG.

Run an article extraction on Apify

When Website Content Extractor is the RAG source

Use Website Content Extractor when the URLs are broader site pages that are not article-shaped. It is a docs and help-center crawl-style cleaner: you still supply the URL list; it does not crawl a full site on its own.

Do not send generateReport or emitExport on this extractor. Those names belong on a linked Site QA or RAG report Actor. Cheapest first paid path on the Website LP is one urls item.

{
  "urls": [
    "https://docs.apify.com/platform/actors"
  ],
  "outputFormat": "markdown",
  "delivery": "dataset",
  "dryRun": false
}

If that run charges one Actor Start event and writes one useful content row, documented PPE on the Website LP is $0.001 + $0.009 = $0.010. The live Store card is from $9.00 / 1,000 results. See docs and help-center RAG. For current Website rates, see the Store pricing card.

Run this 1-URL Website input on Apify

What each source does not do

A practical split

You have Start with
News, blog, newsroom, or press URLs, and you need headline/byline/body Markdown Article Content Extractor
Docs, pricing, policy, product, or help-center URLs, and you need page Markdown Website Content Extractor
Search queries, and you still need article URLs Google News Scraper, then Article cleanup
Publisher feed URLs, and you still need item links RSS & Atom Feed Extractor, then Article cleanup

This is a source choice, not a ranking of tools. Article remains the right read for article-shaped pages. Website is the right read when the page is a site/docs body, not a news article. The four extractors sit on the Content Intelligence pack hub.

Open Article Content Extractor on Apify · Open Website Content Extractor Input on Apify

When should I use Article Content Extractor for RAG?

Use Article Content Extractor when the URLs are article-shaped: public news, blog, newsroom, or press pages you already have. It returns headline, byline, publishedAt, articleText, excerpt, and heroImage as markdown (default) or text. Maximum 300 URLs per run. It does not emit HTML. Leave generateReport and emitExport false on the first paid run. Live PPE is $0.008 per useful article row.

When should I use Website Content Extractor for RAG instead?

Use Website Content Extractor when the URLs are broader site pages that are not article-shaped: docs, product, pricing, policy, and help-center HTML. It returns cleaned markdown (default), text, or HTML plus title, optional metadata, wordCount, sourceUrl, and fetchedAt. Sample rowType is web_content. Maximum 200 URLs per run. The live Store card is from $9.00 / 1,000 results.

How do the first paid inputs differ?

Article first paid is one news or blog URL, outputFormat markdown, generateReport false, emitExport false. Website first paid is one docs URL, outputFormat markdown, delivery dataset, dryRun false. Do not send generateReport or emitExport on Website Content Extractor; those belong on a linked report Actor. Article useful-row PPE is $0.008. Website first paid is documented as 1 docs URL ≈ $0.010 (Actor Start $0.001 plus result $0.009).

Does Article Content Extractor emit HTML?

No. The published outputFormat enum is markdown (default) or text. Website Content Extractor is the sibling that can return html as well as markdown or text.

Does Website Content Extractor emit a Site QA or RAG report?

No. Website Content Extractor returns bounded source rows. Linked report actors have separate report and export events. Article Content Extractor can emit article-content-audit-report ($2.50) and article-batch-export ($5.00) when those flags are true; the first paid path leaves both false.

Should news, blog, newsroom, or press URLs go to Website Content Extractor?

No. Route article-shaped news, blog, newsroom, and press URLs to Article Content Extractor when you need headline, byline, publish date, article body, excerpt, and hero image. Website Content Extractor is for broader site pages that are not article-shaped.

How to run the first paid path

  1. Match the URL shape: article pages to Article Content Extractor; docs and site pages to Website Content Extractor.
  2. Open the matching landing, then the Actor on the Apify Store.
  3. Paste the first-paid JSON above. Leave Article generateReport / emitExport false. Do not add those flags on Website.
  4. Review the default dataset, then raise the URL list after the row is useful.

Article: useful article row $0.008 ($8.00 / 1,000). Website: from $9.00 / 1,000 results; 1 docs URL ≈ $0.010.

Open Article Content Extractor on Apify · Open Website Content Extractor Input on Apify

How is Article Content Extractor priced versus Website Content Extractor?

Both are pay-per-event. Article live events include useful article row $0.008 ($8.00 / 1,000), article-content-audit-report $2.50, article-batch-export $5.00, and Actor Start $0.00005. Website live Store card is from $9.00 / 1,000 results: Actor Start $0.001 and result $0.009 per default-dataset item. Check the Store pricing cards for current rates.

Article pricing on Apify · Website pricing on Apify

Do generateReport and emitExport belong on Website Content Extractor?

No. Those are Article Content Extractor value-event flags, or fields for a linked Site QA or RAG report Actor. Website Content Extractor additionalProperties is false, so unpublished report aliases are not accepted. First paid on Article leaves both flags false.

What fields does an article row include versus a web_content row?

Article README fields include rowType, url, headline, author, publishedAt, articleText, excerpt, heroImage, and sourceUrl. Sample rowType is article. Website README fields include rowType, url, title, markdown, text, wordCount, metadata, sourceUrl, and fetchedAt. Sample rowType is web_content.

Does either Actor crawl a full site or rank pages?

No. You supply the URL list. Article maximum is 300 URLs; Website maximum is 200. Both are bounded source extractors for research, content QA, and RAG ingestion — not ranking tools, full-site crawlers, or claims of publisher endorsement.

Related pages

Public HTTP(S) pages you own or are authorized to fetch. Not a ranking tool or a full-site crawler. Not affiliated with upstream publishers.