Article extractor for news and blog RAG

Use this page when the job is ingesting public news and blog article bodies into RAG. Article Content Extractor is the existing Store Actor for that cleanup step. It is not a new Actor, not Google News search, and not a full-site crawler. Pass article-shaped HTTP(S) URLs you own or are authorized to fetch.

Live Store price: $0.008 per useful article row (apify-default-dataset-item). Actor Start $0.00005. Leave the $2.50 audit report and $5.00 batch export off unless you need those value events.

Open Article Content Extractor on Apify · Cheapest first paid run · URL to Markdown for RAG

What RAG gets from each article row

The published README field list is rowType, url, headline, author, publishedAt, articleText, excerpt, heroImage, and sourceUrl. Sample rowType is article. Markdown is the default outputFormat and the live schema recommends it for first-run proof and LLM ingestion. Raw article rows remain source context for research, content QA, and RAG ingestion — they are not a Site QA or RAG report.

Discover URLs first with Google News Scraper or RSS & Atom Feed Extractor, then send news, blog, newsroom, and press links here. Route docs, pricing, policy, and product pages to Website Content Extractor. The four existing Actors sit on the Content Intelligence pack hub.

Tiny first-paid input

The published Store Quickstart is one public article URL, includeImages false, delivery dataset, and dryRun false. urls is required (maximum 300). Omit generateReport and emitExport.

{
  "urls": [
    "https://example.com/news/example"
  ],
  "includeImages": false,
  "delivery": "dataset",
  "dryRun": false
}

Replace the placeholder with one public news or blog URL. If extraction succeeds, live PPE is Actor Start $0.00005 plus one useful article row at $0.008. Failed or no-content rows stay out of the default dataset. For the PPE table, see Cheapest first paid run for Article Extractor.

Run this article extractor input on Apify

Can this article extractor feed a news and blog RAG corpus?

Yes, as a bounded source step. It fetches public news, blog, newsroom, and press URLs you already have and returns structured article rows: headline, author, publishedAt, articleText, excerpt, heroImage, and sourceUrl. outputFormat defaults to markdown. It is not a search surface, a full-site crawler, or a buyer-facing RAG report Actor. You supply the URL list (maximum 300).

Why not send Google News or RSS output straight into RAG?

Those Actors are discovery layers. Google News Scraper returns headlines and article URLs from queries, without full article text. RSS & Atom Feed Extractor parses known-publisher feed XML and does not crawl page HTML. For RAG you still need the article body. Discover URLs first, then send news, blog, and press links here.

What is the cheapest first paid article extractor input for RAG?

The published Store Quickstart is one urls item, includeImages false, delivery dataset, and dryRun false. Leave generateReport and emitExport off. Live PPE is $0.008 per useful article row plus Actor Start $0.00005.

Open Article Content Extractor on Apify

Related pages