Content Intelligence pack
This hub groups four existing Actors that discover URLs or clean page HTML. It is not a new Actor and it has no Store listing of its own. Pick the landing that matches the job: known-publisher feeds, Google News queries, article-shaped pages, or broader site pages.
- RSS & Atom Feed Extractor — parse public RSS/Atom XML you already have into item URL lists
- Google News Scraper — query Google News RSS for headlines and article URLs
- Article Content Extractor — extract article bodies from news, blog, newsroom, and press URLs
- Website Content Extractor — clean docs, pricing, product, policy, and help-center pages
Discovery versus HTML cleanup
RSS and Google News are discovery layers. They return links (and feed or RSS metadata). They do not replace an HTML extractor. Article and Website extractors fetch live page HTML you already have. They are not search surfaces.
| RSS & Atom Feed Extractor | Google News Scraper | Article Content Extractor | Website Content Extractor | |
|---|---|---|---|---|
| Intent | Discover fresh URLs from known publisher and blog feeds | Discover current article URLs by query | Clean discovered article, news, and blog pages | Clean discovered docs, pricing, policy, or product pages |
| You have | Public RSS or Atom feed URLs | News search queries, not feed URLs | Article-shaped page URLs | Site page URLs that are not article-shaped |
| Input | feedUrls (max 50) |
queries (max 50) |
Page URLs (max 300) | Page URLs (max 200) |
| What it reads | Feed XML — no site crawl | Google News RSS for those queries | Live article HTML | Live page HTML |
| Primary output | Feed items: title, link, dates, summary, encoded content when present | Headline, article URL, publisher, timestamps, snippet, originating query | Article body, byline, date, excerpt, hero image | Cleaned markdown or text plus page metadata |
| HTML cleanup | No. Feed discovery only | No. Clean pages afterward with an extractor | Yes — bounded article extraction | Yes — broader site-page extraction |
| Sample row / fields | source, title, link, pubDate, pubDateISO, description, content, categories |
title, link, source, pubDate, pubDateISO, description, query — no full-article text field |
Sample rowType is article |
Sample rowType is web_content |
| Published Store price | $1.00 / 1,000 RSS or Atom items ($0.001 per successfully parsed public feed item) | $3.00 / 1,000 Google News RSS articles ($0.003 per successfully parsed public RSS article row) | $0.008 per useful article row; $2.50 per article-content-audit-report; $5.00 per article-batch-export |
$0.001 Actor start and $0.009 per useful content row ($9.00 / 1,000 results) |
Prices above are the published Store figures already quoted on those landings. This hub does not add a pack price.
What is the Content Intelligence pack?
A hub on this site that groups four existing Store Actors with distinct jobs: RSS & Atom Feed Extractor (known-publisher feed XML), Google News Scraper (query-based Google News RSS), Article Content Extractor (article HTML cleanup), and Website Content Extractor (docs, pricing, policy, and product HTML cleanup). It is not a new Actor and it has no Store listing of its own.
When should I use RSS & Atom Feed Extractor versus Google News Scraper?
Use RSS & Atom Feed Extractor when you already know the publishers you trust and have their public RSS or Atom feed URLs (feedUrls, maximum 50). Use Google News Scraper when you do not have feed URLs yet and want query-based discovery (queries, maximum 50), localized with language and country.
When should I use Article Content Extractor versus Website Content Extractor?
Use Article Content Extractor for public news, blog, newsroom, and press URLs (maximum 300 per run) when you need headline, byline, publish date, article body, excerpt, and hero image. Use Website Content Extractor for public docs, product, pricing, policy, and help-center pages (maximum 200 per run) that are not article-shaped.
Can RSS or Google News return full article bodies?
No. Both are discovery layers. RSS & Atom Feed Extractor parses feed XML and returns item URLs; it does not crawl website HTML or extract full article bodies from those links. Google News Scraper returns headlines, article URLs, publishers, timestamps, and snippets from Google News RSS; it does not clean HTML or return full article text.
How do these four Actors connect in a pipeline?
- Discover URLs first. Start with RSS & Atom Feed Extractor when you have feed URLs, or Google News Scraper when you have search queries.
- Confirm the dataset shows real article or page URLs, not generic homepages.
- Send news, blog, and press links to Article Content Extractor.
- Send docs, product, pricing, policy, or help-center links to Website Content Extractor.
Each Actor is a separate Store run. This hub does not chain them automatically.
How are the four Actors priced?
Each Actor bills on its own published Store pay-per-event prices. RSS & Atom Feed Extractor: $1.00 per 1,000 RSS or Atom items ($0.001 per successfully parsed public feed item). Google News Scraper: $3.00 per 1,000 Google News RSS articles ($0.003 per successfully parsed public RSS article row). Article Content Extractor: $0.008 per useful article row, $2.50 per article-content-audit-report, and $5.00 per article-batch-export. Website Content Extractor: $0.001 per Actor start and $0.009 per useful content row ($9.00 per 1,000 results). This hub is not billed.
Open the matching Store listing from the landing you chose. There is no pack SKU.
- Open RSS & Atom Feed Extractor on Apify
- Open Google News Scraper on Apify
- Open Article Content Extractor on Apify
- Open Website Content Extractor on Apify
Is this pack a separate Apify Actor?
No. There is no Content Intelligence pack Actor and no pack Store URL. Open the individual Actor landings, then the Store listings those landings already confirm.
Related pages
- RSS & Atom Feed Extractor — known-publisher RSS/Atom feeds
- Google News Scraper — query-based Google News RSS
- Article Content Extractor — news, blog, and press HTML
- Website Content Extractor — docs, pricing, policy, and product HTML
- HHS Healthcare Data Breach Change Scraper — HHS OCR disclosure changes, not this pack
- RSS feed vs iTunes Search API — podcast catalog RSS, not this pack
- Tools