Hacker News
Hacker News top/new/best stories and optional comment trees
This Actor pulls public Hacker News story lists (top, new, best, ask, show, job) and can optionally include comment trees. Default mode is top; includeComments defaults false.
Not Google News, not full article body extraction, and not a sentiment model. Story JSON for RAG/downstream scoring only.
from $1.00 / 1,000 Hacker News story rows ($0.001 per successfully delivered story row). No Actor Start.
Open Hacker News Stories Scraper on Apify
Hacker News Stories Scraper, not a neighboring Actor
Use this page for this Actor’s job. Use Google News Scraper for Google News headlines; Use Article Content Extractor for Clean article body.
| This Actor | Google News Scraper | Article Content Extractor | |
|---|---|---|---|
| Intent | Hacker News Stories Scraper | Google News headlines | Clean article body |
| Primary input | see schema |
queries / topics | article urls |
| What it reads | Public sources listed on the Store page | Google News | article HTML |
| Primary output | Dataset rows billed per live PPE | headline URLs | title/byline/body |
| Not this job | Not Google News, not full article body extraction, and not a sentiment model. Story JSON for RAG/downstream scoring only. | Not Hacker News | Not HN story lists |
Store ID: taroyamada/hacker-news-intelligence. Respect source terms, robots.txt, and rate limits.
Use cases
- HN firehose for RAG
- Sentiment/watch on top stories
- Ask/Show/Job list pulls
How is Hacker News Stories Scraper different from Google News Scraper and Article Content Extractor?
Hacker News Stories Scraper — Comments & RAG JSON (taroyamada/hacker-news-intelligence): This Actor pulls public Hacker News story lists (top, new, best, ask, show, job) and can optionally include comment trees. Default mode is top; includeComments defaults false. Not Google News, not full article body extraction, and not a sentiment model. Story JSON for RAG/downstream scoring only. Google News Scraper is for Google News headlines (input queries / topics; Google News; headline URLs). Not Hacker News. Article Content Extractor is for Clean article body (input article urls; article HTML; title/byline/body). Not HN story lists.
What input is required?
Live required fields: none listed (see live fields). exampleRunInput is mode top, maxItems 10, includeComments false. Schema default maxItems 100 (max 500). No required fields.
| Field | Type | Default | Notes |
|---|---|---|---|
mode |
string | top |
Mode. Operation mode |
maxItems |
integer | 100 |
Max Items. Maximum number of items to return |
minScore |
integer | 0 |
Min Score. Minimum score threshold for filtering |
includeComments |
boolean | false |
Include Comments. Include comments in output |
timeoutMs |
integer | 15000 |
Timeout (ms). Request timeout in milliseconds |
delivery |
string | dataset |
Delivery. Where to send results: dataset or webhook |
webhookUrl |
string | empty |
Webhook URL. Webhook URL to POST results to (if delivery=webhook) |
dryRun |
boolean | false |
Dry Run. Run without saving results (for testing) |
Published Store example run input (omitted fields take schema defaults):
{
"mode": "top",
"maxItems": 10,
"minScore": 0,
"includeComments": false,
"delivery": "dataset",
"dryRun": false
}
Run Hacker News Stories Scraper on Apify
How do dataset, webhook, and dry-run delivery work?
delivery defaults to dataset on the live schema. Dataset output is the billable surface when rows are written. webhookUrl is used when delivery is webhook (and typically not during dryRun). dryRun true validates or samples without the usual dataset/webhook side effects described on the Store schema. Unchanged runs that write zero default-dataset rows typically charge $0.00 on current PPE.
What does a result contain?
Published README Output Example / Sample Output JSON.
{
"id": 12345678,
"title": "Claude 4.5 released with new features",
"url": "https://anthropic.com/news/claude-4-5",
"score": 523,
"by": "user123",
"time": 1712345678,
"descendants": 142,
"type": "story"
}
There is no published output JSON schema on the Store page.
How is Hacker News Stories Scraper priced?
Billing is pay per event. The live Store card is from $1.00 / 1,000 Hacker News story rows ($0.001 per successfully delivered story row). No Actor Start. Current PPE:
| Event | Price | Emitted when |
|---|---|---|
apify-default-dataset-item (Hacker News story row) |
$0.001 | Charged only for one successfully delivered public Hacker News story row. |
The published README Cost block is stale versus the live Store pricing tab. README Cost quotes actor-start $0.01 + dataset-item $0.003. Live: Hacker News story row $0.001, no start. This page quotes live PPE only.
from $1.00 / 1,000 Hacker News story rows ($0.001 per successfully delivered story row). No Actor Start.
See Hacker News Stories Scraper pricing on Apify
Limits to keep in mind
- Public HN API/pages
- maxItems 1–500
- Comments off by default
- Not article extraction
- Respect source terms, robots.txt, and rate limits.
Open Hacker News Stories Scraper on Apify
Related pages
- Google News Scraper — News headlines, not HN.
- Article Content Extractor — Fetch article bodies from HN urls downstream.
- RSS & Atom Feed Extractor — Feed items you already have URLs for
- Tools