Article URL to Markdown for RAG
Extract main article text to Markdown for RAG from a URL
Job: paste a public article URL, get the main body as Markdown you can ingest - headline, byline, date, and text, not the chrome around the story.
outputFormat defaults to markdown. After Start, export the default dataset as CSV or JSON (articleText = Markdown body).
Open Input → Start. 1 URL ≈ $0.008. From $8.00 / 1,000 useful rows. Leave the $2.50 audit report and $5.00 batch export off. You already have the URL - this is not Google News search and not a full-site crawler.
Open Input with one URL - leave audit and batch export off
The paying proof on this page is one path: Try for free → Input → paste one public article URL → Start → export Markdown.
Do these before Start:
- Paste a single public article URL you already have.
- Leave
outputFormaton markdown (the live default). - Leave the $2.50 audit report and the $5.00 batch export off. You need the body row first.
Sign in is required to Start. When the run finishes, open the default dataset and take the Markdown body (articleText). This is not Google News discovery and not a full-site crawl.
1 URL ≈ $0.008. From $8.00 / 1,000 useful rows.
Try for free → open Input (1 URL, report/export off) → Start
Discover URLs first - then extract the body here
| Path | What you get | Next step |
|---|---|---|
| Google News / RSS discovery | Headlines + article URLs (or feed items) | Not full RAG body |
| This guide / Actor | headline, author, publishedAt, articleText (Markdown), sourceUrl |
Export dataset CSV/JSON |
| Website Content Extractor | Docs / help-center / policy pages | Different shape - not article chrome stripping |
Suggested first run:
{
"urls": ["https://en.wikipedia.org/wiki/Web_scraping"],
"outputFormat": "markdown",
"includeImages": false,
"delivery": "dataset",
"dryRun": false
}
Raise URL count only after that row is usable. Failed or no-content fetches stay out of the default dataset.
Start path: one URL to Markdown
- Open
taroyamada/article-content-extractorInput. Sign in (an account is required to run). - Put one public news, blog, newsroom, or press URL in
urls. KeepoutputFormatmarkdown(default),includeImagesfalse,deliverydataset,dryRunfalse. - Click Start. When the run succeeds, open the default dataset and export CSV or JSON.
articleTextis the Markdown body. - Raise URL count only after that row is usable. Leave
generateReportandemitExportfalse.
Failed or no-content fetches stay out of the default dataset, so they are not billed as useful article rows.
Open Input → Start (1 article URL)
Cheapest first Markdown row
Paste this JSON if the form is empty. Swap in a public article URL you are authorized to fetch. Schema prefill on the Store is a Wikipedia Web scraping URL; one URL is enough for a first proof.
{
"urls": [
"https://en.wikipedia.org/wiki/Web_scraping"
],
"outputFormat": "markdown",
"includeImages": false,
"delivery": "dataset",
"dryRun": false
}
Live PPE (copy only; Store prices unchanged): Actor Start $0.00005 + useful article row $0.008. 1 URL ≈ $0.008 if those two events charge. Store card: From $8.00 / 1,000 useful rows. Audit report $2.50 and batch export $5.00 stay off.
| First Markdown extract | Value |
|---|---|
urls |
Required. One article-shaped HTTP(S) page (max 300) |
outputFormat |
markdown (default). Enum: markdown or text. No HTML |
includeImages |
false on the first run |
| Useful row fields | headline, author, publishedAt, articleText, excerpt, heroImage, sourceUrl |
| If 1 Start + 1 useful row | $0.00005 + $0.008 ≈ $0.008 |
Do not send Website fields (includeMetadata, maxChargeUsd), Google News queries, or RSS feedUrls.
What this page is not
- Not a Store mirror. Actor overview: Article Content Extractor.
- Not docs / help-center Markdown. That job is website content extractor for docs RAG.
- Not URL discovery. Headlines first: Google News or newsroom RSS.
- Not a procurement notice queue. Bid CSVs: SAM.gov NAICS alerts.
Can I extract main article text to Markdown from a URL for RAG?
Yes, for article-shaped public HTTP(S) pages you already have (news, blog, newsroom, press; maximum 300 URLs). The run returns structured article rows: headline, author, publishedAt, articleText, excerpt, heroImage, and sourceUrl. outputFormat defaults to markdown, which the live schema recommends for first-run proof and LLM ingestion. You supply the URL list. It is not a search surface or a full-site crawler.
Is markdown the default output?
Yes. The published outputFormat enum is markdown (default) or text. This extractor does not emit HTML. Website Content Extractor is the sibling that can return html as well as markdown or text, and is for docs, pricing, policy, and help-center pages that are not article-shaped. Side-by-side: Article vs Website Content Extractor for RAG.
What is the cheapest first URL-to-Markdown run?
One urls item, includeImages false, delivery dataset, dryRun false. Leave generateReport and emitExport off. Live PPE is Actor Start $0.00005 plus $0.008 per useful article row. 1 URL ≈ $0.008 if those two events charge. From $8.00 / 1,000 useful rows. Failed or no-content rows stay out of the default dataset. Event table: cheapest first paid run.
Does this crawl the whole site?
No. You paste article URLs (maximum 300). Loopback, private, and non-HTML targets are rejected. Default bounds are 3 redirects and 5 MiB per HTML response. Schema prefill is one Wikipedia Web scraping URL. Only fetch public pages you own or are authorized to fetch.
Why not send Google News or RSS output straight into RAG?
Those Actors are discovery layers. Google News Scraper returns headlines and article URLs from queries, without full article text. RSS & Atom Feed Extractor parses known-publisher feed XML and does not crawl page HTML. For RAG you still need the article body. Discover URLs first, then send news, blog, and press links here.
Does this emit HTML?
No. outputFormat is markdown or text only. Do not send html as outputFormat, Website fields such as includeMetadata or maxChargeUsd, Google News queries, or RSS feedUrls. additionalProperties is false on the live schema.
What columns are on an article row after export?
rowType article. Documented fields include url, headline, author, publishedAt, articleText, excerpt, heroImage, and sourceUrl. Markdown is in articleText when outputFormat is markdown. After Start, export the default dataset as CSV or JSON from the dataset view.
Open Input → Start. 1 URL ≈ $0.008. From $8.00 / 1,000 useful rows. Leave the $2.50 audit report and $5.00 batch export off.
Related pages
- Article Content Extractor — product landing
- Cheapest first paid run — PPE table
- Article extractor for news and blog RAG
- Article vs Website Content Extractor for RAG
- Docs and help-center Markdown — not article-shaped URLs
- Content Intelligence pack
- Tools
Public HTTP(S) article pages you own or are authorized to fetch. Not a crawler. Not HTML emit. Failed fetches are not billed as useful rows. Store prices unchanged on this page.