Article URL to Markdown for RAG

Extract main article text to Markdown for RAG from a URL

Job: paste a public article URL, get the main body as Markdown you can ingest - headline, byline, date, and text, not the chrome around the story.

outputFormat defaults to markdown. After Start, export the default dataset as CSV or JSON (articleText = Markdown body).

Open Input → Start. 1 URL ≈ $0.008. From $8.00 / 1,000 useful rows. Leave the $2.50 audit report and $5.00 batch export off. You already have the URL - this is not Google News search and not a full-site crawler.

Open Input → Start on Article Content Extractor - one public article URL, markdown default, report/export off (~$0.008), then Download CSV or JSON from the Dataset.

Open Input with one URL - leave audit and batch export off

The paying proof on this page is one path: Try for free → Input → paste one public article URL → Start → export Markdown.

Do these before Start:

  1. Paste a single public article URL you already have.
  2. Leave outputFormat on markdown (the live default).
  3. Leave the $2.50 audit report and the $5.00 batch export off. You need the body row first.

Sign in is required to Start. When the run finishes, open the default dataset and take the Markdown body (articleText). This is not Google News discovery and not a full-site crawl.

1 URL ≈ $0.008. From $8.00 / 1,000 useful rows.

Try for free → open Input (1 URL, report/export off) → Start

Discover URLs first - then extract the body here

Path What you get Next step
Google News / RSS discovery Headlines + article URLs (or feed items) Not full RAG body
This guide / Actor headline, author, publishedAt, articleText (Markdown), sourceUrl Export dataset CSV/JSON
Website Content Extractor Docs / help-center / policy pages Different shape - not article chrome stripping

Suggested first run:

{
  "urls": ["https://en.wikipedia.org/wiki/Web_scraping"],
  "outputFormat": "markdown",
  "includeImages": false,
  "delivery": "dataset",
  "dryRun": false
}

Raise URL count only after that row is usable. Failed or no-content fetches stay out of the default dataset.

Start path: one URL to Markdown

  1. Open taroyamada/article-content-extractor Input. Sign in (an account is required to run).
  2. Put one public news, blog, newsroom, or press URL in urls. Keep outputFormat markdown (default), includeImages false, delivery dataset, dryRun false.
  3. Click Start. When the run succeeds, open the default dataset and export CSV or JSON. articleText is the Markdown body.
  4. Raise URL count only after that row is usable. Leave generateReport and emitExport false.

Failed or no-content fetches stay out of the default dataset, so they are not billed as useful article rows.

Open Input → Start (1 article URL)

Cheapest first Markdown row

Paste this JSON if the form is empty. Swap in a public article URL you are authorized to fetch. Schema prefill on the Store is a Wikipedia Web scraping URL; one URL is enough for a first proof.

{
  "urls": [
    "https://en.wikipedia.org/wiki/Web_scraping"
  ],
  "outputFormat": "markdown",
  "includeImages": false,
  "delivery": "dataset",
  "dryRun": false
}

Live PPE (copy only; Store prices unchanged): Actor Start $0.00005 + useful article row $0.008. 1 URL ≈ $0.008 if those two events charge. Store card: From $8.00 / 1,000 useful rows. Audit report $2.50 and batch export $5.00 stay off.

First Markdown extract Value
urls Required. One article-shaped HTTP(S) page (max 300)
outputFormat markdown (default). Enum: markdown or text. No HTML
includeImages false on the first run
Useful row fields headline, author, publishedAt, articleText, excerpt, heroImage, sourceUrl
If 1 Start + 1 useful row $0.00005 + $0.008 ≈ $0.008

Do not send Website fields (includeMetadata, maxChargeUsd), Google News queries, or RSS feedUrls.

Open Input → Start

What this page is not

Can I extract main article text to Markdown from a URL for RAG?

Yes, for article-shaped public HTTP(S) pages you already have (news, blog, newsroom, press; maximum 300 URLs). The run returns structured article rows: headline, author, publishedAt, articleText, excerpt, heroImage, and sourceUrl. outputFormat defaults to markdown, which the live schema recommends for first-run proof and LLM ingestion. You supply the URL list. It is not a search surface or a full-site crawler.

Is markdown the default output?

Yes. The published outputFormat enum is markdown (default) or text. This extractor does not emit HTML. Website Content Extractor is the sibling that can return html as well as markdown or text, and is for docs, pricing, policy, and help-center pages that are not article-shaped. Side-by-side: Article vs Website Content Extractor for RAG.

What is the cheapest first URL-to-Markdown run?

One urls item, includeImages false, delivery dataset, dryRun false. Leave generateReport and emitExport off. Live PPE is Actor Start $0.00005 plus $0.008 per useful article row. 1 URL ≈ $0.008 if those two events charge. From $8.00 / 1,000 useful rows. Failed or no-content rows stay out of the default dataset. Event table: cheapest first paid run.

Does this crawl the whole site?

No. You paste article URLs (maximum 300). Loopback, private, and non-HTML targets are rejected. Default bounds are 3 redirects and 5 MiB per HTML response. Schema prefill is one Wikipedia Web scraping URL. Only fetch public pages you own or are authorized to fetch.

Why not send Google News or RSS output straight into RAG?

Those Actors are discovery layers. Google News Scraper returns headlines and article URLs from queries, without full article text. RSS & Atom Feed Extractor parses known-publisher feed XML and does not crawl page HTML. For RAG you still need the article body. Discover URLs first, then send news, blog, and press links here.

Does this emit HTML?

No. outputFormat is markdown or text only. Do not send html as outputFormat, Website fields such as includeMetadata or maxChargeUsd, Google News queries, or RSS feedUrls. additionalProperties is false on the live schema.

What columns are on an article row after export?

rowType article. Documented fields include url, headline, author, publishedAt, articleText, excerpt, heroImage, and sourceUrl. Markdown is in articleText when outputFormat is markdown. After Start, export the default dataset as CSV or JSON from the dataset view.

Open Input → Start. 1 URL ≈ $0.008. From $8.00 / 1,000 useful rows. Leave the $2.50 audit report and $5.00 batch export off.

Open Input → Start on Article Content Extractor - one public article URL, markdown default, report/export off (~$0.008), then Download CSV or JSON from the Dataset.

Related pages

Public HTTP(S) article pages you own or are authorized to fetch. Not a crawler. Not HTML emit. Failed fetches are not billed as useful rows. Store prices unchanged on this page.