OpenAlex works

Extract OpenAlex works, authors, citations, and OA signals

OpenAlex Scholarly Works Scraper queries the official OpenAlex API for works, authors, institutions, sources, citations, topics, DOI, and open-access signals. It is bibliometrics and scholarly discovery — not NCBI PubMed E-utilities, not NIH RePORTER awards, and not Google News RSS.

This independent Actor is not affiliated with, sponsored by, or endorsed by OpenAlex.

$4.00 / 1,000 openalex research rows ($0.004 per delivered work, author, or source row)

Open OpenAlex Scholarly Works Scraper on Apify

Official OpenAlex API, not PubMed E-utilities or news RSS

Use this page when the job is OpenAlex works and related rows. Use PubMed Literature Watch & Research Report for NCBI citation metadata and literature-update reports. Use NIH RePORTER Funding Landscape Report for NIH awards. Use PubMed & Clinical Trials Evidence Gap Report for trial↔PubMed gaps. Use Google News Scraper for news headlines.

This Actor PubMed Literature Watch NIH RePORTER Funding Landscape
Intent OpenAlex works, authors, sources, OA, citations PubMed citation watch, reports, new-PMID alerts NIH award landscape, alerts, exports
Source Official OpenAlex API Official NCBI ESearch/ESummary Official NIH RePORTER v2 Project API
Primary input Search terms or OpenAlex IDs PubMed searchTerms or pmids watches[] (terms, FY, activity codes)
Reports / exports No. Work/author/source rows only Yes — literature-update report and citation export Yes — landscape report and export events

Use cases

How is OpenAlex Scholarly Works Scraper different from PubMed Literature Watch and Google News Scraper?

This Actor queries the official OpenAlex API for works, authors, institutions, sources, citations, topics, DOI, and open-access signals. Store ID taroyamada/openalex-research-intelligence. PubMed Literature Watch queries NCBI ESearch/ESummary for PubMed citation metadata, literature-update reports, and new-PMID alerts; it is not OpenAlex. Google News Scraper is query-based Google News RSS of headlines and article URLs. Article Content Extractor cleans news/blog HTML. NIH RePORTER Funding Landscape Report is NIH awards, not scholarly works. Use this Actor for OpenAlex bibliometrics and discovery, then confirm a focused PubMed query when needed.

What input is required?

The live schema required array is empty. additionalProperties is false. Supply at least one of searchTerms, workIds, authorIds, institutionIds, or conceptIds.

Field Type Default Notes
searchTerms string[] prefill LLM/RAG Work search keywords
workIds / authorIds / institutionIds / conceptIds string[] — OpenAlex IDs or URLs (example work id W2741809807)
fromDate / toDate string — YYYY-MM-DD
sort string cited_by_count:desc Also publication_date:desc, publication_date:asc, relevance_score:desc. Relevance is search-only
limitPerSource integer 25 Works per source input
maxWorks integer 100 Global unique work cap
includeAbstract boolean false Reconstruct abstract from inverted index
mailto string — OpenAlex polite-pool email
timeoutMs integer 20000 Per-request timeout
delivery string dataset dataset or webhook
monitor boolean false Emit only unseen works when true
monitorKey string — Blank → derived from query
initialRunMode string baseline_only Used when monitor is true
dryRun boolean false Skip dataset/webhook

Published Store Quickstart (free baseline monitor):

{
  "searchTerms": ["retrieval augmented generation"],
  "fromDate": "2025-01-01",
  "sort": "publication_date:desc",
  "limitPerSource": 10,
  "maxWorks": 20,
  "monitor": true,
  "monitorKey": "rag-literature-watch",
  "initialRunMode": "baseline_only",
  "delivery": "dataset",
  "dryRun": false
}

Published example run input (non-monitor extract):

{
  "searchTerms": ["retrieval augmented generation"],
  "workIds": [],
  "authorIds": [],
  "institutionIds": [],
  "conceptIds": [],
  "fromDate": "2024-01-01",
  "limitPerSource": 10,
  "maxWorks": 20,
  "includeAbstract": false,
  "dryRun": false
}

Run OpenAlex Scholarly Works Scraper on Apify

What does a result contain?

Published README row types:

Sample JSON lives in docs/sample-output.json on the Actor, not inlined on the Store README. This Actor does not emit report, export, or synthetic alert rows; companion Actors do.

Does this Actor return abstracts or fabricate reports?

includeAbstract default is false. When true, abstracts are reconstructed from the OpenAlex inverted index and increase payload size. OpenAlex is not Crossref, PubMed, Semantic Scholar, or a publisher API. Citation counts can lag. Auth-free official API. It does not fabricate literature reports or funding landscapes; use PubMed Literature Watch and NIH RePORTER Funding Landscape Report for those jobs. Fail-closed on state/charge receipt errors.

How is OpenAlex Scholarly Works Scraper priced?

Billing is pay per event. The live Store card is $4.00 / 1,000 openalex research rows. The billed event is OpenAlex research row (apify-default-dataset-item) at $0.004, charged for one delivered source-linked OpenAlex work, author, or source row. There is no Actor Start on the current pricing tab. A first baseline_only monitor and a later unchanged poll cost $0.00. emit_backfill bills the current baseline.

$4.00 per 1,000 openalex research rows ($0.004 per delivered row)

See OpenAlex Scholarly Works Scraper pricing on Apify

Limits to keep in mind

Open OpenAlex Scholarly Works Scraper on Apify

Related pages