YouTube RAG

Authorized transcript text into RAG chunks and readiness reports

This Actor audits authorized transcript text for RAG readiness, produces retrieval chunks and source-linked reports, and can optionally try public YouTube captions without login or cookies.

Not a bulk caption dump of arbitrary channels without authorization. Not YouTube Channel Analytics (subscribers/uploads). Not a website HTML extractor. Public captions only; no login or cookies.

from $6.00 / 1,000 Transcript RAG chunk ($0.006 per delivered transcript_rag_chunk); $9.00 per Transcript corpus snapshot report; $29 per YouTube RAG readiness report. No Actor Start.

Open YouTube Transcript Corpus Audit for RAG on Apify

YouTube Transcript Corpus Audit for RAG, not a neighboring Actor

Use this page for this Actor’s job. Use YouTube Transcript Scraper for Bulk captions / SRT / VTT; Use YouTube Channel Analytics Scraper for Public channel stats; Use Website Content Extractor for One-shot page-to-markdown.

This ActorYouTube Transcript ScraperYouTube Channel Analytics ScraperWebsite Content Extractor
IntentYouTube Transcript Corpus Audit for RAGBulk captions / SRT / VTTPublic channel statsOne-shot page-to-markdown
Primary inputvideoUrls / transcriptTextvideo URLschannel URLsurls
What it readsAuthorized text and optional public captionsPublic captionsPublic channel pagesPage HTML
Primary outputRAG chunks plus optional snapshot/readiness reportsTranscript rowsChannel analytics rowsMarkdown/text
Not this jobNot a bulk caption dump of arbitrary channels without authorization. Not YouTube Channel Analytics (subscribers/uploads). Not a website HTML extractor. Public captions only; no login or cookies.Not a RAG readiness report.Not transcript chunks.Not recurring pricing/legal diffs or RAG audits.

Store ID: taroyamada/youtube-channel-transcript-rag-intelligence. Respect source terms, robots.txt, and rate limits.

Use cases

How is YouTube Transcript Corpus Audit for RAG different from YouTube Transcript Scraper and YouTube Channel Analytics Scraper?

YouTube Transcript Corpus Audit for RAG (taroyamada/youtube-channel-transcript-rag-intelligence): This Actor audits authorized transcript text for RAG readiness, produces retrieval chunks and source-linked reports, and can optionally try public YouTube captions without login or cookies. Not a bulk caption dump of arbitrary channels without authorization. Not YouTube Channel Analytics (subscribers/uploads). Not a website HTML extractor. Public captions only; no login or cookies. YouTube Transcript Scraper is for Bulk captions / SRT / VTT (input video URLs; Public captions; Transcript rows). Not a RAG readiness report. YouTube Channel Analytics Scraper is for Public channel stats (input channel URLs; Public channel pages; Channel analytics rows). Not transcript chunks. Website Content Extractor is for One-shot page-to-markdown (input urls; Page HTML; Markdown/text). Not recurring pricing/legal diffs or RAG audits.

What input is required?

Live required fields: none listed (see live fields). Published exampleRunInput is shown below.

Field Type Default Notes
videoUrls any[] ["https://www.youtube.com/watch?v=dQw4w9WgXcQ"] Video URLs. YouTube watch, Shorts, embed, live, or youtu.be URLs.
videoIds any[] empty Video IDs. Direct 11-character YouTube video IDs.
channelUrls any[] empty Channel URLs or @handles. YouTube channel URLs or bare @handles. The actor expands public channel pages into visible video IDs.
playlistUrls any[] empty Playlist URLs. Optional public playlist URLs to expand into visible video IDs.
transcriptText string empty Transcript Text. Optional transcript text that you own or are authorized to process. This path does not fetch YouTube and is the most reliable input for RAG audits.
transcriptTitle string User-provided transcript Transcript Title. Title attached to user-provided transcript chunks and reports.
transcriptSourceUrl string empty Transcript Source URL. Optional source URL used as evidence. The actor does not fetch this URL when transcriptText is provided.
language string en Preferred Caption Language. Preferred caption language code such as en, ja, or es.
includeAutoGenerated boolean true Include Auto-generated Captions. Allow YouTube auto-generated captions when manual captions are unavailable.
translationLanguage string empty Translation Language. Optional YouTube transcript translation target language code.
maxVideos integer 5 Max Videos. Maximum videos to process in one run. minimum=1 maximum=1000
chunkSize integer 1200 Chunk Size. Approximate maximum transcript characters per RAG chunk. minimum=250 maximum=8000
chunkOverlap integer 150 Chunk Overlap. Approximate overlap characters carried from the previous chunk, segment-aligned. minimum=0 maximum=2000
maxTranscriptChars integer 100000 Max Transcript Characters. Maximum transcript text characters to chunk per video. minimum=500 maximum=500000
timeoutMs integer 15000 HTTP Timeout. Request timeout in milliseconds. minimum=3000 maximum=60000
reportTier string enum corpus_snapshot Report Tier. Public report tiers produce one decision-ready audit report. Use chunks only when you need legacy per-chunk output; watch summaries are planned and proof-gated. enum: corpus_snapshot, rag_readiness, chunks
maxChargeUsd number 9 Max Charge USD. Safety cap checked before report-tier charging. If the selected report price exceeds this cap, the actor returns a no-charge limit_reached row. minimum=0 maximum=500
maxReports integer 1 Max Reports. Report-mode safety limit. V1 emits one corpus report per run. minimum=1 maximum=10
demoMode boolean false Demo Mode. Return a no-charge sample report without external requests.
dedupeVideos boolean true Dedupe Videos. Remove duplicate video IDs after combining video, playlist, and channel sources.
limit integer 50 Result Row Limit. Maximum chunk/warning rows to emit. minimum=1 maximum=10000
delivery string enum dataset Delivery Mode. Choose whether to write rows to the dataset, post the final payload to webhookUrl, or both. enum: dataset, webhook, dataset_and_webhook
webhookUrl string empty Webhook URL. HTTPS endpoint used when delivery is webhook or dataset_and_webhook.
dryRun boolean false Dry Run. Set true to validate input and emit sample RAG rows without fetching YouTube or delivering dataset/webhook output.

Published Store example run input (omitted fields take schema defaults):

{
  "videoUrls": [
    "https://www.youtube.com/watch?v=dQw4w9WgXcQ"
  ],
  "videoIds": [],
  "channelUrls": [],
  "playlistUrls": [],
  "language": "en",
  "includeAutoGenerated": true,
  "reportTier": "corpus_snapshot",
  "maxChargeUsd": 9,
  "maxVideos": 1,
  "chunkSize": 1200,
  "chunkOverlap": 150,
  "limit": 20,
  "delivery": "dataset",
  "webhookUrl": "",
  "dryRun": false
}

Run YouTube Transcript Corpus Audit for RAG on Apify

How do dataset, webhook, and dry-run delivery work?

delivery defaults to dataset on the live schema. Dataset output is the billable surface when rows are written. webhookUrl is used when delivery is webhook (and typically not during dryRun). dryRun true validates or samples without the usual dataset/webhook side effects described on the Store schema. Unchanged runs that write zero default-dataset rows typically charge $0.00 on current PPE.

What does a result contain?

README output fence is truncated or invalid JSON on Store. Illustration head: { "meta": { "actorName": "youtube-channel-transcript-rag-intelligence", "actorTitle": "YouTube Channel Transcript RAG Intelligence", "fetchedAt": "2026-05-09T00:00:00.000Z", "totalRows": 2 }, "rows": [ { "rowType": "corpus_audit_report", "reportTier": "rag_readiness", "status": "… Treat README samples as illustrations, not a live coverage guarantee. There is no published output JSON schema on the Store page.

How is YouTube Transcript Corpus Audit for RAG priced?

Billing is pay per event. The live Store card is from $6.00 / 1,000 Transcript RAG chunk ($0.006 per delivered transcript_rag_chunk); $9.00 per Transcript corpus snapshot report; $29 per YouTube RAG readiness report. No Actor Start. Current PPE:

Event Price Emitted when
transcript_rag_chunk (Transcript RAG chunk) $0.006 Charged for each retrieval-ready chunk from a public-caption transcript.
corpus_snapshot_report (Transcript corpus snapshot report) $9.00 Charged once for a successful transcript coverage and corpus-risk report.
rag_readiness_report (YouTube RAG readiness report) $29 Charged once for a successful RAG readiness report with source-linked findings and actions.

from $6.00 / 1,000 Transcript RAG chunk ($0.006 per delivered transcript_rag_chunk); $9.00 per Transcript corpus snapshot report; $29 per YouTube RAG readiness report. No Actor Start.

See YouTube Transcript Corpus Audit for RAG pricing on Apify

Limits to keep in mind

Open YouTube Transcript Corpus Audit for RAG on Apify

Related pages