YouTube RAG
Authorized transcript text into RAG chunks and readiness reports
This Actor audits authorized transcript text for RAG readiness, produces retrieval chunks and source-linked reports, and can optionally try public YouTube captions without login or cookies.
Not a bulk caption dump of arbitrary channels without authorization. Not YouTube Channel Analytics (subscribers/uploads). Not a website HTML extractor. Public captions only; no login or cookies.
from $6.00 / 1,000 Transcript RAG chunk ($0.006 per delivered transcript_rag_chunk); $9.00 per Transcript corpus snapshot report; $29 per YouTube RAG readiness report. No Actor Start.
Open YouTube Transcript Corpus Audit for RAG on Apify
YouTube Transcript Corpus Audit for RAG, not a neighboring Actor
Use this page for this Actor’s job. Use YouTube Transcript Scraper for Bulk captions / SRT / VTT; Use YouTube Channel Analytics Scraper for Public channel stats; Use Website Content Extractor for One-shot page-to-markdown.
| This Actor | YouTube Transcript Scraper | YouTube Channel Analytics Scraper | Website Content Extractor | |
|---|---|---|---|---|
| Intent | YouTube Transcript Corpus Audit for RAG | Bulk captions / SRT / VTT | Public channel stats | One-shot page-to-markdown |
| Primary input | videoUrls / transcriptText | video URLs | channel URLs | urls |
| What it reads | Authorized text and optional public captions | Public captions | Public channel pages | Page HTML |
| Primary output | RAG chunks plus optional snapshot/readiness reports | Transcript rows | Channel analytics rows | Markdown/text |
| Not this job | Not a bulk caption dump of arbitrary channels without authorization. Not YouTube Channel Analytics (subscribers/uploads). Not a website HTML extractor. Public captions only; no login or cookies. | Not a RAG readiness report. | Not transcript chunks. | Not recurring pricing/legal diffs or RAG audits. |
Store ID: taroyamada/youtube-channel-transcript-rag-intelligence. Respect source terms, robots.txt, and rate limits.
Use cases
- Chunk authorized transcript text for retrieval
- Optional public caption fetch on known video URLs
- Snapshot vs readiness report tiers; keep expensive reports off on first run
How is YouTube Transcript Corpus Audit for RAG different from YouTube Transcript Scraper and YouTube Channel Analytics Scraper?
YouTube Transcript Corpus Audit for RAG (taroyamada/youtube-channel-transcript-rag-intelligence): This Actor audits authorized transcript text for RAG readiness, produces retrieval chunks and source-linked reports, and can optionally try public YouTube captions without login or cookies. Not a bulk caption dump of arbitrary channels without authorization. Not YouTube Channel Analytics (subscribers/uploads). Not a website HTML extractor. Public captions only; no login or cookies. YouTube Transcript Scraper is for Bulk captions / SRT / VTT (input video URLs; Public captions; Transcript rows). Not a RAG readiness report. YouTube Channel Analytics Scraper is for Public channel stats (input channel URLs; Public channel pages; Channel analytics rows). Not transcript chunks. Website Content Extractor is for One-shot page-to-markdown (input urls; Page HTML; Markdown/text). Not recurring pricing/legal diffs or RAG audits.
What input is required?
Live required fields: none listed (see live fields). Published exampleRunInput is shown below.
| Field | Type | Default | Notes |
|---|---|---|---|
videoUrls |
any[] | ["https://www.youtube.com/watch?v=dQw4w9WgXcQ"] |
Video URLs. YouTube watch, Shorts, embed, live, or youtu.be URLs. |
videoIds |
any[] | empty |
Video IDs. Direct 11-character YouTube video IDs. |
channelUrls |
any[] | empty |
Channel URLs or @handles. YouTube channel URLs or bare @handles. The actor expands public channel pages into visible video IDs. |
playlistUrls |
any[] | empty |
Playlist URLs. Optional public playlist URLs to expand into visible video IDs. |
transcriptText |
string | empty |
Transcript Text. Optional transcript text that you own or are authorized to process. This path does not fetch YouTube and is the most reliable input for RAG audits. |
transcriptTitle |
string | User-provided transcript |
Transcript Title. Title attached to user-provided transcript chunks and reports. |
transcriptSourceUrl |
string | empty |
Transcript Source URL. Optional source URL used as evidence. The actor does not fetch this URL when transcriptText is provided. |
language |
string | en |
Preferred Caption Language. Preferred caption language code such as en, ja, or es. |
includeAutoGenerated |
boolean | true |
Include Auto-generated Captions. Allow YouTube auto-generated captions when manual captions are unavailable. |
translationLanguage |
string | empty |
Translation Language. Optional YouTube transcript translation target language code. |
maxVideos |
integer | 5 |
Max Videos. Maximum videos to process in one run. minimum=1 maximum=1000 |
chunkSize |
integer | 1200 |
Chunk Size. Approximate maximum transcript characters per RAG chunk. minimum=250 maximum=8000 |
chunkOverlap |
integer | 150 |
Chunk Overlap. Approximate overlap characters carried from the previous chunk, segment-aligned. minimum=0 maximum=2000 |
maxTranscriptChars |
integer | 100000 |
Max Transcript Characters. Maximum transcript text characters to chunk per video. minimum=500 maximum=500000 |
timeoutMs |
integer | 15000 |
HTTP Timeout. Request timeout in milliseconds. minimum=3000 maximum=60000 |
reportTier |
string enum | corpus_snapshot |
Report Tier. Public report tiers produce one decision-ready audit report. Use chunks only when you need legacy per-chunk output; watch summaries are planned and proof-gated. enum: corpus_snapshot, rag_readiness, chunks |
maxChargeUsd |
number | 9 |
Max Charge USD. Safety cap checked before report-tier charging. If the selected report price exceeds this cap, the actor returns a no-charge limit_reached row. minimum=0 maximum=500 |
maxReports |
integer | 1 |
Max Reports. Report-mode safety limit. V1 emits one corpus report per run. minimum=1 maximum=10 |
demoMode |
boolean | false |
Demo Mode. Return a no-charge sample report without external requests. |
dedupeVideos |
boolean | true |
Dedupe Videos. Remove duplicate video IDs after combining video, playlist, and channel sources. |
limit |
integer | 50 |
Result Row Limit. Maximum chunk/warning rows to emit. minimum=1 maximum=10000 |
delivery |
string enum | dataset |
Delivery Mode. Choose whether to write rows to the dataset, post the final payload to webhookUrl, or both. enum: dataset, webhook, dataset_and_webhook |
webhookUrl |
string | empty |
Webhook URL. HTTPS endpoint used when delivery is webhook or dataset_and_webhook. |
dryRun |
boolean | false |
Dry Run. Set true to validate input and emit sample RAG rows without fetching YouTube or delivering dataset/webhook output. |
Published Store example run input (omitted fields take schema defaults):
{
"videoUrls": [
"https://www.youtube.com/watch?v=dQw4w9WgXcQ"
],
"videoIds": [],
"channelUrls": [],
"playlistUrls": [],
"language": "en",
"includeAutoGenerated": true,
"reportTier": "corpus_snapshot",
"maxChargeUsd": 9,
"maxVideos": 1,
"chunkSize": 1200,
"chunkOverlap": 150,
"limit": 20,
"delivery": "dataset",
"webhookUrl": "",
"dryRun": false
}
Run YouTube Transcript Corpus Audit for RAG on Apify
How do dataset, webhook, and dry-run delivery work?
delivery defaults to dataset on the live schema. Dataset output is the billable surface when rows are written. webhookUrl is used when delivery is webhook (and typically not during dryRun). dryRun true validates or samples without the usual dataset/webhook side effects described on the Store schema. Unchanged runs that write zero default-dataset rows typically charge $0.00 on current PPE.
What does a result contain?
README output fence is truncated or invalid JSON on Store. Illustration head: { "meta": { "actorName": "youtube-channel-transcript-rag-intelligence", "actorTitle": "YouTube Channel Transcript RAG Intelligence", "fetchedAt": "2026-05-09T00:00:00.000Z", "totalRows": 2 }, "rows": [ { "rowType": "corpus_audit_report", "reportTier": "rag_readiness", "status": "… Treat README samples as illustrations, not a live coverage guarantee. There is no published output JSON schema on the Store page.
How is YouTube Transcript Corpus Audit for RAG priced?
Billing is pay per event. The live Store card is from $6.00 / 1,000 Transcript RAG chunk ($0.006 per delivered transcript_rag_chunk); $9.00 per Transcript corpus snapshot report; $29 per YouTube RAG readiness report. No Actor Start. Current PPE:
| Event | Price | Emitted when |
|---|---|---|
transcript_rag_chunk (Transcript RAG chunk) |
$0.006 | Charged for each retrieval-ready chunk from a public-caption transcript. |
corpus_snapshot_report (Transcript corpus snapshot report) |
$9.00 | Charged once for a successful transcript coverage and corpus-risk report. |
rag_readiness_report (YouTube RAG readiness report) |
$29 | Charged once for a successful RAG readiness report with source-linked findings and actions. |
from $6.00 / 1,000 Transcript RAG chunk ($0.006 per delivered transcript_rag_chunk); $9.00 per Transcript corpus snapshot report; $29 per YouTube RAG readiness report. No Actor Start.
See YouTube Transcript Corpus Audit for RAG pricing on Apify
Limits to keep in mind
- No login, cookies, or private videos
- maxVideos / chunkSize caps
- reportTier controls billed reports
- Treat auto-generated captions as noisy
- Respect source terms, robots.txt, and rate limits.
Open YouTube Transcript Corpus Audit for RAG on Apify
Related pages
- YouTube Transcript Scraper — Not a RAG readiness report.
- YouTube Channel Analytics Scraper — Not transcript chunks.
- Website Content Extractor — Not recurring pricing/legal diffs or RAG audits.
- Tools