YouTube transcript extractor for RAG
Use this page when the job is ingesting public YouTube captions into RAG. YouTube Transcript Scraper is the existing Store Actor for that caption step. It is not a new Actor, not a video-file downloader, and not a website HTML cleaner. Pass public video URLs, video IDs, playlist URLs, or channel URLs.
Live Store price: from $2.50 / 1,000 results (Result at $0.0025 per default-dataset item). Actor Start $0.00005. Cheapest first paid path: 1 video URL ≈ $0.00255.
Try the cheapest first paid YouTube transcript run on Apify Input
Open taroyamada/youtube-transcript-bulk-api on the Apify Store · YouTube Transcript Scraper landing
What RAG gets from each transcript row
The published README field list is videoId, videoUrl, videoTitle, channelTitle, sourceType, status, chargedEvent, sourceUrls, errors, language, sourceLanguage, isAutoGenerated, captionTrackName, segmentCount, fullText, segments, formattedTranscript, errorCode, errorMessage, and scrapedAt. One row per video. outputFormat defaults to json and can emit an extra formatted field as text, srt, or vtt. Raw transcript rows are source context for embeddings and subtitle export — they are not a corpus RAG report.
For docs, pricing, policy, and help-center HTML, use Website Content Extractor and the docs and help-center RAG page. For news and blog article bodies, use article extractor for news and blog RAG.
Tiny first-paid input
The published Store input example is one public watch URL, language en, includeAutoGenerated true, outputFormat json, delivery dataset, and dryRun false. The schema prefill is the watch URL below. Replace it with one public captioned video you are authorized to process.
{
"videoUrls": [
"https://www.youtube.com/watch?v=dQw4w9WgXcQ"
],
"language": "en",
"includeAutoGenerated": true,
"outputFormat": "json",
"delivery": "dataset",
"dryRun": false
}
If extraction succeeds, live PPE is Actor Start $0.00005 plus one Result at $0.0025 = $0.00255. Age-restricted, private, deleted, or captionless videos return an error row instead of failing the full run. Playlist and channel expansion discovers visible video IDs only. README recommended event names such as transcript_extracted are not the live Store event titles; the live primary item event is Result.
Run this YouTube transcript input on Apify
Can this YouTube transcript extractor feed a RAG corpus?
Yes, as a bounded source step. It resolves public video URLs, video IDs, playlist URLs, or channel URLs into visible video IDs, fetches public caption tracks, and returns one dataset row per video with fullText, timed segments, and formatted SRT, VTT, or plain text. It is HTTP-first and does not use browser automation. It is not a video-file downloader, a website HTML cleaner, or a buyer-facing corpus RAG report Actor. Only public videos with public caption tracks are supported.
Why not use Website Content Extractor for video text?
Website Content Extractor fetches public docs, pricing, policy, and help-center page URLs and returns cleaned markdown. It does not read YouTube caption tracks. Use this Actor for public captions. Use Website Content Extractor for page-body markdown.
What is the cheapest first paid YouTube transcript input for RAG?
The published Store input example is one videoUrls item, language en, includeAutoGenerated true, outputFormat json, delivery dataset, and dryRun false. Live PPE is $0.0025 per Result plus Actor Start $0.00005. If those two events charge, documented PPE is $0.00255.
Open YouTube Transcript Scraper Input on Apify
Related pages
- YouTube Transcript Scraper — product landing and field list
- YouTube Channel Analytics Scraper — channel metadata, not caption tracks
- Website content extractor for docs and help-center RAG
- Article extractor for news and blog RAG
- Website Content Extractor — docs and product HTML, not YouTube captions
- Article Content Extractor
- Tools