The problem it solves
Manual transcription limits studies to a few dozen videos, and ad-hoc scraping scripts break and are hard to reproduce. Researchers need a dependable way to collect complete transcripts for a defined sample of channels and dates, store them with metadata, and share their method so others can replicate it.
What to include in your first version
- A documented sample of channels and a date range
- Complete transcripts stored as raw JSON with metadata
- Caption source recorded for each video (manual, automatic or AI)
- Export to CSV or text for R, Python, NVivo or ATLAS.ti
- A re-runnable collection script for replication
How to build it, step by step
- 01
Define the sample
List the channels and the time window you'll study, and record your inclusion criteria for the methods section.
- 02
Resolve video lists
Call /channels/resolve for each channel and filter the results by published date.
- 03
Collect transcripts
Send video IDs to /batch and store the full JSON response, including the language and caption source of each transcript.
- 04
Prepare for analysis
Export transcripts as text or CSV with video ID, channel, date and timestamps, then analyze them with topic models, keyword frequencies or qualitative coding.
Starter code
This idea is built on the "Recipe 3: Bulk-load a playlist into a knowledge base" pattern. Swap in your own channels, prompts and storage.
import requests
API = "https://www.youtubetranscript.dev/api/v2"
HEADERS = {"Authorization": "Bearer yt_sk_live_YOUR_KEY"}
playlist = requests.post(f"{API}/playlists/resolve", headers=HEADERS,
json={"playlist_url": "https://www.youtube.com/playlist?list=PLAYLIST_ID",
"limit": 100}).json()
ids = [v["video_id"] for v in playlist["items"]]
batch = requests.post(f"{API}/batch", headers=HEADERS, json={
"video_ids": ids,
"format": {"timestamp": True, "paragraphs": True},
}).json()
# Large batches run async: poll GET /batch/{batch_id} until status is "completed"
for r in batch.get("results", []):
if r["status"] == "completed":
for p in r["data"]["transcript"].get("paragraphs") or []:
store(video_id=r["video_id"], start=p["start"], text=p["text"]) # your vector DBWhy it's worth building
For researchers the return is a reproducible dataset at a fraction of the cost and time of manual transcription. Teams also build paid datasets or analytics services for media monitoring firms and think tanks.
Mistakes to avoid
- Mixing automatic and human captions without noting it: record the caption source because accuracy differs.
- Not saving raw responses: store the original JSON so your analysis can be re-run.
- Ignoring ethics review and platform terms: check your institution's requirements and YouTube's terms before publishing datasets.
Questions about this idea
Can I use YouTube transcripts for academic research?
Many researchers do. Check your institution's ethics requirements, YouTube's terms and copyright law, especially before sharing a dataset publicly.
How accurate are automatic YouTube captions?
They vary with audio quality, accents and topic. The API reports whether each transcript came from manual captions, automatic captions or AI speech recognition so you can account for it.
How many videos can I collect at once?
A single batch request accepts up to 3,000 video IDs on the Business plan. Larger studies can run several batches.