BUILD GUIDE · RESEARCHERS

How to build a YouTube transcript corpus for discourse analysis

YouTube is a primary venue for political commentary, news, education and culture, but researchers have mostly studied it through titles, comments and view counts. Transcripts let you analyze what is actually said, at scale, with timestamps for every quote.

Who it's for
Communication, media studies, political science and linguistics researchers
Build time
1–2 weeks
API endpoints
/channels/resolve/batch/transcripts

The problem it solves

Manual transcription limits studies to a few dozen videos, and ad-hoc scraping scripts break and are hard to reproduce. Researchers need a dependable way to collect complete transcripts for a defined sample of channels and dates, store them with metadata, and share their method so others can replicate it.

What to include in your first version

  • A documented sample of channels and a date range
  • Complete transcripts stored as raw JSON with metadata
  • Caption source recorded for each video (manual, automatic or AI)
  • Export to CSV or text for R, Python, NVivo or ATLAS.ti
  • A re-runnable collection script for replication

How to build it, step by step

  1. 01

    Define the sample

    List the channels and the time window you'll study, and record your inclusion criteria for the methods section.

  2. 02

    Resolve video lists

    Call /channels/resolve for each channel and filter the results by published date.

  3. 03

    Collect transcripts

    Send video IDs to /batch and store the full JSON response, including the language and caption source of each transcript.

  4. 04

    Prepare for analysis

    Export transcripts as text or CSV with video ID, channel, date and timestamps, then analyze them with topic models, keyword frequencies or qualitative coding.

Starter code

This idea is built on the "Recipe 3: Bulk-load a playlist into a knowledge base" pattern. Swap in your own channels, prompts and storage.

import requests

API = "https://www.youtubetranscript.dev/api/v2"
HEADERS = {"Authorization": "Bearer yt_sk_live_YOUR_KEY"}

playlist = requests.post(f"{API}/playlists/resolve", headers=HEADERS,
                         json={"playlist_url": "https://www.youtube.com/playlist?list=PLAYLIST_ID",
                               "limit": 100}).json()

ids = [v["video_id"] for v in playlist["items"]]
batch = requests.post(f"{API}/batch", headers=HEADERS, json={
    "video_ids": ids,
    "format": {"timestamp": True, "paragraphs": True},
}).json()

# Large batches run async: poll GET /batch/{batch_id} until status is "completed"
for r in batch.get("results", []):
    if r["status"] == "completed":
        for p in r["data"]["transcript"].get("paragraphs") or []:
            store(video_id=r["video_id"], start=p["start"], text=p["text"])  # your vector DB

Why it's worth building

For researchers the return is a reproducible dataset at a fraction of the cost and time of manual transcription. Teams also build paid datasets or analytics services for media monitoring firms and think tanks.

Mistakes to avoid

  • Mixing automatic and human captions without noting it: record the caption source because accuracy differs.
  • Not saving raw responses: store the original JSON so your analysis can be re-run.
  • Ignoring ethics review and platform terms: check your institution's requirements and YouTube's terms before publishing datasets.

Questions about this idea

Can I use YouTube transcripts for academic research?

Many researchers do. Check your institution's ethics requirements, YouTube's terms and copyright law, especially before sharing a dataset publicly.

How accurate are automatic YouTube captions?

They vary with audio quality, accents and topic. The API reports whether each transcript came from manual captions, automatic captions or AI speech recognition so you can account for it.

How many videos can I collect at once?

A single batch request accepts up to 3,000 video IDs on the Business plan. Larger studies can run several batches.

Start building today

10 free credits every month. No credit card needed.

How to Build a YouTube Transcript Corpus for Research | YouTubeTranscript.dev