BUILD GUIDE · RESEARCHERS

How to build a multilingual speech dataset from YouTube transcripts

Written text corpora underrepresent how people actually speak, especially in languages and dialects with little published text. YouTube holds conversational speech in nearly every language. Transcripts with word-level timing turn it into usable research data.

Who it's for
Computational linguists, NLP labs, AI evaluation teams
Build time
1 month
API endpoints
/transcripts/{video_id}/languagesasr_options.codeSwitchingformat.words

The problem it solves

Researchers studying spoken language, code-switching or low-resource languages struggle to find enough natural, conversational data. Many videos lack captions, and the ones that exist mix human and automatic sources of varying quality. Teams need a consistent pipeline that handles many languages and records where each transcript came from.

What to include in your first version

  • Detection of available caption languages per video
  • AI speech recognition in 99 languages for videos without captions
  • Code-switching support for speech that mixes languages
  • Word-level timestamps for alignment
  • Metadata on language and caption source for every transcript

How to build it, step by step

  1. 01

    Find candidate videos

    Resolve channels or playlists known to feature your target languages and call /transcripts/{video_id}/languages to see which caption tracks exist.

  2. 02

    Use captions where available

    Fetch manual captions where they exist, since they're usually the most accurate.

  3. 03

    Fill gaps with speech recognition

    For videos without captions, request ASR with the language set or auto-detected, and enable codeSwitching for mixed-language speech.

  4. 04

    Store with alignment data

    Request format.words for word-level timestamps and store the language and source of each transcript alongside the text.

Starter code

This idea is built on the "Recipe 3: Bulk-load a playlist into a knowledge base" pattern. Swap in your own channels, prompts and storage.

import requests

API = "https://www.youtubetranscript.dev/api/v2"
HEADERS = {"Authorization": "Bearer yt_sk_live_YOUR_KEY"}

playlist = requests.post(f"{API}/playlists/resolve", headers=HEADERS,
                         json={"playlist_url": "https://www.youtube.com/playlist?list=PLAYLIST_ID",
                               "limit": 100}).json()

ids = [v["video_id"] for v in playlist["items"]]
batch = requests.post(f"{API}/batch", headers=HEADERS, json={
    "video_ids": ids,
    "format": {"timestamp": True, "paragraphs": True},
}).json()

# Large batches run async: poll GET /batch/{batch_id} until status is "completed"
for r in batch.get("results", []):
    if r["status"] == "completed":
        for p in r["data"]["transcript"].get("paragraphs") or []:
            store(video_id=r["video_id"], start=p["start"], text=p["text"])  # your vector DB

Why it's worth building

Labs and AI teams gain evaluation and research data in languages that are expensive to collect otherwise. Clean, well-documented datasets for specific languages or domains are also valuable to model builders.

Mistakes to avoid

  • Assuming automatic captions are ground truth: label sources and validate a sample manually.
  • Mixing dialects without metadata: record the channel and region so results can be interpreted.
  • Overlooking licensing: check platform terms and copyright before redistributing any dataset.

Questions about this idea

How many languages does the speech recognition support?

AI speech recognition supports 99 languages, with automatic language detection and optional code-switching for mixed-language speech.

Can I get word-level timestamps?

Yes. Request format.words to receive timing for each word, which is useful for alignment and phonetic research.

Can I redistribute a dataset built from YouTube?

That depends on YouTube's terms, copyright law and your institution's policies. Many researchers share video IDs and code rather than the transcripts themselves.

Start building today

10 free credits every month. No credit card needed.

How to Build a Multilingual Speech Dataset from YouTube | YouTubeTranscript.dev