The problem it solves
Researchers studying spoken language, code-switching or low-resource languages struggle to find enough natural, conversational data. Many videos lack captions, and the ones that exist mix human and automatic sources of varying quality. Teams need a consistent pipeline that handles many languages and records where each transcript came from.
What to include in your first version
- Detection of available caption languages per video
- AI speech recognition in 99 languages for videos without captions
- Code-switching support for speech that mixes languages
- Word-level timestamps for alignment
- Metadata on language and caption source for every transcript
How to build it, step by step
- 01
Find candidate videos
Resolve channels or playlists known to feature your target languages and call /transcripts/{video_id}/languages to see which caption tracks exist.
- 02
Use captions where available
Fetch manual captions where they exist, since they're usually the most accurate.
- 03
Fill gaps with speech recognition
For videos without captions, request ASR with the language set or auto-detected, and enable codeSwitching for mixed-language speech.
- 04
Store with alignment data
Request format.words for word-level timestamps and store the language and source of each transcript alongside the text.
Starter code
This idea is built on the "Recipe 3: Bulk-load a playlist into a knowledge base" pattern. Swap in your own channels, prompts and storage.
import requests
API = "https://www.youtubetranscript.dev/api/v2"
HEADERS = {"Authorization": "Bearer yt_sk_live_YOUR_KEY"}
playlist = requests.post(f"{API}/playlists/resolve", headers=HEADERS,
json={"playlist_url": "https://www.youtube.com/playlist?list=PLAYLIST_ID",
"limit": 100}).json()
ids = [v["video_id"] for v in playlist["items"]]
batch = requests.post(f"{API}/batch", headers=HEADERS, json={
"video_ids": ids,
"format": {"timestamp": True, "paragraphs": True},
}).json()
# Large batches run async: poll GET /batch/{batch_id} until status is "completed"
for r in batch.get("results", []):
if r["status"] == "completed":
for p in r["data"]["transcript"].get("paragraphs") or []:
store(video_id=r["video_id"], start=p["start"], text=p["text"]) # your vector DBWhy it's worth building
Labs and AI teams gain evaluation and research data in languages that are expensive to collect otherwise. Clean, well-documented datasets for specific languages or domains are also valuable to model builders.
Mistakes to avoid
- Assuming automatic captions are ground truth: label sources and validate a sample manually.
- Mixing dialects without metadata: record the channel and region so results can be interpreted.
- Overlooking licensing: check platform terms and copyright before redistributing any dataset.
Questions about this idea
How many languages does the speech recognition support?
AI speech recognition supports 99 languages, with automatic language detection and optional code-switching for mixed-language speech.
Can I get word-level timestamps?
Yes. Request format.words to receive timing for each word, which is useful for alignment and phonetic research.
Can I redistribute a dataset built from YouTube?
That depends on YouTube's terms, copyright law and your institution's policies. Many researchers share video IDs and code rather than the transcripts themselves.