BUILD GUIDE · APP DEVELOPERS

How to build a multilingual subtitle pipeline for YouTube videos

Most viewers of a popular video don't speak its original language. A subtitle pipeline that translates transcripts and exports timed subtitle files helps creators and course platforms reach them without hiring translators for every upload.

Who it's for
Localization tools, international creators, e-learning platforms
Build time
Weekend
API endpoints
/transcripts/{video_id}/languages/transcripts/{video_id}/translate

The problem it solves

Professional subtitle translation is slow and priced per minute, so most creators publish in one language only. YouTube's auto-translate is inconsistent and can't be exported or edited. Teams need editable, timed subtitle files in several languages that they can review and upload themselves.

What to include in your first version

  • Detection of caption languages that already exist
  • Translation into many target languages at once
  • SRT and VTT export with original timings preserved
  • A side-by-side editor for reviewing translations
  • A glossary for brand names and technical terms

How to build it, step by step

  1. 01

    Check existing languages

    Call /transcripts/{video_id}/languages to see which caption tracks already exist so you don't pay to translate what's available.

  2. 02

    Fetch the source transcript

    Call /transcribe with format.timestamp to get segments with start and end times.

  3. 03

    Translate

    Call /transcripts/{video_id}/translate for each target language. Segment timings stay aligned with the original.

  4. 04

    Export subtitle files

    Write each segment as a numbered SRT block or a VTT cue using its start and end times, then let users download or upload them to YouTube.

Starter code

This idea is built on the "Recipe 1: One video in, structured output out" pattern. Swap in your own channels, prompts and storage.

import requests

API = "https://www.youtubetranscript.dev/api/v2"
HEADERS = {"Authorization": "Bearer yt_sk_live_YOUR_KEY"}

res = requests.post(f"{API}/transcribe", headers=HEADERS, json={
    "video": "https://www.youtube.com/watch?v=VIDEO_ID",
    "format": {"timestamp": True, "paragraphs": True},
}).json()

transcript = res["data"]["transcript"]
print(transcript["text"][:500])        # plain text for your LLM prompt
for seg in transcript["segments"][:3]:  # timestamps for deep links
    print(seg["start"], seg["text"])

Why it's worth building

Price per video minute per language, or by monthly subscription tiers. E-learning platforms and agencies that localize at volume are the best customers because the value grows with every language added.

Mistakes to avoid

  • Translating line by line without context, which breaks sentences that span several subtitle cues.
  • Ignoring reading speed: translated text can run longer than the original, so check characters per second.
  • Paying to translate languages that already have human-made captions.

Questions about this idea

Can I download translated YouTube subtitles as SRT?

Yes. Translate the transcript through the API, then write each timed segment out in SRT or VTT format.

How much does transcript translation cost?

Translation costs about 1 credit per 2,500 characters of transcript text. Fetching native captions costs 1 credit per video.

Can this be used for dubbing?

Yes. The translated, timed script is the input most AI voice and dubbing tools need.

Start building today

10 free credits every month. No credit card needed.

How to Build a YouTube Subtitle Translator (SRT and VTT) | YouTubeTranscript.dev