The problem it solves
Large video libraries are effectively unsearchable. Fans scrub through hours of footage to find a moment they half remember, course students rewatch whole lessons to find one explanation, and companies lose institutional knowledge buried in recorded talks. Indexing transcripts with timestamps turns every library into something you can search like a document, with results that jump straight to the right moment.
What to include in your first version
- Phrase and keyword search across every video in a channel or playlist
- Results that show the matching sentence with a link to the exact timestamp
- Semantic search so "how do I price my product" also finds "setting your rates"
- Filters by date, video and speaker
- Automatic indexing of new uploads
How to build it, step by step
- 01
Collect the video list
Call /channels/resolve or /playlists/resolve with the channel handle or playlist URL to get every video ID and title.
- 02
Fetch transcripts in bulk
Send the IDs to /batch with format.timestamp and format.paragraphs enabled. Large batches run asynchronously, so poll /batch/{batch_id} or use a webhook.
- 03
Index the segments
Store each paragraph with its video ID and start time. Postgres full-text search is enough for keyword search; add embeddings in pgvector or a vector database for meaning-based search.
- 04
Build the search UI
Return the top matches with surrounding text and a link of the form youtube.com/watch?v=ID&t=SECONDS so users land on the exact moment.
- 05
Keep it fresh
Run a scheduled job that resolves the channel again, finds new video IDs and transcribes only those.
Starter code
This idea is built on the "Recipe 3: Bulk-load a playlist into a knowledge base" pattern. Swap in your own channels, prompts and storage.
import requests
API = "https://www.youtubetranscript.dev/api/v2"
HEADERS = {"Authorization": "Bearer yt_sk_live_YOUR_KEY"}
playlist = requests.post(f"{API}/playlists/resolve", headers=HEADERS,
json={"playlist_url": "https://www.youtube.com/playlist?list=PLAYLIST_ID",
"limit": 100}).json()
ids = [v["video_id"] for v in playlist["items"]]
batch = requests.post(f"{API}/batch", headers=HEADERS, json={
"video_ids": ids,
"format": {"timestamp": True, "paragraphs": True},
}).json()
# Large batches run async: poll GET /batch/{batch_id} until status is "completed"
for r in batch.get("results", []):
if r["status"] == "completed":
for p in r["data"]["transcript"].get("paragraphs") or []:
store(video_id=r["video_id"], start=p["start"], text=p["text"]) # your vector DBWhy it's worth building
Sell it to the owner of the library rather than the viewer: podcast networks, course platforms and companies with internal video archives will pay a monthly fee for search that keeps their audience engaged. For fan communities, a free public search with a paid tier for alerts or API access also works.
Mistakes to avoid
- Indexing whole transcripts as one document: search results become useless without paragraph-level timestamps.
- Re-transcribing the entire channel on every update instead of tracking which video IDs you've already processed.
- Relying only on exact keyword match: spoken language is messy, so add semantic search early.
Questions about this idea
Do I need a vector database to search YouTube transcripts?
No. Postgres full-text search handles keyword search well for most libraries. Add embeddings when you want results that match meaning rather than exact words.
How do I link a search result to the exact moment in the video?
Each transcript segment includes a start time in seconds. Append &t= followed by that number to the YouTube URL.
Can I index videos that don't have captions?
Yes. Enable allow_asr with a webhook URL and the API transcribes the audio with AI speech recognition when captions are missing.