The problem it solves
When users paste a YouTube link into an assistant, it either refuses, hallucinates from the title, or fails on long videos. Teams building agents for research, sales or support need a dependable way to turn any video into text the model can reason over, without maintaining scrapers that break whenever YouTube changes.
What to include in your first version
- A tool the model can call with any YouTube URL or video ID
- Clean transcript text plus timestamped segments for citations
- Language selection and translation for non-English videos
- Chunking for long videos so they fit the model's context
- Caching so the same video isn't fetched twice
How to build it, step by step
- 01
Choose your integration
If your assistant supports the Model Context Protocol, connect the MCP server and the tools appear automatically. Otherwise define a function tool that calls /transcribe.
- 02
Define the tool contract
Accept a video URL and optional language. Return the transcript text and a list of segments with start times.
- 03
Handle long videos
For videos longer than your context window, split segments into chunks and summarize each before answering, or retrieve only the chunks relevant to the question.
- 04
Cite the source
Instruct the model to quote timestamps from the segments so users can verify every claim.
Starter code
This idea is built on the "Recipe 1: One video in, structured output out" pattern. Swap in your own channels, prompts and storage.
import requests
API = "https://www.youtubetranscript.dev/api/v2"
HEADERS = {"Authorization": "Bearer yt_sk_live_YOUR_KEY"}
res = requests.post(f"{API}/transcribe", headers=HEADERS, json={
"video": "https://www.youtube.com/watch?v=VIDEO_ID",
"format": {"timestamp": True, "paragraphs": True},
}).json()
transcript = res["data"]["transcript"]
print(transcript["text"][:500]) # plain text for your LLM prompt
for seg in transcript["segments"][:3]: # timestamps for deep links
print(seg["start"], seg["text"])Why it's worth building
This is usually a feature, not a product: it makes your existing assistant, bot or copilot noticeably more useful. As a standalone product, a Slack or Discord bot that answers questions about shared videos can charge per workspace.
Mistakes to avoid
- Stuffing a three-hour transcript into one prompt: retrieve or summarize in chunks instead.
- Dropping timestamps, which makes answers impossible to verify.
- Fetching the same popular video repeatedly instead of caching transcripts.
Questions about this idea
Can ChatGPT or Claude read YouTube videos?
Not reliably on their own. Connecting a transcript tool or MCP server lets them read the full transcript of any public video with captions.
What is an MCP server?
The Model Context Protocol is a standard way to give AI assistants tools. Our MCP server exposes transcript tools that compatible assistants can call without custom code.
How do I handle very long videos?
Split the timestamped segments into chunks, then summarize each chunk or retrieve only the chunks relevant to the user's question.