BUILD GUIDE · APP DEVELOPERS

How to give your AI agent the ability to watch YouTube

Language models can't watch video, and most can't fetch YouTube transcripts reliably on their own. Give them a transcript tool and they can summarize a talk, compare two reviews or answer questions about a lecture, with citations back to the moment it was said.

Who it's for
Agent builders, internal copilots, research assistants, Slack and Discord bots
Build time
Afternoon
API endpoints
/transcribeMCP server

The problem it solves

When users paste a YouTube link into an assistant, it either refuses, hallucinates from the title, or fails on long videos. Teams building agents for research, sales or support need a dependable way to turn any video into text the model can reason over, without maintaining scrapers that break whenever YouTube changes.

What to include in your first version

  • A tool the model can call with any YouTube URL or video ID
  • Clean transcript text plus timestamped segments for citations
  • Language selection and translation for non-English videos
  • Chunking for long videos so they fit the model's context
  • Caching so the same video isn't fetched twice

How to build it, step by step

  1. 01

    Choose your integration

    If your assistant supports the Model Context Protocol, connect the MCP server and the tools appear automatically. Otherwise define a function tool that calls /transcribe.

  2. 02

    Define the tool contract

    Accept a video URL and optional language. Return the transcript text and a list of segments with start times.

  3. 03

    Handle long videos

    For videos longer than your context window, split segments into chunks and summarize each before answering, or retrieve only the chunks relevant to the question.

  4. 04

    Cite the source

    Instruct the model to quote timestamps from the segments so users can verify every claim.

Starter code

This idea is built on the "Recipe 1: One video in, structured output out" pattern. Swap in your own channels, prompts and storage.

import requests

API = "https://www.youtubetranscript.dev/api/v2"
HEADERS = {"Authorization": "Bearer yt_sk_live_YOUR_KEY"}

res = requests.post(f"{API}/transcribe", headers=HEADERS, json={
    "video": "https://www.youtube.com/watch?v=VIDEO_ID",
    "format": {"timestamp": True, "paragraphs": True},
}).json()

transcript = res["data"]["transcript"]
print(transcript["text"][:500])        # plain text for your LLM prompt
for seg in transcript["segments"][:3]:  # timestamps for deep links
    print(seg["start"], seg["text"])

Why it's worth building

This is usually a feature, not a product: it makes your existing assistant, bot or copilot noticeably more useful. As a standalone product, a Slack or Discord bot that answers questions about shared videos can charge per workspace.

Mistakes to avoid

  • Stuffing a three-hour transcript into one prompt: retrieve or summarize in chunks instead.
  • Dropping timestamps, which makes answers impossible to verify.
  • Fetching the same popular video repeatedly instead of caching transcripts.

Questions about this idea

Can ChatGPT or Claude read YouTube videos?

Not reliably on their own. Connecting a transcript tool or MCP server lets them read the full transcript of any public video with captions.

What is an MCP server?

The Model Context Protocol is a standard way to give AI assistants tools. Our MCP server exposes transcript tools that compatible assistants can call without custom code.

How do I handle very long videos?

Split the timestamped segments into chunks, then summarize each chunk or retrieve only the chunks relevant to the user's question.

Start building today

10 free credits every month. No credit card needed.

How to Give an AI Agent Access to YouTube Videos (MCP and Tools) | YouTubeTranscript.dev