The problem it solves
Coaches and educators answer the same questions in comments and DMs every day, even though they've covered them in videos. Their audience can't find the right video, and the creator can't scale personal answers. A channel chatbot answers in the creator's voice and drives viewers back to the original content.
What to include in your first version
- Answers drawn only from the creator's own videos
- Citations that link to the exact video and timestamp
- Automatic updates when new videos are published
- A free tier with a question limit and a paid tier for members
- Embeddable widget for the creator's website or community
How to build it, step by step
- 01
Ingest the channel
Resolve the channel or a curated playlist and send all video IDs to /batch with format.paragraphs and format.timestamp.
- 02
Chunk and embed
Store each paragraph with its video title, URL and start time, and create embeddings in a vector database.
- 03
Retrieve and answer
For each question, retrieve the most relevant paragraphs and ask an LLM to answer using only those, citing the video and timestamp.
- 04
Stay current
Check the channel for new uploads on a schedule and add only the new transcripts to the index.
Starter code
This idea is built on the "Recipe 3: Bulk-load a playlist into a knowledge base" pattern. Swap in your own channels, prompts and storage.
import requests
API = "https://www.youtubetranscript.dev/api/v2"
HEADERS = {"Authorization": "Bearer yt_sk_live_YOUR_KEY"}
playlist = requests.post(f"{API}/playlists/resolve", headers=HEADERS,
json={"playlist_url": "https://www.youtube.com/playlist?list=PLAYLIST_ID",
"limit": 100}).json()
ids = [v["video_id"] for v in playlist["items"]]
batch = requests.post(f"{API}/batch", headers=HEADERS, json={
"video_ids": ids,
"format": {"timestamp": True, "paragraphs": True},
}).json()
# Large batches run async: poll GET /batch/{batch_id} until status is "completed"
for r in batch.get("results", []):
if r["status"] == "completed":
for p in r["data"]["transcript"].get("paragraphs") or []:
store(video_id=r["video_id"], start=p["start"], text=p["text"]) # your vector DBHow it makes money
Charge the creator a monthly fee for the hosted chatbot, or share revenue on a paid tier they sell to their audience. Membership communities and course sellers can bundle it as a premium perk.
Mistakes to avoid
- Letting the model answer from general knowledge: restrict it to retrieved transcript passages so it sounds like the creator.
- Using fixed-size chunks that cut sentences in half: paragraph groupings keep context intact.
- Launching without creator approval: build with the creator, since it's their voice and content.
Questions about this idea
Can I train a chatbot on a YouTube channel?
Yes. Rather than training a model, you retrieve relevant transcript passages and have an LLM answer from them. This is called retrieval-augmented generation.
How do I make the chatbot cite its sources?
Store the video URL and start time with every transcript chunk and instruct the model to include them in each answer.
How much does it cost to index a full channel?
Videos with native captions cost 1 credit each. A 300-video channel costs about 300 credits to index once, and only new uploads after that.