The problem it solves
Influencer teams vet creators by watching a handful of recent videos. They miss a competitor sponsorship from six months ago, a controversial remark, or a habit of mocking sponsors on camera. Discovering this after a campaign launches is expensive. Reviewing hundreds of hours manually is impossible.
What to include in your first version
- Bulk analysis of a creator's recent videos
- Detection of past sponsors and competitor brands
- Content-safety flags for profanity, sensitive topics and controversy
- Sentiment toward sponsors and product categories
- A shareable report with quotes and timestamped evidence
How to build it, step by step
- 01
Pull the back catalogue
Call /channels/resolve with the creator's handle and a limit of 100 to 500 to get recent video IDs.
- 02
Transcribe in one batch
Send the IDs to /batch. For videos without captions, ASR options such as content moderation and sentiment analysis add structured safety signals.
- 03
Extract sponsors and risks
Run an LLM over each transcript to list sponsors mentioned, competitor brands and statements that match the brand's risk categories.
- 04
Score and report
Combine findings into a score per category and generate a report where every flag links to the quote and timestamp.
Starter code
This idea is built on the "Recipe 3: Bulk-load a playlist into a knowledge base" pattern. Swap in your own channels, prompts and storage.
import requests
API = "https://www.youtubetranscript.dev/api/v2"
HEADERS = {"Authorization": "Bearer yt_sk_live_YOUR_KEY"}
playlist = requests.post(f"{API}/playlists/resolve", headers=HEADERS,
json={"playlist_url": "https://www.youtube.com/playlist?list=PLAYLIST_ID",
"limit": 100}).json()
ids = [v["video_id"] for v in playlist["items"]]
batch = requests.post(f"{API}/batch", headers=HEADERS, json={
"video_ids": ids,
"format": {"timestamp": True, "paragraphs": True},
}).json()
# Large batches run async: poll GET /batch/{batch_id} until status is "completed"
for r in batch.get("results", []):
if r["status"] == "completed":
for p in r["data"]["transcript"].get("paragraphs") or []:
store(video_id=r["video_id"], start=p["start"], text=p["text"]) # your vector DBHow it makes money
Charge per report for brands that vet occasionally, and sell seat-based plans to agencies and creator marketplaces that vet every week. The price is easy to justify against the cost of one failed sponsorship.
Mistakes to avoid
- Flagging without evidence: every risk needs a quote and timestamp so a human can judge context.
- Treating all profanity as equally risky: let each brand set its own thresholds.
- Analyzing only recent videos: competitor deals often appear further back.
Questions about this idea
How do brands vet YouTube influencers?
Most watch a few videos manually. Transcript analysis lets you review a creator's entire recent catalogue for sponsors, sensitive topics and tone in minutes.
Can I detect which sponsors a YouTuber has worked with?
Yes. Sponsorship reads are spoken, so an LLM can extract brand names and ad segments from transcripts.
How many videos should a vetting report cover?
A useful default is the last 100 videos or 12 months, whichever is larger. Batch requests can include up to 3,000 videos on the Business plan.