The problem it solves
Fact-checkers can search text platforms, but video is largely opaque. Claims get repeated with slightly different wording across dozens of channels, and documenting their spread means watching hours of footage. Researchers need searchable transcripts, semantic matching for paraphrases and a verifiable record of each occurrence.
What to include in your first version
- Monitoring of a defined set of channels
- Semantic search for paraphrases of a claim, not just exact wording
- A timeline showing first appearance and spread
- Quotes with timestamped links for verification
- Exportable evidence for reports and publications
How to build it, step by step
- 01
Collect transcripts
Resolve the channels in scope and batch their videos for your time window. Keep monitoring new uploads if the study is ongoing.
- 02
Embed the segments
Create embeddings for each transcript paragraph so you can find statements with similar meaning.
- 03
Search for the claim
Embed several phrasings of the claim, retrieve similar passages and have a reviewer or an LLM confirm which ones actually make the claim.
- 04
Build the timeline
Order confirmed matches by publish date and export them with channel, quote and timestamped link.
Starter code
This idea is built on the "Recipe 2: Watch channels and alert on new uploads" pattern. Swap in your own channels, prompts and storage.
import requests
API = "https://www.youtubetranscript.dev/api/v2"
HEADERS = {"Authorization": "Bearer yt_sk_live_YOUR_KEY"}
WATCHLIST = ["@somechannel", "@anotherchannel"]
KEYWORDS = ["your brand", "competitor"]
seen = set() # persist this in a database in production
for handle in WATCHLIST:
latest = requests.post(f"{API}/channels/resolve", headers=HEADERS,
json={"handle": handle, "limit": 10}).json()
for video in latest["items"]:
if video["video_id"] in seen:
continue
seen.add(video["video_id"])
t = requests.post(f"{API}/transcribe", headers=HEADERS,
json={"video": video["video_id"]}).json()
text = t.get("data", {}).get("transcript", {}).get("text", "").lower()
hits = [k for k in KEYWORDS if k in text]
if hits:
print(f"ALERT {video['title']}: {hits}")Why it's worth building
Fact-checking organizations and research groups gain speed and verifiable evidence. Trust and safety teams and media monitoring companies will pay for tooling that covers video.
Mistakes to avoid
- Treating automated matches as conclusions: always have a human confirm before publishing.
- Searching only exact phrases: claims mutate, so use semantic search with several phrasings.
- Losing provenance: keep the video ID, timestamp and retrieval date for every quote.
Questions about this idea
How do fact-checkers search YouTube videos?
By transcribing videos and searching the text. Semantic search finds paraphrases that keyword search would miss.
Can I prove when a claim was first made on YouTube?
You can document the earliest occurrence within the channels you monitor, with the publish date and timestamp. Coverage determines how complete that picture is.
What about videos without captions?
Enable AI speech recognition so videos without captions are still transcribed and searchable.