BUILD GUIDE · RESEARCHERS

How to build a claim and misinformation tracker for YouTube

False and misleading claims often spread through spoken video before they appear in text. A claim tracker searches transcripts across many channels to show who said what, when it first appeared and how the wording changed, with links to verify every quote.

Who it's for
Fact-checkers, journalism schools, trust and safety researchers
Build time
1–2 weeks
API endpoints
/batch/transcripts/{video_id}

The problem it solves

Fact-checkers can search text platforms, but video is largely opaque. Claims get repeated with slightly different wording across dozens of channels, and documenting their spread means watching hours of footage. Researchers need searchable transcripts, semantic matching for paraphrases and a verifiable record of each occurrence.

What to include in your first version

  • Monitoring of a defined set of channels
  • Semantic search for paraphrases of a claim, not just exact wording
  • A timeline showing first appearance and spread
  • Quotes with timestamped links for verification
  • Exportable evidence for reports and publications

How to build it, step by step

  1. 01

    Collect transcripts

    Resolve the channels in scope and batch their videos for your time window. Keep monitoring new uploads if the study is ongoing.

  2. 02

    Embed the segments

    Create embeddings for each transcript paragraph so you can find statements with similar meaning.

  3. 03

    Search for the claim

    Embed several phrasings of the claim, retrieve similar passages and have a reviewer or an LLM confirm which ones actually make the claim.

  4. 04

    Build the timeline

    Order confirmed matches by publish date and export them with channel, quote and timestamped link.

Starter code

This idea is built on the "Recipe 2: Watch channels and alert on new uploads" pattern. Swap in your own channels, prompts and storage.

import requests

API = "https://www.youtubetranscript.dev/api/v2"
HEADERS = {"Authorization": "Bearer yt_sk_live_YOUR_KEY"}

WATCHLIST = ["@somechannel", "@anotherchannel"]
KEYWORDS = ["your brand", "competitor"]
seen = set()  # persist this in a database in production

for handle in WATCHLIST:
    latest = requests.post(f"{API}/channels/resolve", headers=HEADERS,
                           json={"handle": handle, "limit": 10}).json()
    for video in latest["items"]:
        if video["video_id"] in seen:
            continue
        seen.add(video["video_id"])
        t = requests.post(f"{API}/transcribe", headers=HEADERS,
                          json={"video": video["video_id"]}).json()
        text = t.get("data", {}).get("transcript", {}).get("text", "").lower()
        hits = [k for k in KEYWORDS if k in text]
        if hits:
            print(f"ALERT {video['title']}: {hits}")

Why it's worth building

Fact-checking organizations and research groups gain speed and verifiable evidence. Trust and safety teams and media monitoring companies will pay for tooling that covers video.

Mistakes to avoid

  • Treating automated matches as conclusions: always have a human confirm before publishing.
  • Searching only exact phrases: claims mutate, so use semantic search with several phrasings.
  • Losing provenance: keep the video ID, timestamp and retrieval date for every quote.

Questions about this idea

How do fact-checkers search YouTube videos?

By transcribing videos and searching the text. Semantic search finds paraphrases that keyword search would miss.

Can I prove when a claim was first made on YouTube?

You can document the earliest occurrence within the channels you monitor, with the publish date and timestamp. Coverage determines how complete that picture is.

What about videos without captions?

Enable AI speech recognition so videos without captions are still transcribed and searchable.

Start building today

10 free credits every month. No credit card needed.

How to Build a YouTube Claim and Misinformation Tracker | YouTubeTranscript.dev