frameworkyoutube-transcripts · v21099:codex-work-06ad2026-09-16served from databaseAll documents

FrameWork Youtube Transcripts

FrameWork Youtube Transcripts

For: every JBNX agent that needs the words from a YouTube video. Slug: youtube-transcripts Page: https://projects.jbnx.io/framework/youtube-transcripts Do not re-learn this from scratch. Follow the ladder. Stop when you have the captions.

Sister framework (what to do with a strategy transcript): FrameWork Youtube Viral Videos.


0. Outcome

A timestamped caption file covering the full runtime, plus the real title, channel, length, and chapter list. Auto-captions are ASR: names will be wrong. Keep them, note the caveats, do not "fix" quotes unless you watched that second.


1. Identify the video first (cheap, works from cloud IPs)

Cloud / datacenter IPs are treated as bots. The watch page and InnerTube player often return LOGIN_REQUIRED / "Sign in to confirm you're not a bot" with empty captions. That is not "no captions exist."

Do this first:

curl -sS "https://www.youtube.com/oembed?url=https://www.youtube.com/watch?v=VIDEO_ID&format=json"

You get title, author_name, author_url, thumbnail. Length is not in oembed.

Optional InnerTube that often returns metadata only (title, lengthSeconds, description + chapters) even when playability is UNPLAYABLE:

POST https://www.youtube.com/youtubei/v1/player
clientName: ANDROID_TESTSUITE / clientVersion: 1.9

Read videoDetails.title, lengthSeconds, shortDescription (chapters live at the bottom of the description on long Open Residency / podcast uploads).

Never trust a web-search snippet for which video an 11-character id is. Similar Open Residency titles collide (Kallaway vs Paddy Galloway). oembed / videoDetails wins.


2. What will fail from a Cloud Agent VM (do not loop)

Skip these once you have seen bot-block or empty timedtext. They burn the hour:

PathTypical failure
youtube-transcript-apiRequestBlocked (cloud IP)
yt-dlp without cookies + JS runtimeSign in to confirm you're not a bot
InnerTube ANDROID / IOS / TVHTML5 / WEB_EMBEDDEDLOGIN_REQUIRED or 400
https://www.youtube.com/api/timedtext?v=…&lang=enempty 200
Invidious / Piped public instances401/403/502 or shutdown
Jina / many "transcript generator" sitesCloudflare challenge
Watch-page WebFetchtimeout or bot interstitial

A second Cloud Agent will hit the same IP reputation. Do not spawn one to "verify."


3. What actually works (use in this order)

A. Browser "Show transcript" (primary)

Open https://www.youtube.com/watch?v=VIDEO_ID in a real browser (computer-use / headed Chrome). Accept cookies. Do not Google-sign-in.

  1. Expand description → Show transcript (or the transcript control).
  2. Select the transcript panel only and copy. Scroll so every chapter loads.
  3. Dump the raw clipboard to a file. Do not reformat timestamps yet.

YouTube glues the visible stamp to an accessibility label:

0:000 secondsWithin a day of changing…
0:1414 secondsVideos that have 10, 50…
1:041 minute, 4 secondsYouTube strategy.
2:44:252 hours, 44 minutes, 25 secondsbigger. So, I think…
Chapter 12: Ten Titles, Three Thumbnails, Every Video

That is:

Parser (copy this, do not reinvent):

import re
a11y = re.compile(
    r'^(?P<ts>\d{1,2}:\d{2}(?::\d{2})?)'
    r'(?:'
      r'(?:(?P<h>\d+) hours?, )?'
      r'(?:(?P<m>\d+) minutes?, )?'
      r'(?P<s>\d+) seconds?'
    r'|'
      r'(?:(?P<h2>\d+) hours?, )?'
      r'(?P<m2>\d+) minutes?'
    r')'
    r'(?P<text>.*)$'
)
ch_re = re.compile(r'^Chapter \d+:\s*(.+)$')

def norm_ts(ts: str) -> str:
    parts = [int(x) for x in ts.split(":")]
    if len(parts) == 2:
        mm, ss = parts
        return f"0:{mm:02d}:{ss:02d}"  # first hour is M:SS
    h, mm, ss = parts
    return f"{h}:{mm:02d}:{ss:02d}"

B. Podcast twin (when the upload is also an episode)

Many long YouTube interviews ship the same audio to Apple/Spotify the same day.

curl -sS "https://itunes.apple.com/search?term=VIDEO_TITLE&entity=podcastEpisode&limit=5"

You get feedUrl, episodeUrl / previewUrl (mp3). Use this when you need audio (Whisper) because captions never loaded — not as a substitute for captions when the panel worked.

Episode pages (example: https://openresidency.com/<guest-kebab>) give chapters even when YouTube is blocked.

C. Last resort

Export cookies from a human browser (yt-dlp --cookies) only if the operator offered them. Do not burn a Google account. Whisper on the podcast mp3 if captions truly do not exist (rare on long English uploads).


4. Package the delivery

Write three artifacts, then answer in chat:

  1. Raw dump (*-raw.txt) — clipboard, untouched.
  2. Clean file (*-final.txt) — [H:MM:SS] text plus ## Chapter headings.
  3. Identity block: title, channel, length, published date, chapter list, ASR caveat.

Do not commit a 40k-word transcript into the JBNX repo or public/framework/. The framework stays here; the transcript stays in chat / /tmp / artifacts.


5. Worked example (2026-09-14)

FieldValue
URLhttps://www.youtube.com/watch?v=Z2uoA3bhJT0
oembed title2-Hour Youtube Masterclass From The World's Highest-Paid Strategist
ChannelOpen Residency
GuestPaddy Galloway (ASR: "Patty")
Length9908 s = 2:45:08
Method that workedBrowser Show transcript → a11y parse
Captions1346 lines, 24 chapters, 0:00:00 → 2:45:01, ~38.8k words
Search trapWeb search matched a different Open Residency cut (Kallaway, VcqQmrGqthg). oembed prevented mixing them.

6. Agent rules that still apply

Transcript validation update — 2026-09-16

When the browser runtime advertises tab.content.exportYouTubeTranscript(), use that supported export on the identified watch page before manual panel copying. Verify the exported video ID, caption language, timestamp coverage and runtime. This worked for u_yvc7NTYvI (Dubibubi, 16:20; English ASR 0:00–16:17) after HTTP timedtext returned empty responses. Empty captions from one access path do not prove the video has no transcript.

Separate transcript-confirmed statements from independently verified technical facts. Preserve the difference between headline maxima and the actual demo result; do not validate an unseen screenshot from an AI description. See the completed rule review. Never publish the full copyrighted transcript as a framework.