An agent skill that turns a YouTube video — or any web page linking to one (podcast episode pages like Lenny's Newsletter, newsletter posts, articles with embedded players) — into a clean, timestamped Markdown transcript, with optional paragraph-by-paragraph bilingual translation.
Works with any agent that supports the .agents/skills/ convention (ZCode, Claude Code, Cursor, Codex, and more).
Share a link and ask for the transcript. The agent finds the video's public YouTube captions, converts them into ~30-second timestamped paragraphs, asks you to confirm the target language, and produces an interleaved bilingual file:
[00:25] into a cave for about a month with the sole objective of build an amazing
knowledge work product that brings agents to the rest of the company. >> It's been
only 3 weeks since launch...
[00:25] 闭关大约一个月,唯一目标就是打造一款出色的知识工作产品,把 agent 带给
公司里更多的人。>> 发布才三周……
Both files are kept: the original-language transcript and the bilingual version (each original paragraph followed by its translation under the same timestamp).
# via the skills.sh CLI (project scope)
npx skills add SherlockShemol/youtube-transcript
# global (available in all projects)
npx skills add -g SherlockShemol/youtube-transcript
# or manually
git clone https://github.com/SherlockShemol/youtube-transcript ~/.agents/skills/youtube-transcriptNote: the skills CLI collects anonymous install telemetry; set DISABLE_TELEMETRY=1 to opt out.
-
Python 3.8+ — the two helper scripts use only the standard library.
-
yt-dlp, latest version — this is the #1 gotcha: old yt-dlp versions fail against YouTube's bot check ("Sign in to confirm you're not a bot" / HTTP 429). Install current:
uv tool install yt-dlp # or: pipx install yt-dlp / pip install --user yt-dlp
page URL (or YouTube URL)
→ find the YouTube version (pages' "Listen on" links, or web search)
→ yt-dlp downloads the <lang>-orig auto-caption track (video itself is skipped)
→ scripts/json3_to_transcript.py → timestamped Markdown
→ agent asks which language to translate into (skip is an option)
→ agent writes only the translations, one line per paragraph
→ scripts/merge_bilingual.py interleaves original + translation
The model never retypes the original text while translating — a merge script enforces a strict 1:1 paragraph↔translation count and refuses to write output on mismatch, so long transcripts can't silently drop or misalign paragraphs.
- Requires the content to have a public YouTube version with captions. Spotify/Apple Podcasts links don't expose downloadable captions.
- Auto-generated captions have no punctuation or speaker names;
>>marks speaker changes. - Translation is AI-generated, paragraph by paragraph.
- Paywalled transcripts (e.g., paid Substack posts) are never bypassed — the skill uses the same audio's public YouTube captions instead.