AncherOpen Ancher →
YouTube video

Make a podcast

Turn a YouTube video into a podcast

A listenable version of something you would otherwise have to sit and watch.

No link handy? Try one:

Free to try, no card. Your YouTube video opens in Ancher.

Andrej Karpathy — Intro to LLMsY Combinator — How to get startup ideas

Your YouTube video

Paste a YouTube link. Nothing to upload.

Caption track with timingsKey frames from the videoDescription and linked sources

Kept in parts

Every one of these stays separately addressable instead of collapsing into one block of text.

ScriptSegmentsVoiceLength

Podcast

A narrated version you can listen to instead of read.

Some videos are worth an hour of your attention and some are worth twenty minutes of your commute. Ancher writes a narration script from the argument rather than the transcript, describes the on-screen material that a listener cannot see, marks the segment breaks where the subject actually turns, and puts the creator's credit in the script itself so the source travels with the audio.

A listener cannot see the chart, so the script has to say what was on it

The reason a read-aloud transcript is unlistenable is not the voice. It is that every reference to something visual becomes a hole — this graph here, look at what happens, the thing on the right. A script that works as audio has to convert what was shown into what can be said, and that is only possible if the frames were read in the first place.

One second of one video, at 14:03, read two ways.Illustrative. The figure is invented; the gap it falls into is not.

What a caption-only tool gets14:03
14:03

...and you can see right here what happens to the second line — — which is the whole point.

On the graph — not read

The specific never reaches the draft

What Ancher also gets14:03
Podcast On the graph Line B crosses at month 7 from 14:03

Lands in the draft, with its timecode

  • References to on-screen material are rewritten as description, using what was actually in the frame.
  • Segment breaks follow the turns in the argument, so the audio can be paused and resumed at a sensible place.
  • Credit to the original creator and a link to the video are written into the script, not left to a show-notes field.

Podcast, section by section

A narrated version you can listen to instead of read.

Script

Written to be spoken rather than read, with the visual references converted into description.

Video-to-podcast-script prompt

The link on its own gets a summary. This is the instruction that produces the 5 sections above, in that order, with the rules that keep them honest. Paste it with your YouTube video — in Ancher, or in whatever assistant you already use.

Write a narration script from this video. (1) Rewrite it to be spoken to a listener, not read from a transcript. (2) Every reference to something on screen must be converted into a description of what was actually shown. (3) Mark segment breaks where the argument turns. (4) Aim for roughly half the original runtime and list what you cut. (5) Write the creator credit and the source link into the script itself. Do not imitate the original speaker's voice or manner.
How each part of a YouTube video becomes a section
From the YouTube videoIntoWhy
Caption track with timingsScriptThe transcript is the raw material, but it is rewritten rather than read, because speech written for a viewer is not speech that works for a listener.
Key frames from the videoScriptEvery on-screen reference is a hole in audio, and the frame is the only way to fill it with something true.
Caption track with timingsSegmentsSubject changes in the timings mark the breaks, which is more reliable than cutting on a fixed interval.
Title, channel, duration, dateSource creditTitle, channel and date make the credit specific enough to be worth saying aloud.
Title, channel, duration, dateLengthThe source duration sets what a realistic target runtime is before any cutting decisions get made.
Everything Ancher keeps from a YouTube video

Caption track with timings

Published and auto-generated captions both come through, and every line keeps the second it was spoken.

Key frames from the video

Frames are sampled and scored, so a number shown on a slide survives even though it was never said out loud.

Description and linked sources

The creator's own cited links come along, so a claim can be traced past the video.

Likes and comment counts

Reception is a weak signal about correctness and a strong one about which claims got attention.

Title, channel, duration, date

Basic provenance, kept so the finished work can attribute the video properly.

The four steps
  1. Add the YouTube video. Paste a YouTube URL. Ancher pulls the caption track with its timings, samples key frames from the video itself, and keeps the description links and engagement counts on the same source.
  2. Nothing gets flattened. Caption track with timings, key frames from the video, description and linked sources, likes and comment counts, title, channel, duration, date stay separately addressable instead of collapsing into one block of text.
  3. Ask for a podcast. Ancher fills script, segments, voice, length, source credit using the mapping above, not a generic summary pass.
  4. Check it before you use it. Every specific keeps a link back to where it came from in the YouTube video, so a claim can be verified without reopening the source.
Where this goes wrong
  • This is a retelling of one person's video. Publishing it as your own show is a licensing problem, and the credit in the script does not by itself make redistribution permitted.
  • Content that is mostly demonstration does not survive the move to audio. If most of the value was on screen, the script will be description of a thing rather than the thing.
  • Anything read straight from a transcript will sound like a transcript. If the script still contains "as you can see here", it was not rewritten.
Questions people ask
Does it sound like the original creator?

No, and deliberately so. The script is written for a neutral reading voice rather than as an imitation of the speaker, because cloning someone's voice to retell their own work is a problem no amount of credit fixes.

How much shorter does it get?

Usually about half, because speech to camera carries a lot of repetition that a listener does not need repeated. The cuts are listed with the script, so you can put anything back that you disagree with.

Can I publish this?

Only with the creator's permission. What it is genuinely good for is a listenable version of something you need to absorb — a talk, a briefing, a long explainer — where the audience is you rather than a feed.

Bring your own YouTube videos

Everything you save lives in one workspace, so the podcast is built from your sources — not from a model's memory of the internet.

Open Ancher →

Not a YouTube video? The same podcast also comes from X thread, LinkedIn post, Instagram post.