AncherOpen Ancher →
YouTube video

Make a transcript

Turn a YouTube video into a transcript

Paste a link. You get clean, timestamped text that stays searchable next to everything else you have saved.

No link handy? Try one:

Free to try, no card. Your YouTube video opens in Ancher.

Andrej Karpathy — Intro to LLMsY Combinator — How to get startup ideas

Your YouTube video

Paste a YouTube link. Nothing to upload.

Caption track with timingsKey frames from the videoDescription and linked sources

Kept in parts

Every one of these stays separately addressable instead of collapsing into one block of text.

Full textTimestampsFiller removedExport formats

Transcript

The spoken words, cleaned up, with a timestamp on every segment.

Ancher pulls the caption track where one exists and transcribes the audio where it does not, then strips the filler and keeps a timestamp on every segment. What separates it from the dozens of free transcript sites is where the text lands: in a workspace alongside your other sources, so a phrase you half-remember is findable across forty videos rather than in one downloaded file you will never open again.

A transcript in a text file is a transcript you will lose

Getting the words out of a video is close to a solved problem — the free tools do it fine. The part nobody solves is what happens next. A .txt on your desktop is unsearchable in practice, disconnected from every other thing you saved on the subject, and useless six weeks later when you can only remember that someone, somewhere, said the thing.

One second of one video, at 09:33, read two ways.Illustrative. The figure is invented; the gap it falls into is not.

What a caption-only tool gets09:33
09:33

...the number we ended up shipping against was and it held all quarter.

Shown on the slide, never said — not read

The specific never reaches the draft

What Ancher also gets09:33
Transcript Shown on the slide, never said p95 = 240ms from 09:33

Lands in the draft, with its timecode

  • Transcripts land in your workspace, so one search crosses every video, article and PDF you have saved rather than one file.
  • Timestamps stay attached, so a line found by search is one click from the moment it was said.
  • Text that only ever appeared on screen is captured separately, so the record is not limited to what was spoken aloud.

Transcript, section by section

The spoken words, cleaned up, with a timestamp on every segment.

Full text

Clean running text with the crutch words removed, so it reads as prose rather than as a dump of everything the microphone caught.

YouTube transcript prompt

The link on its own gets a summary. This is the instruction that produces the 5 sections above, in that order, with the rules that keep them honest. Paste it with your YouTube video — in Ancher, or in whatever assistant you already use.

Produce a clean transcript of this video. (1) Use the published caption track if one exists; otherwise transcribe the audio and say so. (2) Remove filler and false starts, but keep hedges such as "I think" or "roughly" — they change the claim. (3) Keep a start timestamp on every segment. (4) Mark any text that appeared on screen but was never spoken as ON-SCREEN. (5) Flag names, acronyms and figures you are unsure of rather than guessing. Do not summarise, condense or reorder anything.
How each part of a YouTube video becomes a section
From the YouTube videoIntoWhy
Caption track with timingsFull textA published caption track beats machine transcription on proper nouns, so Ancher uses it whenever the creator provided one.
Caption track with timingsTimestampsCaption cues already carry timings, so the timestamps come from the source rather than being estimated after the fact.
Key frames from the videoFiller removedText visible on screen is marked as shown rather than spoken, so a reader knows which lines were never actually said.
Title, channel, duration, dateExport formatsChannel, title and date ride along with the export, so a transcript pasted into a doc still knows where it came from.
Description and linked sourcesSearchable in your workspaceDescription links are indexed with the transcript, so searching a company name also finds the page the creator was citing.
Everything Ancher keeps from a YouTube video

Caption track with timings

Published and auto-generated captions both come through, and every line keeps the second it was spoken.

Key frames from the video

Frames are sampled and scored, so a number shown on a slide survives even though it was never said out loud.

Description and linked sources

The creator's own cited links come along, so a claim can be traced past the video.

Likes and comment counts

Reception is a weak signal about correctness and a strong one about which claims got attention.

Title, channel, duration, date

Basic provenance, kept so the finished work can attribute the video properly.

The four steps
  1. Add the YouTube video. Paste a YouTube URL. Ancher pulls the caption track with its timings, samples key frames from the video itself, and keeps the description links and engagement counts on the same source.
  2. Nothing gets flattened. Caption track with timings, key frames from the video, description and linked sources, likes and comment counts, title, channel, duration, date stay separately addressable instead of collapsing into one block of text.
  3. Ask for a transcript. Ancher fills full text, timestamps, filler removed, export formats, searchable in your workspace using the mapping above, not a generic summary pass.
  4. Check it before you use it. Every specific keeps a link back to where it came from in the YouTube video, so a claim can be verified without reopening the source.
Where this goes wrong
  • Auto-generated captions are unreliable on names, acronyms and numbers. For anything you plan to quote, check the line against the frame before it leaves the workspace.
  • There are no speaker labels. An interview or a panel comes back as one continuous voice, so attribution has to be added by hand.
  • A transcript is not a summary and should not be read as one. Length is not a proxy for substance, and a two-hour transcript rarely contains two hours of usable material.
Questions people ask
Why not use one of the free transcript sites?

For a single one-off transcript, use one — the extraction is comparable. The difference shows up on the fortieth video, when you need to find a phrase and cannot remember which file it was in. Ancher's transcripts are searchable together, which is the only part that compounds.

Does it work when captions are turned off?

Yes. Ancher transcribes the audio directly in that case. Accuracy is generally lower than a creator-supplied caption track, particularly on technical vocabulary, so treat names and figures as needing a check.

Can I get subtitles out of it?

The export includes SRT with the original cue timings, so it can go straight into a video editor. That is also the format to choose if you plan to re-time or translate it later.

Bring your own YouTube videos

Everything you save lives in one workspace, so the transcript is built from your sources — not from a model's memory of the internet.

Open Ancher →

Not a YouTube video? The same transcript also comes from Instagram post, TikTok video, recording.