Your Instagram post
Paste your Instagram post. Nothing to upload.
Make a transcript
The spoken words, plus the words that were only ever on the screen.
Free to try, no card. Your Instagram post opens in Ancher.
Paste your Instagram post. Nothing to upload.
Every one of these stays separately addressable instead of collapsing into one block of text.
The spoken words, cleaned up, with a timestamp on every segment.
A reel transcript that contains only speech is missing half the reel, because creators put the number on the screen and say something vaguer out loud. Ancher transcribes the audio with timings and separately recovers the text burned into the frames, so a claim that exists only as an on-screen caption still lands in the transcript instead of disappearing.
The honest version
The format's convention is to say the emotional line and show the factual one. Speech carries the delivery; the overlay carries the price, the dosage, the timeline. A speech-only transcript is therefore a document of the delivery, which is the half nobody needed, and it looks complete enough that the gap is easy to miss.
One post, at 0:14, read two ways.Illustrative. The figure is invented; the gap it falls into is not.
and the third one is the one that actually matters, which is
Burned into the reel — not readThe specific never reaches the draft
Lands in the draft, with its timecode
What comes out
The spoken words, cleaned up, with a timestamp on every segment.
The spoken words, cleaned of filler, with on-screen text interleaved and labelled.
A timing on every segment, so any line can be found in the reel in seconds.
Filler and repetition removed, with a note on what was cut so nothing looks like a paraphrase.
Plain text you can paste, rather than a viewer you have to keep open.
Sits with your other sources, so a search crosses this reel and everything near it.
The link on its own gets a summary. This is the instruction that produces the 5 sections above, in that order, with the rules that keep them honest. Paste it with your Instagram post — in Ancher, or in whatever assistant you already use.
Transcribe this reel. (1) Give the spoken words cleaned of filler, with a timing on every segment. (2) Separately recover any text burned into the video and interleave it, clearly marked ON-SCREEN rather than spoken. (3) Note what you removed in cleanup. (4) Flag any number that came from audio alone as unverified against the frame. (5) Do not attribute lines to speakers — this source has no speaker labels.
The detail, if you want it
| From the Instagram post | Into | Why |
|---|---|---|
| Reel audio, transcribed | Full text | Spoken-over reels carry detail that appears nowhere in the caption, and the audio is the only route to it. |
| Frames from the video | Full text | Text burned into a reel is invisible to a caption scrape but survives frame sampling, and it is usually the specific. |
| Reel audio, transcribed | Timestamps | Segment timings are what make a fifteen-second claim findable inside a ninety-second reel. |
| Caption | Filler removed | The caption sets the subject, which is what tells you whether a garbled word was a term of art or noise. |
| Frames from the video | Searchable in your workspace | On-screen text indexed alongside speech is what lets a search find a reel by a number that was never said aloud. |
The caption is where the actual claim lives; the visual is usually the hook.
Carousels are decks in disguise — slide seven is often the one with the substance.
Spoken-over reels carry detail that appears nowhere in the caption.
Text burned into a reel is invisible to a caption scrape but survives frame sampling.
Reach tells you which format the audience actually responded to.
Because on this platform it is where the facts are. A creator says "it costs about this much" out loud and types the actual figure on the screen, so speech-only output loses the number and keeps the hedge.
Good on clear speech to camera and worse with a music bed or rapid delivery, which is most of the format. Anything numeric should be checked against the frame, and that is why both are in the output.
You get timed segments, which is the substance of a subtitle file. What it will not do is separate two speakers, so a conversational reel comes back as one voice regardless of how many people are in it.
Everything you save lives in one workspace, so the transcript is built from your sources — not from a model's memory of the internet.
Open Ancher →Not an Instagram post? The same transcript also comes from YouTube video, TikTok video, recording.