What the TXT file contains
The download starts with the video title and channel name, a "--- Transcript ---" divider, then one caption line per row. Nothing else: no HTML, no styling, no cue numbers. That makes it the easiest format to read on any device, paste into a document, or hand to another tool.
You choose one of two versions from the Download menu. "With timestamps" puts a [MM:SS] marker (or [HH:MM:SS] past the first hour) in front of each line. "Without timestamps" keeps only the spoken text.
Sample TXT output, with timestamps (original example text, not from a real video)
Example Video Title Channel: Example Channel --- Transcript --- [00:00] First, write down the question you want the video to answer. [00:04] Next, search the transcript for a distinctive phrase. [00:08] Finally, check the matching moment in the source video.
With or without timestamps?
- Keep timestamps when you will need to go back to the video: checking a quote, writing show notes, or citing a lecture.
- Drop them when the text is the end product: reading, editing into an article, or pasting into an AI chat where markers only take up space.
- Each line is still one caption segment either way. Caption segments are short, so the text reads as many short lines rather than paragraphs. Use the Transcript Cleaner if you want merged paragraphs.
Things to check before you use the text
- Auto-generated captions often miss punctuation and misspell names, brands, and technical terms. Skim for these before publishing or quoting.
- Sound cues such as [Music] or [Applause] come through exactly as YouTube captioned them.
- The file is UTF-8, so non-Latin scripts (Arabic, Hindi, Japanese, and others) are preserved as-is.
- The export reflects the caption language you selected. If a video has several caption tracks, pick the one you want before downloading.