
Transcript export formats: which file to choose
You have a transcript, and now you have to decide which file to take it away in. The words are the same in every one; what changes is the wrapping, and the wrapping is what makes a file useful in a notes app, a spreadsheet, a video editor or a vector database. This page is the short version: a table to choose from, then one real example of each format with a link to its reference page for the exact fields.
You pick a video once. From that single extraction the transcript exports as seven formats, or nine downloads if you count the timestamped variants of plain text and Markdown as their own option. Everything is free except RAG JSON, and only plain text can be downloaded without an account.
Pick a format
| Format | What it's for | Cost | Account |
|---|---|---|---|
| Plain text (TXT) | Read a video like a document, or start your own writing from it | Free | Not needed |
| Markdown | Notes apps: Obsidian, Notion, Logseq. Frontmatter in the header | Free | Free account |
| SRT | Subtitles for a video editor: Premiere, DaVinci Resolve, CapCut | Free | Free account |
| VTT | Subtitles for the web and course platforms: Canvas, Moodle | Free | Free account |
| CSV | One row per segment, for analysis in a spreadsheet or script | Free | Free account |
| JSON | Segments with timestamps and a metadata wrapper, for developers | Free | Free account |
| RAG JSON | Chunked, with deep links and token estimates, for a vector database | 1 credit / 10 min | Free account |
Plain text and Markdown each come in two shapes, plain and with timestamps, which is why seven formats add up to nine download options in the export menu. The rest of this page takes each format in turn.
Plain text (TXT)
Choose plain text to read a video through, or to hand the words to an AI tool without any markup in the way. INDXR groups the raw caption fragments into paragraphs on the natural pauses in speech, so it reads like a document rather than a wall of two-second lines. There is also a variant with a timestamp on every line, for when you need to point at the exact moment something was said. It is the one format you can download without an account.
Funding for this program is provided by: Additional funding provided by
This is a course about Justice and we begin
with a story suppose you're the driver of a trolley car, and your trolley car is hurdling down
the track at sixty miles an hourThe full field-by-field behaviour is on the plain text reference page.
Markdown
Choose Markdown for a notes app. The file opens with a YAML frontmatter block that Obsidian reads as Properties and drops straight into a vault; Notion imports the headings as a page outline. Fields the video does not carry, like a channel or a publish date, are left out rather than filled with blanks. The timestamps variant adds a clickable ## [HH:MM:SS] heading per section that links back to that second of the video.
---
title: "Justice: What's The Right Thing To Do? Episode 01 …"
url: "https://www.youtube.com/watch?v=kBdfcR-8hEY"
duration: 3296
language: "en"
transcript_source: "YouTube captions"
created: "2026-08-07"
type: youtube
tags: [youtube, transcript]
---
# Justice: What's The Right Thing To Do? Episode 01 …
## [00:00:04](https://youtu.be/kBdfcR-8hEY?t=4)
Funding for this program is provided by: Additional funding provided byThe frontmatter keys and the heading format are on the Markdown reference page. For the full note-taking workflow, from summary to vault, see YouTube to notes.
SRT and VTT subtitles
Choose SRT for a desktop video editor and VTT for a web player or a course platform. Both carry the same cues; SRT writes the time with a comma and VTT opens with a WEBVTT header that also holds the title and language. INDXR rebuilds the transcript into readable subtitle blocks rather than copying the raw fragments across, so no line runs past 42 characters, no block carries more than 2 lines, and none stays on screen longer than 7 seconds, following the Netflix timed-text guideline that most of the industry works to.
1
00:00:04,200 --> 00:00:10,723
Funding for this program is provided by:
Additional
2
00:00:10,723 --> 00:00:15,240
funding provided by
The exact timing format for each is on the SRT reference page and the VTT reference page. For how the blocks are built, and how to make subtitles from an audio file with no video, see the SRT generator.
CSV
Choose CSV to work with the transcript as data: one row per segment, ready for a spreadsheet or a script. Each row carries the segment index, its start and end time, its duration, a word count and the text. The end_time of a row is the start of the next segment, so the rows join up with no gaps. A short block of comment lines at the top records the video title, URL, duration, language and source, and the file is written with a UTF-8 byte-order mark so Excel opens non-Latin scripts without garbling them.
# title: Justice: What's The Right Thing To Do? Episode 01 …
# url: https://www.youtube.com/watch?v=kBdfcR-8hEY
# duration_seconds: 3296
# language: en
# transcript_source: YouTube captions
# extracted: 2026-08-07
segment_index,start_time,end_time,duration,word_count,text
0,4.2,8.24,4.04,7,"Funding for this program is provided by:"
1,8.24,33.51,7,4,"Additional funding provided by"The full column list is on the CSV reference page.
JSON
Choose standard JSON when code will read the transcript. It is a metadata wrapper around the segments: each segment has its text, a start_time and an end_time, and the wrapper carries the video id, title, duration and, when the source is a YouTube video, the channel, language and publish date. It is free for any captioned video.
{
"metadata": {
"video_id": "kBdfcR-8hEY",
"title": "Justice: What's The Right Thing To Do? Episode 01 …",
"duration_seconds": 3296,
"extracted_at": "2026-08-07T15:18:48.032Z",
"language": "en",
"extraction_method": "youtube_captions"
},
"segments": [
{ "text": "Funding for this program is provided by:", "start_time": 4.2, "end_time": 8.24 },
{ "text": "Additional funding provided by", "start_time": 8.24, "end_time": 33.51 }
]
}The full schema, including what a diarised transcript adds, is on the JSON reference page.
RAG JSON
Choose RAG JSON when the transcript is going into a vector database. It is the one export with choices in it, and the one that costs credits. Where standard JSON hands you the raw two to five second segments, RAG JSON merges them into larger chunks sized for an embedding model, which works best on a few hundred tokens of coherent text rather than a stream of short fragments (Vectara NAACL 2025, NVIDIA).
Each chunk carries what the other formats do not: a chunk_id, a deep_link that opens the video at the second the chunk starts, a token_count_estimate, and a flat block of metadata that loads into a vector database without reshaping. The file also records the chunk size, the overlap and the total chunk count in its chunking_config.
{
"metadata": {
"video_id": "kBdfcR-8hEY",
"duration_seconds": 3296,
"language": "en",
"extraction_method": "youtube_captions",
"chunking_config": {
"chunk_size_seconds": 60,
"overlap_seconds": 9,
"overlap_strategy": "segment_boundary",
"total_chunks": 60
}
},
"chunks": [
{
"chunk_index": 0,
"chunk_id": "kBdfcR-8hEY_chunk_000",
"text": "Funding for this program is provided by: Additional funding provided by This is a course about Justice and we begin with a story suppose you're the driver of a trolley car …",
"start_time": 4.2,
"end_time": 65.08,
"deep_link": "https://youtu.be/kBdfcR-8hEY?t=4",
"token_count_estimate": 128,
"metadata": { "…": "flat, per-chunk: video id, title, channel, timestamps, language, total_chunks" }
}
]
}You choose the chunk length when you export:
| Preset | Length | Approx. tokens |
|---|---|---|
| Quote | 30s | ~100 tokens |
| Balanced (default) | 60s | ~200 tokens |
| Precise | 90s | ~300 tokens |
| Context | 120s | ~390 tokens |
RAG JSON costs 1 credit per 10 minutes of video, rounded up, minimum 1. Re-downloading a transcript you have already exported to RAG JSON is free, and credits never expire. There are no free RAG exports: the credit is charged the first time you export a given transcript. The schema and the field-by-field detail are on the JSON reference page. For choosing a chunk size, see how to chunk YouTube transcripts for RAG; for loading the chunks into a store, see YouTube transcripts in a vector database; and for building a searchable base from a whole channel, see a YouTube channel knowledge base.
Exporting a whole library at once
Every format is available in bulk. In your library, select the transcripts you want, pick a format, and download them together. You get a ZIP with one file per video, each named after its video, in the format you chose. It is separate files bundled together, not a single merged document, so a folder of Markdown notes stays a folder of notes and a set of subtitle files stays one file per video.
Playlists feed this directly: the Playlist tab processes every selected video in one job, and the results land in your library ready to select and export as a batch.
Getting started
Everything you extract is saved to your library and stays there to re-export in any format whenever you need it. A free account includes 50 credits and takes a minute to make, with no card. For credit packages, see the pricing page; for the whole extraction pipeline end to end, see how INDXR.AI works.




