Skip to content
INDXR.AI
five brass measuring cups lined up casting long shadows
Formats

Transcript export formats: which file to choose

IE
INDXR.AI Editorial
Published April 16, 2026 · Updated August 27, 2026

You have a transcript, and now you have to decide which file to take it away in. The words are the same in every one; what changes is the wrapping, and the wrapping is what makes a file useful in a notes app, a spreadsheet, a video editor or a vector database. This page is the short version: a table to choose from, then one real example of each format with a link to its reference page for the exact fields.

You pick a video once. From that single extraction the transcript exports as seven formats, or nine downloads if you count the timestamped variants of plain text and Markdown as their own option. Everything is free except RAG JSON, and only plain text can be downloaded without an account.

Pick a format

FormatWhat it's forCostAccount
Plain text (TXT)Read a video like a document, or start your own writing from itFreeNot needed
MarkdownNotes apps: Obsidian, Notion, Logseq. Frontmatter in the headerFreeFree account
SRTSubtitles for a video editor: Premiere, DaVinci Resolve, CapCutFreeFree account
VTTSubtitles for the web and course platforms: Canvas, MoodleFreeFree account
CSVOne row per segment, for analysis in a spreadsheet or scriptFreeFree account
JSONSegments with timestamps and a metadata wrapper, for developersFreeFree account
RAG JSONChunked, with deep links and token estimates, for a vector database1 credit / 10 minFree account

Plain text and Markdown each come in two shapes, plain and with timestamps, which is why seven formats add up to nine download options in the export menu. The rest of this page takes each format in turn.

Plain text (TXT)

Choose plain text to read a video through, or to hand the words to an AI tool without any markup in the way. INDXR groups the raw caption fragments into paragraphs on the natural pauses in speech, so it reads like a document rather than a wall of two-second lines. There is also a variant with a timestamp on every line, for when you need to point at the exact moment something was said. It is the one format you can download without an account.

Funding for this program is provided by: Additional funding provided by

This is a course about Justice and we begin
with a story suppose you're the driver of a trolley car, and your trolley car is hurdling down
the track at sixty miles an hour

The full field-by-field behaviour is on the plain text reference page.

Markdown

Choose Markdown for a notes app. The file opens with a YAML frontmatter block that Obsidian reads as Properties and drops straight into a vault; Notion imports the headings as a page outline. Fields the video does not carry, like a channel or a publish date, are left out rather than filled with blanks. The timestamps variant adds a clickable ## [HH:MM:SS] heading per section that links back to that second of the video.

---
title: "Justice: What's The Right Thing To Do? Episode 01 …"
url: "https://www.youtube.com/watch?v=kBdfcR-8hEY"
duration: 3296
language: "en"
transcript_source: "YouTube captions"
created: "2026-08-07"
type: youtube
tags: [youtube, transcript]
---

# Justice: What's The Right Thing To Do? Episode 01 …

## [00:00:04](https://youtu.be/kBdfcR-8hEY?t=4)
Funding for this program is provided by: Additional funding provided by

The frontmatter keys and the heading format are on the Markdown reference page. For the full note-taking workflow, from summary to vault, see YouTube to notes.

SRT and VTT subtitles

Choose SRT for a desktop video editor and VTT for a web player or a course platform. Both carry the same cues; SRT writes the time with a comma and VTT opens with a WEBVTT header that also holds the title and language. INDXR rebuilds the transcript into readable subtitle blocks rather than copying the raw fragments across, so no line runs past 42 characters, no block carries more than 2 lines, and none stays on screen longer than 7 seconds, following the Netflix timed-text guideline that most of the industry works to.

1
00:00:04,200 --> 00:00:10,723
Funding for this program is provided by:
Additional

2
00:00:10,723 --> 00:00:15,240
funding provided by
An exported SRT subtitle file: numbered cues, each with a start and end timestamp and one or two short lines of text.
An exported SRT file: numbered cues, timestamps, and lines rebuilt to subtitle length rather than raw transcript fragments.

The exact timing format for each is on the SRT reference page and the VTT reference page. For how the blocks are built, and how to make subtitles from an audio file with no video, see the SRT generator.

CSV

Choose CSV to work with the transcript as data: one row per segment, ready for a spreadsheet or a script. Each row carries the segment index, its start and end time, its duration, a word count and the text. The end_time of a row is the start of the next segment, so the rows join up with no gaps. A short block of comment lines at the top records the video title, URL, duration, language and source, and the file is written with a UTF-8 byte-order mark so Excel opens non-Latin scripts without garbling them.

# title: Justice: What's The Right Thing To Do? Episode 01 …
# url: https://www.youtube.com/watch?v=kBdfcR-8hEY
# duration_seconds: 3296
# language: en
# transcript_source: YouTube captions
# extracted: 2026-08-07
segment_index,start_time,end_time,duration,word_count,text
0,4.2,8.24,4.04,7,"Funding for this program is provided by:"
1,8.24,33.51,7,4,"Additional funding provided by"

The full column list is on the CSV reference page.

JSON

Choose standard JSON when code will read the transcript. It is a metadata wrapper around the segments: each segment has its text, a start_time and an end_time, and the wrapper carries the video id, title, duration and, when the source is a YouTube video, the channel, language and publish date. It is free for any captioned video.

{
  "metadata": {
    "video_id": "kBdfcR-8hEY",
    "title": "Justice: What's The Right Thing To Do? Episode 01 …",
    "duration_seconds": 3296,
    "extracted_at": "2026-08-07T15:18:48.032Z",
    "language": "en",
    "extraction_method": "youtube_captions"
  },
  "segments": [
    { "text": "Funding for this program is provided by:", "start_time": 4.2, "end_time": 8.24 },
    { "text": "Additional funding provided by", "start_time": 8.24, "end_time": 33.51 }
  ]
}

The full schema, including what a diarised transcript adds, is on the JSON reference page.

RAG JSON

Choose RAG JSON when the transcript is going into a vector database. It is the one export with choices in it, and the one that costs credits. Where standard JSON hands you the raw two to five second segments, RAG JSON merges them into larger chunks sized for an embedding model, which works best on a few hundred tokens of coherent text rather than a stream of short fragments (Vectara NAACL 2025, NVIDIA).

Each chunk carries what the other formats do not: a chunk_id, a deep_link that opens the video at the second the chunk starts, a token_count_estimate, and a flat block of metadata that loads into a vector database without reshaping. The file also records the chunk size, the overlap and the total chunk count in its chunking_config.

{
  "metadata": {
    "video_id": "kBdfcR-8hEY",
    "duration_seconds": 3296,
    "language": "en",
    "extraction_method": "youtube_captions",
    "chunking_config": {
      "chunk_size_seconds": 60,
      "overlap_seconds": 9,
      "overlap_strategy": "segment_boundary",
      "total_chunks": 60
    }
  },
  "chunks": [
    {
      "chunk_index": 0,
      "chunk_id": "kBdfcR-8hEY_chunk_000",
      "text": "Funding for this program is provided by: Additional funding provided by This is a course about Justice and we begin with a story suppose you're the driver of a trolley car …",
      "start_time": 4.2,
      "end_time": 65.08,
      "deep_link": "https://youtu.be/kBdfcR-8hEY?t=4",
      "token_count_estimate": 128,
      "metadata": { "…": "flat, per-chunk: video id, title, channel, timestamps, language, total_chunks" }
    }
  ]
}

You choose the chunk length when you export:

PresetLengthApprox. tokens
Quote30s~100 tokens
Balanced (default)60s~200 tokens
Precise90s~300 tokens
Context120s~390 tokens

RAG JSON costs 1 credit per 10 minutes of video, rounded up, minimum 1. Re-downloading a transcript you have already exported to RAG JSON is free, and credits never expire. There are no free RAG exports: the credit is charged the first time you export a given transcript. The schema and the field-by-field detail are on the JSON reference page. For choosing a chunk size, see how to chunk YouTube transcripts for RAG; for loading the chunks into a store, see YouTube transcripts in a vector database; and for building a searchable base from a whole channel, see a YouTube channel knowledge base.

Exporting a whole library at once

Every format is available in bulk. In your library, select the transcripts you want, pick a format, and download them together. You get a ZIP with one file per video, each named after its video, in the format you chose. It is separate files bundled together, not a single merged document, so a folder of Markdown notes stays a folder of notes and a set of subtitle files stays one file per video.

Playlists feed this directly: the Playlist tab processes every selected video in one job, and the results land in your library ready to select and export as a batch.

Getting started

Everything you extract is saved to your library and stays there to re-export in any format whenever you need it. A free account includes 50 credits and takes a minute to make, with no card. For credit packages, see the pricing page; for the whole extraction pipeline end to end, see how INDXR.AI works.

Frequently Asked Questions

Do I need an account to export?
Only for everything except plain text. A TXT download works with no account. Every other format needs a free account, which comes with 50 credits. It is a sign-in wall on the richer formats, not a paywall.
Which formats are free?
All of them except RAG JSON. Plain text, Markdown, CSV, SRT, VTT and standard JSON cost no credits once the transcript exists. RAG JSON is the one paid export.
What does RAG JSON cost?
1 credit per 10 minutes of video, rounded up, minimum 1. Re-downloading a transcript you have already exported to RAG JSON is free, and credits never expire. There are no free RAG exports; the credit is charged the first time.
Can I export a whole playlist at once?
Yes. Select the transcripts in your library and download them together. You get a ZIP with one file per video in the format you chose, not a single merged file.
What is the difference between standard JSON and RAG JSON?
Standard JSON is the raw segments with a metadata wrapper, free. RAG JSON merges the segments into larger chunks and adds per-chunk deep links, token estimates and flat metadata for a vector database. Standard JSON is a data format; RAG JSON is a pipeline-ready input.
Which format should I use for Obsidian or Notion?
Markdown. It carries a YAML frontmatter block that Obsidian reads as Properties, and Notion imports the headings as a page outline. Use the timestamps variant if you want each section to link back to the video.

Sources

See also