Skip to content
INDXR.AI
All docs

SRT

SRT — short for SubRip — is the subtitle format almost every video player and editor reads. An SRT export turns your transcript into timed, numbered subtitle cues you can drop straight into YouTube, Premiere, CapCut, DaVinci Resolve, or VLC. INDXR re-segments the transcript into readable cues and wraps long lines, so you get clean subtitles instead of the raw caption fragments most tools hand back.

Reach for SRT when you want subtitles on a video — burned in during editing, or added as a selectable track. If you're publishing to the web instead, its sibling VTT is the HTML5 equivalent. The rest of this page is the exact shape of the file.

Timestamp format

SRT uses a comma before the milliseconds: HH:MM:SS,mmm (this is the difference from VTT, which uses a dot). Cues are numbered from 1.

Re-segmentation & line wrapping

Segments are re-cut into cues rather than shown one raw fragment at a time. INDXR builds a per-word timeline, then packs words into a cue until the text would need more than 2 lines of 42 characters or the cue would run past 7 seconds. It prefers to end a cue on a sentence boundary: if a cue stopped mid-sentence but a sentence ended earlier inside it, the cue is cut back to that boundary, so a sentence is never split across cues unless it is itself too long for one. A change of speaker always starts a new cue.

Each cue is then held on screen long enough to read: at least 1 second, lengthened toward 20 characters per second by filling the silent gap before the next cue (which never shifts the timeline), and never left above 21 characters per second. A single word longer than 42 characters is left intact rather than broken.

When a transcript has speaker labels, the name shows on the first cue of each turn. SRT has no speaker field, so the name is baked in as a Name: prefix that counts against the 42-character line budget — unlike VTT, which carries it out of budget as a <v Name> voice tag. Because SRT spends characters on the name and VTT does not, a turn-opening cue fits less spoken text in SRT than in VTT, so the two files break into cues slightly differently.

3
00:00:33,509 --> 00:00:38,799
This is a course about Justice and we
begin with a story suppose you're the

4
00:00:38,799 --> 00:00:43,760
driver of a trolley car, and your trolley
car is hurdling down the track at sixty

Sources

  • Matroska / SubRip (.srt) — the SRT cue + comma-millisecond timestamp convention
  • Netflix Timed Text Style Guide — the 42-char / 2-line / 7s cue conventions and reading-speed target
  • INDXR (own code) — segmentation constants, sentence-aware cue packing, 42-char line wrapVerified against packages/shared/src/lib/subtitleConfig.ts + packages/shared/src/utils/formatTranscript.ts (generateSrt, buildSubtitleCues, wrapLines)