VTT
VTT is the web-native subtitle format — it's what an HTML5 video player loads for on-screen captions through a <track> element. Reach for it when your video plays in a browser; for a desktop editor, SRT is the sibling to use. The file starts with a required WEBVTT header, an optional NOTE block carrying the title and language, then numbered cues with HH:MM:SS.mmm time ranges — a dot before the milliseconds, unlike SRT's comma.
Timestamp format & header
VTT uses a dot before the milliseconds: HH:MM:SS.mmm (SRT uses a comma). The file always opens with WEBVTT; the NOTE block is added when a title or language is known.
Re-segmentation & line wrapping
The cue text is segmented the same way as SRT: words are packed into a cue until they would need more than 2 lines of 42 characters or run past 7 seconds; cues prefer to end on a sentence boundary; a change of speaker starts a new cue; and each cue stays on screen between 1 second and 7 seconds — lengthened toward 20 characters per second by filling the silent gap before the next cue, and never left above 21 characters per second.
Where a transcript has speakers, VTT carries the name as a native <v Name> voice tag on the first cue of each turn. That tag is zero-width on screen, so the full 42 characters stay available for spoken text — unlike SRT, which spends part of the line on a Name: prefix. Because VTT keeps the name out of the line budget and SRT does not, a turn-opening cue fits more spoken text in VTT, so the two files break into cues slightly differently.
WEBVTT
NOTE
title: Justice: What's The Right Thing To Do? Episode 01 "THE MORAL SIDE OF MURDER"
language: en
3
00:00:33.509 --> 00:00:38.799
This is a course about Justice and we
begin with a story suppose you're the
4
00:00:38.799 --> 00:00:43.760
driver of a trolley car, and your trolley
car is hurdling down the track at sixtySources
- W3C — the WebVTT format (WEBVTT header, dot-millisecond timestamps, <v> voice tag)
- Netflix Timed Text Style Guide — the 42-char / 2-line / 7s cue conventions and reading-speed target
- INDXR (own code) — NOTE block, segmentation constants, <v> voice tag, line wrappingVerified against packages/shared/src/lib/subtitleConfig.ts + packages/shared/src/utils/formatTranscript.ts (generateVtt, buildSubtitleCues)



