SRT and WebVTT carry almost the same information: numbered or unnumbered cues, a start and end time, and one or two lines of text. Use WebVTT for HTML5 video on the web, SRT when a platform or editing tool asks for it, and generate both from the same word timestamps so they never drift apart.
Why there are two formats
SRT (SubRip Text) comes from a DVD-ripping tool of the early 2000s. It was never formally standardised, but it is so simple that nearly every player, editor and upload form reads it. WebVTT (Web Video Text Tracks) is a W3C format built for the browser. It started as a cleaned-up SRT and added what the web needed: a file header, a MIME type, positioning, styling and speaker tags.
The practical difference is small. A cue in both formats is a time range plus text. The details that trip people up are the separators and the header.
The syntax side by side
An SRT file is a sequence of blocks separated by blank lines. Each block has a counter, a time range with a comma before the milliseconds, and the text:
1
00:00:04,400 --> 00:00:07,300
Welcome to the quarterly review
of Muster GmbH.
2
00:00:07,400 --> 00:00:10,900
First item: the RE-2026-0042 dispute.
The same cues in WebVTT. The file starts with the line WEBVTT, the milliseconds use a period, and the counter is optional:
WEBVTT
00:00:04.400 --> 00:00:07.300
Welcome to the quarterly review
of Muster GmbH.
00:00:07.400 --> 00:00:10.900
First item: the RE-2026-0042 dispute.
WebVTT can do more if you want it: cue settings after the time range (line:, position:, align:), NOTE comments, STYLE blocks and voice spans such as <v Anna>. Most generated subtitles use none of these, which is why converting between the two is mostly a matter of swapping the comma and adding or removing the header.
Where each format is used
- HTML5 video: the
<track>element expects WebVTT, served astext/vtt. Browsers do not play SRT natively; you would have to convert it in JavaScript. - YouTube and most video platforms: accept both SRT and WebVTT as uploads. If you only produce one file for a platform upload, SRT is the safer choice because older tools and editors read it too.
- Video editors: most desktop editors import SRT; WebVTT support varies by tool and version. Check yours before you commit to a pipeline.
- Broadcast: TV delivery usually uses other formats (for example EBU STL, EBU-TT or TTML in Europe, CEA-608/708 captions in North America). SRT and WebVTT are rarely the final broadcast deliverable, but their line-length and timing conventions follow the same readability logic.
The 42 characters, two lines convention
The most widely used convention, from Netflix and many other streaming style guides, is at most 42 characters per line and at most two lines per cue. Broadcast guidelines are similar or stricter; BBC subtitles, for example, allow 37 characters per line. It exists because a longer line forces the eye to travel too far across the screen, and a third line covers too much of the picture. Most streaming style guides use these numbers for Latin-script languages.
A few rules follow from it:
- Break lines at natural points: after punctuation, before a conjunction or preposition, never between an article and its noun if you can avoid it.
- Prefer a short top line and a longer bottom line, or two lines of similar length. The exact preference differs between style guides.
- A speaker prefix (
Anna:) counts towards the 42 characters.
Timing rules
Line length is only half of readability. The other half is time on screen. The common guidance, with some variation between style guides:
- Maximum duration of a cue around 6 to 7 seconds. Longer cues make viewers re-read or lose track.
- Minimum duration around 0.8 to 1 second, even for a single word, so the text is actually seen.
- Reading speed roughly 15 to 20 characters per second for adult audiences. The exact number depends on language, audience and the guideline you follow.
- Gaps: leave a small gap between consecutive cues so the screen visibly changes. Do not overlap cues from the same speaker.
- Sync to speech: a cue should appear when the words start, not a second early.
Generating both from word timestamps
Segment timestamps from a speech recognition model usually cover whole sentences or pauses, and a sentence often runs longer than two lines. Word timestamps let you cut cues exactly where the line budget runs out. The algorithm is a greedy walk over the words:
- Start a cue with the first word.
- Add the next word if the cue still fits in two lines of 42 characters (including any speaker prefix), stays under 7 seconds, and the speaker has not changed.
- Close the cue early when the previous word ended a sentence and the cue already has a reasonable amount of text.
- Pad cues that are shorter than the minimum duration, but never into the next cue.
- Render the same cue list twice: once with commas and counters (SRT), once with the header and periods (WebVTT).
The last step is the important one. If both files come from one cue list, they cannot disagree.
klangzeile does this on delivery. When you upload with timestamps=word, the final GET can return a subtitle file instead of JSON:
curl -X POST https://klangzeile.com/v1/transcriptions \
-H "Authorization: Bearer $KEY" \
-F "file=@review.mp3" -F "language=en" \
-F "timestamps=word" -F "speakers=true"
curl "https://klangzeile.com/v1/transcriptions/tr_…?format=vtt&speaker_names=S1:Anna,S2:Ben" \
-H "Authorization: Bearer $KEY" -o review.vtt
The generated files follow 42 characters per line, two lines and at most 7 seconds per cue, and prefix speakers as Anna: in plain text rather than with WebVTT voice tags. One limitation to know: the result is delivered once, so you get either JSON or one subtitle file from that fetch. If you need SRT and WebVTT, fetch format=json, take the words array and render both yourself with the steps above.
A worked example
Muster GmbH records a 20-minute product demo and wants subtitles on its website and on YouTube. They upload once with word timestamps, fetch the JSON, and run their own renderer over words:
demo.vttgoes into the<track kind="subtitles" srclang="en">element on the website.demo.srtgoes to YouTube and to the video editor for a cut-down version.
Because both files come from one cue list, a correction in the transcript (say, a product name) is fixed once and appears in both.
FAQ
Can I just rename an .srt file to .vtt?
No. You need at least the WEBVTT header and periods instead of commas in the timestamps. Some players tolerate the commas, but the specification does not.
Should subtitles contain speaker names? For interviews and meetings, yes, when the speaker is not visible or changes often. For a single narrator, leave them out; they only cost characters.
What about translated subtitles? They usually keep the timing of the original segments, because word-level timing does not map across languages. Expect longer lines in German than in English and check the reading speed.
Related: Speaker diarization explained · Meeting transcription with a summary API · Custom vocabulary for speech recognition · GDPR-compliant transcription in the EU