What is an VTT file?
WebVTT is the caption format built for the web. It starts with a WEBVTT header and uses dots in timestamps (00:01:02.500 --> 00:01:05.000); browsers display it natively through the <track> element.
What the file looks like
WEBVTT 00:00:00.000 --> 00:00:03.200 Today we continue from the last session. 00:00:03.200 --> 00:00:06.900 Price is the signal, not the cause.
Where it's used
- HTML5 video via <track kind="subtitles">
- Video.js, Plyr, JW Player and similar web players
- learning platforms and LMS video players
- Vimeo and Wistia uploads
Tips
- Serve .vtt files with the text/vtt content type, or some browsers will ignore the track.
- If the captions live on another domain, the video element needs crossorigin and the server must send CORS headers.
- Add one <track> per language and set default on the one you want shown first.
What a WebVTT file can contain
WebVTT looks like SRT at first glance but it's a richer format, and knowing what it allows helps when you edit. The file must begin with the word WEBVTT on the first line, optionally followed by a space and a title. After a blank line come the cues. Each cue may open with an identifier (a number or a name such as intro), then the timing line, then the text. Hours are optional, so 01:15.000 is valid, and a period sits before the milliseconds.
- NOTE block
- a comment for editors that players ignore
- Cue settings
- words after the timing line, such as line:0 to move a cue to the top or align:start
- Voice tag
- <v Maria> before the text marks who is speaking
- STYLE block
- CSS for ::cue, placed before the first cue
- Formatting tags
- <i>, <b>, <u> and <c.name> for styled spans
mydubly's VTT contains timing and text only, so all of these are additions you make by hand when a page needs them. Converting a plain SRT to WebVTT is mechanical (add the header, change the comma before the milliseconds to a period), but converting back drops cue settings, voice tags and styling, so do your WebVTT-only edits last and keep them in the .vtt.
Captions or subtitles: choosing the track kind
The HTML <track> element takes a kind attribute, and it tells the player's menu and assistive technology what the track is. kind="subtitles" is a translation for viewers who can hear but don't understand the language; kind="captions" is same-language text for viewers who can't hear, which ideally also marks speaker changes and important sounds. A VTT made with the spoken language selected is a same-language transcript of the speech: call it captions once you've added sound cues and voice tags, and subtitles until then. A translated VTT is subtitles. Captions vs subtitles explains the distinction and its regional variations.
Styling and positioning limits
- The ::cue selector accepts a limited set of properties (colour, background, font, text decoration and a few others). Layout properties such as margin and padding are ignored.
- On an iPhone in fullscreen, captions may be drawn natively in the viewer's system caption style, ignoring your ::cue rules. Test there before signing off.
- Viewers can override caption appearance in their operating system's accessibility settings. Treat that as intended and never rely on colour alone to carry meaning.
- When a lower-third graphic hides one cue, add line:0 to that cue rather than moving every cue. Subtitle positioning covers safe areas.
Three languages on one course video
Run the lesson once with English selected for the spoken English, untranslated, once with Spanish and once with French. Give each track a srclang (en, es, fr) and a label written in that language, such as Español and Français, so viewers recognise their own language in the menu. An eight-lesson module of 12-minute videos costs 96 credits (9.6¢) per language, or 288 credits (28.8¢) for all three.
For course platforms, online course translation covers the rest of the workflow, multilingual video on your website covers page structure, and accessible video players helps you pick a player that exposes caption controls properly. If the destination is an editor or an upload form instead, the SRT generator page is the better fit; each job offers both files.
Interactive transcripts from the same file
Because the browser parses a VTT into cue objects, a short script can turn the caption file into a clickable transcript beside the player. Listen for the track's cuechange event to highlight the current line, and set the video's currentTime when someone clicks a line. Set the track's mode to hidden if you want the cues available to your script without drawing them over the picture. It's a common pattern on course and documentation sites and needs nothing beyond the VTT you already have.
Using the tool on this page
- The spoken language is detected automatically; you select the output language — the spoken language itself for an untranslated transcript, or another language for a translation. Files can be up to 2 hours long.
- Only the audio is sent for processing — the video stays on your device — and audio and text are deleted within 30 minutes of delivery (how files are protected).
- You can download a timestamped transcript, subtitles (SRT and VTT) and plain text. Every output, step and limit is explained on the subtitle generator page.
Frequently asked questions
How do I add the VTT to my website's video?
Add <track src="captions.vtt" kind="subtitles" srclang="en" label="English" default> inside your <video> element.
Should I use VTT or SRT?
Use VTT for web players and HTML5 video; use SRT for editors, desktop players and most upload forms. Every job gives you both.
Does the VTT include styling or positioning?
No — cues contain timing and text only. Add VTT cue settings by hand if you need positioning.
The VTT loads but no captions appear. What should I check?
Check that the very first line is exactly WEBVTT (a byte-order mark before it is allowed, any other text isn't), that timing lines use periods rather than commas, and that the track isn't disabled. In the browser console, inspecting video.textTracks[0].cues shows whether the cues were parsed at all.
Can I use the VTT with HLS streaming?
HLS carries WebVTT as segmented subtitle playlists. Packaging tools such as ffmpeg can segment a single VTT for you, and most video hosts that accept caption uploads handle the packaging themselves.
Can WebVTT hold chapters as well as captions?
Yes. A track with kind="chapters" uses the same file format, with each cue's text being a chapter title and its timing covering that section. The downloaded VTT is a caption track, but the timestamped transcript from the same job is a quick way to find where topics change and write a chapters file by hand. Player support for chapter tracks varies, so check yours.