What is an SRT file?
SRT (SubRip) is the most widely supported subtitle format: a plain-text file of numbered cues, each with a start and end time (00:01:02,500 --> 00:01:05,000) and the line of text.
What the file looks like
1 00:00:00,000 --> 00:00:03,200 Today we continue from the last session. 2 00:00:03,200 --> 00:00:06,900 Price is the signal, not the cause.
Where it's used
- YouTube Studio and Vimeo uploads
- Premiere Pro, DaVinci Resolve, Final Cut (via import)
- VLC, MPV and most desktop players
- LinkedIn and Facebook video uploads
Tips
- Cues follow the speech segments the model detects. Very long sentences become long cues — split them in any subtitle editor if a platform enforces line limits.
- SRT has no styling. If you need positioning or colours, convert to another format after editing.
- Select another language to get a translated SRT; select the spoken language for untranslated subtitles.
Anatomy of a cue
An SRT file is plain text made of cues separated by blank lines. Each cue is laid out like this:
- First line
- the cue number: 1, 2, 3 and so on, in order
- Second line
- start and end time as hours:minutes:seconds,milliseconds with an arrow between them, e.g. 00:00:04,100 --> 00:00:07,350
- Next line or lines
- the subtitle text, usually one or two lines
- Blank line
- closes the cue; without it many players merge two cues into one
Three details cause most broken files: the separator before milliseconds is a comma (the period is WebVTT's convention), hours are always written even when they're 00, and the arrow needs a space on each side. Formatting is limited to a few HTML-like tags such as <i> for italics, which many players honour and some platforms strip. Most players don't care whether cue numbers run without gaps, but some upload validators do, so renumber after deleting cues.
How mydubly's SRT is produced
Cues come from the speech segments found during transcription, so a cue starts when a phrase starts and ends when it ends rather than at fixed intervals. Which text you get depends on the mode:
- Transcript mode with the spoken language selected: untranslated subtitles in the spoken language.
- Transcript mode with a target language: the same segmenting, with each cue's text translated.
- Dub mode: the SRT follows what the AI voice actually says, including lines shortened to fit the timing, so it matches the dubbed audio rather than a literal translation of the original speech.
The SRT is always a separate file and the picture itself is never changed, so viewers can switch subtitles off and you can keep editing them after publishing.
Editing without breaking the file
Cue 41 reads "the new sequel database" at 00:12:03,400 --> 00:12:08,900. Change the text to "the new SQL database" and leave the timing line alone. If the cue is too long to read comfortably, split it at the comma into two cues and let a subtitle editor renumber the rest.
- Use a plain-text editor or a dedicated subtitle editor such as Subtitle Edit or Aegisub, not a word processor, which can insert smart quotes and change line endings.
- Save as UTF-8. Accented letters and Cyrillic, Arabic or Chinese characters turn into gibberish when an SRT is re-saved in a legacy encoding; strange characters in subtitles explains the fix.
- If you trim an intro from the video after making subtitles, shift every cue with the editor's timing-offset tool instead of retyping hundreds of timestamps.
- Check reading speed after translating. Translated lines can run longer than the original, and many style guides cap characters per second, so split or condense cues that flash past.
How to edit an SRT file goes through each step.
Getting it to show up
Upload forms on YouTube, Vimeo and LinkedIn take the .srt directly; the YouTube subtitles workflow covers language tracks there. Desktop players such as VLC and MPV load subtitles automatically when the file sits beside the video with the same base name (lecture.mp4 and lecture.srt), and many media servers read a language code before the extension, as in lecture.es.srt. Editors import SRT as a caption track you can restyle. To carry subtitles inside the video file, MKV accepts SRT as it is, while MP4 needs it converted to its own text-track format while muxing (in ffmpeg, the mov_text subtitle codec does this). When a platform rejects the file, SRT file not working lists the usual causes, and for web players built on WebVTT see the VTT generator.
What a subtitle track costs
Subtitles are priced as transcription. A 25-minute video costs 25 credits (2.5¢) for an SRT in the spoken language, and each further language is another run at the same price. A 3-minute clip is billed at the 5-credit minimum, 5 credits (0.5¢). Dubbing the 25-minute video costs 1,250 credits ($1.25) and includes an SRT matched to the new voice.
Using the tool on this page
- The spoken language is detected automatically; you select the output language — the spoken language itself for an untranslated transcript, or another language for a translation. Files can be up to 2 hours long.
- Only the audio is sent for processing — the video stays on your device — and audio and text are deleted within 30 minutes of delivery (how files are protected).
- You can download a timestamped transcript, subtitles (SRT and VTT) and plain text. Every output, step and limit is explained on the subtitle generator page.
Frequently asked questions
Can I upload the SRT to YouTube?
Yes. In YouTube Studio open Subtitles → Add language → Upload file → With timing, and choose the .srt.
Can I get SRT subtitles in two languages?
One job produces one subtitle track: in the spoken language, or in the target language you pick. Run the file again with another language for a second track.
Can I upload an existing SRT to translate it?
Not currently. mydubly creates subtitles from the audio of a video or audio file; it doesn't import subtitle files.
Can a professional translator work from the SRT?
Yes. A timed source-language SRT is a common starting point: the translator replaces the text and keeps the timing, so every language shares the same cues. Subtitle templates explains the approach.
Will the SRT include sound descriptions like [music] or [laughter]?
Expect speech only. Recognition transcribes words, so tags for music, laughter or off-screen sounds need adding by hand if you're captioning for deaf and hard-of-hearing viewers. Captioning sound effects and music covers the conventions.
Why doesn't the SRT from a dub job match the translated transcript word for word?
In dub mode the subtitles follow what the voice actually says, and some lines are shortened so the speech fits the original timing. That keeps subtitles and audio consistent for viewers who watch with both. If you want subtitles that follow a fuller translation instead, run the file in transcript mode with the same target language.