An SRT made from the speech
The SRT is generated from what people say in the MP4: the audio is transcribed, timed and written out as numbered cues. It doesn't extract a subtitle track that is already stored inside the file, and it can't read text burned into the picture. If your MP4 already contains a soft subtitle track, a tool like ffmpeg can copy it out directly (ffmpeg -i video.mp4 -map 0:s:0 subtitles.srt), which is faster and free. Embedded vs sidecar subtitles explains the difference between the two kinds.
What you get back
Every run returns the SRT, the same cues as a VTT for web players, a timestamped transcript and a plain transcript. An SRT cue looks like this:
12 — 00:01:04,200 --> 00:01:07,900 — And that's why we moved the launch to March.
Each cue has a number, a start and end time with milliseconds after a comma, and the text. The file is plain UTF-8 text, so accented letters, Devanagari, Arabic and Chinese characters are stored correctly and you can edit it in any text editor. How to edit an SRT file shows how to fix a word or a timing.
Getting the SRT into your video
- YouTube
- Studio → Subtitles → Add language → Upload file → With timing
- Premiere Pro
- Import the SRT and drag it onto the timeline as a caption track
- DaVinci Resolve
- File → Import → Subtitle, then drag it onto the timeline
- VLC and desktop players
- Give the SRT the same name as the MP4 and keep both in one folder, or drag the SRT onto the playing video
To burn subtitles into the picture for platforms that need it, import the SRT into an editor and export with captions burned in; mydubly delivers the file only. The guide to adding subtitles to a video has steps for more platforms.
Line length and timing
Cues follow natural speech segments, so their length follows the speaker. Fast talkers produce long cues; common guidelines suggest about 42 characters per line and two lines per cue. Where a cue is too long, split it in the SRT, keeping the timings in order. Cue times come from the audio, so variable frame rate in phone recordings doesn't push the subtitles out of sync. Subtitle segmentation covers where to break lines.
A translated SRT from the same MP4
Pick a different language from the one spoken and the SRT comes back translated, with transcripts in both languages. Each run makes one subtitle language, so run the MP4 again for each extra language. Subtitles in Czech and Vietnamese work too, even though those two have no dubbed voice. For several languages and how to choose them, see translated subtitles.
An SRT in the spoken language costs 30 credits (3¢). A Spanish SRT from the same file is another 30 credits (3¢).
Files are billed by length, at 1 credit per minute with a 5-credit minimum, up to 2 hours per file.
Using the tool on this page
- The spoken language is detected automatically; you select the output language — the spoken language itself for an untranslated transcript, or another language for a translation. Files can be up to 2 hours long.
- Only the audio is sent for processing — the video stays on your device — and audio and text are deleted within 30 minutes of delivery (how files are protected).
- You can download a timestamped transcript, subtitles (SRT and VTT) and plain text. Every output, step and limit is explained on the subtitle generator page.
See also
Frequently asked questions
Can it pull out subtitles that are already inside my MP4?
No. The SRT is generated from the speech. If the MP4 already has a soft subtitle track, extract it with a tool like ffmpeg or MKVToolNix instead.
Do I get a VTT file as well?
Yes. Every run returns the subtitles as both SRT and VTT, plus a timestamped transcript and plain text.