Subtitle generator

AI Subtitle Generator

Create timed subtitles from your video automatically. Get SRT and VTT files in the spoken language, or translated, ready for YouTube, web players and editors.

01 — Upload

Your video

Full translation outputTimestamped transcript only — no video or audio in results.

The video file stays in this browser. Language is required: pick the spoken language for a plain transcript, or another language to translate it.

02 — Result

Your transcript

Your timestamped transcript will appear here.

Waiting for upload

25 languages available

Subtitle files you can use anywhere

SRT
For YouTube, Vimeo, LinkedIn, editors and desktop players — SRT generator
VTT
For HTML5 video and web players — VTT generator
Timestamped TXT
For reading and quoting — transcript with timestamps

Cues follow natural speech segments. Subtitles are delivered as files, not burned into the picture; editors like Premiere Pro or DaVinci Resolve can burn them in if a platform requires it.

Translated subtitles

Choose a target language to get subtitles in that language instead of the spoken one. Each run produces one subtitle track, so run the video again for each extra language.

Example

Subtitles for a 24-minute episode cost 24 credits (2.4¢) per language.

Adding subtitles to your video

  1. YouTube: Studio → Subtitles → Add language → Upload file → With timing.
  2. Website: add a <track> element pointing at the VTT file.
  3. Editors: import the SRT onto your timeline.

Step-by-step for each platform in how to add subtitles to a video.

How it works

  1. Drop in a video. Your browser reads it and pulls out only the audio — the video file never leaves your device.
  2. The audio is processed in 30-second chunks: speech is transcribed and the spoken language is detected automatically.
  3. Download the results. Audio and text on our servers are deleted after delivery, within 30 minutes at most.

Supported files

  • Video: MP4, WebM, MOV, MKV, M4V — whatever your browser can read.
  • Length: up to 2 hours per file.
  • Languages: speech in 21 languages is recognised automatically; translation and voices cover the same 21.

Privacy

Your video is never uploaded — only its audio is, over HTTPS in small chunks. Audio, transcripts and voice files are deleted after delivery, within 30 minutes at most, and are never used to train models. Your account keeps only the file name, length, languages and credits used. Details in the privacy policy.

Good to know

  • No speaker labels — transcripts don't say who is speaking.
  • Files only: no links from YouTube or social apps, and no live audio.
  • Subtitles come as SRT/VTT files rather than burned into the picture.

Captions in the spoken language, or subtitles in another

There are two different jobs here. Same-language captions serve viewers who watch muted or can't hear the audio; run transcript mode with the spoken language selected and the SRT and VTT files come back untranslated, in the language spoken. Translated subtitles serve viewers who don't understand the speech; pick a different target language and the files come out in that language instead.

For accessibility, captions usually do more than repeat the words: deaf and hard-of-hearing viewers rely on cues such as [applause] or speaker names. These files contain spoken words only, so add those by hand where they matter. Captioning sound effects and music covers the conventions.

Transcript-mode or dub-mode subtitle files

Transcript mode, spoken language selected
Captions in the spoken language, untranslated, 1 credit per minute
Transcript mode with a target language
The full translation of each segment, 1 credit per minute
Dub mode
Subtitles that follow what the translated voice says, including lines shortened to fit timing; included in the 50 credits per minute

If you publish a dubbed version and also offer subtitles, use the dub-mode file so text and voice agree. If you publish subtitles over the original audio, the transcript-mode file is the fuller translation.

A sample cue, and what to check before publishing

One SRT cue

A number (12), a timing line (00:01:04,200 --> 00:01:07,900), the text (Press the blue button to save your draft.) and a blank line before the next cue. VTT uses a dot instead of a comma in the times and starts with a WEBVTT header.

  • Reading time: a long sentence spoken quickly can flash past. Split it into two cues in a subtitle editor.
  • Line length: players wrap long lines differently, so preview on the device your audience uses.
  • Names and terms: correct them once in the file rather than in every platform's caption editor.
  • Right-to-left languages: Arabic and Hebrew subtitles need a player that handles right-to-left text; check punctuation at line ends. The Arabic page covers this.
  • Odd characters: if accents turn into symbols, the player is reading the file with the wrong text encoding.

Editing SRT files covers fixing text and shifting timing, and splitting text into cues covers where line breaks should fall.

Generate subtitles from the final cut

Subtitle timings are tied to the exact file they were made from. If you trim the intro, move a scene or add a title card after generating them, every cue after the change drifts. The fix is ordering: lock the edit, export, then run that export. If you deliver several versions, such as a full episode and a short cut for social, run each export separately rather than trying to re-time one file for both.

For long projects edited in stages, a useful pattern is to generate a quick draft from the rough cut for reviewers to read, then a clean run on the final export for publishing. Before release, a short subtitle quality check catches timing, spelling and reading-speed problems in one pass.

Building a set of subtitle languages

Each run produces one subtitle language. A 30-minute episode with English captions plus Spanish, French and German subtitles is four runs of 30 credits: 120 credits (12¢). Name the files consistently, such as episode.en.srt and episode.es.srt, which many desktop players pick up automatically when the file sits next to the video. On YouTube, each file is uploaded as its own language; the YouTube subtitles use case and the guide to translating a YouTube video cover that route. Showing two languages at once is possible on some players; see dual-language subtitles.

When a sidecar file isn't what you need

Some social apps and ad platforms don't accept subtitle files, and others display them inconsistently. In those cases the text has to be part of the picture, which mydubly doesn't do; import the SRT into an editor and export a captioned version from there. Embedded, sidecar or burned-in subtitles compares where each works. Subtitles also arrive after the recording, never during it: there are no captions for live streams.

If you only need the text for reading, the timestamped transcript is easier to work with than a subtitle file.

Translated subtitles in…

Subtitles in the spoken language

Subtitle and transcript formats

Learn more

Articles that go deeper on the technology and workflows behind this tool.

Frequently asked questions

Can I upload an existing SRT to translate it?

Not currently. Subtitles are generated from the audio of your video or audio file.

Do subtitles include sound descriptions?

No. They contain the spoken words only, without [music] or speaker tags.

Why do my dub-mode subtitles differ from the transcript-mode ones?

In dub mode some lines are shortened so the voice fits the original timing, and the subtitles follow the voice. Transcript-mode subtitles keep the full translation because nothing has to be spoken in time.

Can I adjust the subtitle timing?

Yes, in any subtitle editor. You can shift every cue at once if the whole file is early or late, or drag individual cues. The timing comes from where speech was detected, so large offsets usually mean the video was edited after the subtitles were made.

Can I make subtitles for an audio-only podcast?

Yes. Audio files also produce SRT and VTT, which is useful for video versions of an episode, audiograms and web players that show captions under the audio.