Speech translation and on-screen text are separate problems
A speech pipeline hears the video; it does not read it. Recognition turns the soundtrack into text, translation converts that text, and the output is a voice track or a subtitle file. The pixels are never analyzed. Text that appears in the picture, such as a slide headline, a speaker's name in a lower third, a menu label in a screen recording or a sign in the background, passes through untouched.
Reading text out of video frames is possible in principle with optical character recognition, but replacing it convincingly means redrawing the frame, matching fonts, handling motion and fitting longer translations into the same space. That is closer to motion graphics work than to translation, which is why most workflows handle it with editing rather than automation.
Start by listing the text in your video
Before choosing a method, scrub through the video and note every piece of text and what job it does. A five-minute pass usually reveals that only a handful of items actually need translating.
- Write down each slide, title card, lower third, callout and caption, with its timestamp.
- Mark whether the speaker also says the same thing aloud. If they do, the speech translation already covers it.
- Mark whether the text is essential to understanding (a step number, a warning, a price) or decorative (a logo, a background sign).
- Note whether you still have the source files: the slide deck, the editing project or the graphics template.
That list decides everything else. Essential text with source files available can be re-exported; essential text without source files needs an overlay or a subtitle note; decorative text can usually stay as it is.
Option 1: re-export slides and graphics in the target language
If the video was built from a slide deck or graphics templates, translate the source and render it again. For a lecture or webinar recorded as slides plus voice, this gives the most polished result: the translated slides replace the originals frame for frame, and the dubbed or subtitled audio sits on top.
The translated transcript is a useful starting point here, because it already contains the speaker's terminology in the target language. Keep slide text consistent with the words the voice uses, or viewers will see one term and hear another. Expect some layout work, since translations are often longer and may not fit the original text boxes.
Option 2: narrate the text so it is translated with the speech
For new recordings, the simplest fix is for the presenter to say what the screen shows. "Step three: open Settings and choose Privacy" spoken aloud becomes speech, and speech is translated. Viewers of the dubbed version hear the instruction in their language even though the screen still shows the original words.
This works especially well for screen recordings and software tutorials, where the interface cannot be changed anyway. It also improves accessibility in the original language, because viewers who cannot see the screen clearly get the same information. Guidance for this kind of content is in screen recording translation.
Option 3: overlay translated text in a video editor
When you cannot regenerate the source, cover or sit alongside the original text in an editor. Common patterns:
- Place a translated lower third directly over the original one, using a solid background so the old text is hidden.
- Add a translated label next to an interface element instead of hiding it, so viewers can match it to the real software.
- Blur or crop out a burned-in caption line and add new text in the same position.
- Use a translated title card at the start of a section instead of altering every frame of an animated title.
Overlays take the most manual effort per language, so reserve them for text that is essential and repeated on screen long enough to read.
Option 4: put the text in the subtitle file
Subtitle files can carry more than speech. Film and TV releases use a practice often called forced narrative subtitles: short subtitles that translate signs, letters and on-screen captions for viewers who do not otherwise need subtitles. You can do the same thing by adding cues to the translated SRT or VTT file at the moment the text appears.
A product demo shows the on-screen caption "Dr. Ana Ruiz, Head of Research" from 00:00:04 to 00:00:07 while she starts speaking. In the Spanish SRT file, add a cue for that window reading "[En pantalla: Dra. Ana Ruiz, directora de investigación]" just before her first spoken line, then save and re-upload the file.
Keep these cues short and visually distinct, for example in square brackets, so viewers can tell them apart from dialogue. If the cue overlaps a spoken line, put the note first or merge both into one two-line cue. This is cheap, fast and works on any platform that accepts subtitle files, though it only helps viewers who turn subtitles on.
Which option suits which kind of text
- Slide headlines and bullet points
- Re-export from the deck; narrate the key points in new recordings
- Lower thirds with names and titles
- Overlay in an editor or add a bracketed subtitle cue
- Software interface labels
- Narrate the action; add a translated label beside the element where it helps
- Burned-in captions of the dialogue
- Cover or crop them, then supply translated subtitles as a file
- Warnings, prices and legal text
- Always translate, by re-export or overlay, and have it reviewed
- Background signs and decorative text
- Usually leave as is, or add a subtitle note if the plot depends on it
Trade-offs and what can't be fixed in post
Every option has a cost. Re-exporting needs the source files and layout time per language. Narration only helps future recordings. Overlays are manual and can look patched if fonts and colors do not match. Subtitle notes are invisible to anyone watching with subtitles off, which includes most viewers of a dubbed version.
Some text is effectively stuck. Burned-in captions under moving footage, text on fast-moving products and handwriting on a whiteboard are hard to cover cleanly. For those, a bracketed subtitle note or a spoken explanation in a re-recorded intro may be the only realistic route. For new projects, the cheapest fix is upstream: making videos translation-ready explains how to keep text editable from the start.
Where mydubly fits: speech only
mydubly translates speech. It does not read, translate or replace text that appears in the picture, and it never alters the video frames; the translated MP4 keeps your original video stream and swaps only the audio. Subtitles are delivered as separate SRT and VTT files rather than burned in, which is exactly what you need for Option 4: open the translated SRT in a text editor, add your bracketed cues, and upload it alongside the video.
The translated transcript is also useful for Options 1 and 3, as a reference for the terms the voice uses, so your slides and overlays match what viewers hear. For videos where most meaning is on screen, a subtitle-only run in transcript mode costs 1 credit per minute and gives you the files to edit; a 30-minute slide webinar costs 30 credits, or 3¢. The use case for webinar translation covers slide-heavy recordings further.
Next step: build your subtitle file first
Whichever option you choose for the picture, translated subtitles are the backbone, because they are where you add notes and check terms. Generate them with the subtitle generator, add cues for essential on-screen text, and then decide which items deserve a re-export or an overlay. If you are unsure how to attach the edited file on your platform, the guide on adding subtitles to a video walks through it.
Frequently asked questions
Can AI automatically translate text that appears in a video?
Some research and specialist tools attempt it by reading the text from frames and redrawing it, but results depend heavily on fonts, motion and backgrounds. Speech-focused translators, including mydubly, leave the picture untouched, so on-screen text needs one of the editing approaches described here.
How do I translate burned-in subtitles in a video?
You cannot edit burned-in subtitles as text, because they are part of the image. Cover or crop the original line in an editor, then supply translated subtitles as a separate SRT or VTT file generated from the speech.
What are forced narrative subtitles?
They are subtitles that only translate on-screen text and foreign-language moments, such as signs or a caption naming a location, for viewers who do not otherwise use subtitles. You can mimic them by adding short bracketed cues to your translated subtitle file.
Should I translate software interface text in a tutorial?
Usually not in the picture, because viewers will see the original labels in the real product unless it is localized too. Narrate each action clearly so the instruction is translated, and add a translated label beside an element only where confusion is likely.
Will the translated voice read my slide text aloud?
Only if the speaker said it in the original recording. The voice track translates what was spoken; slide text that was never read aloud stays silent and untranslated.