Where a wrong name can come from
In a speech-to-speech dub, a name passes through three stages before you hear it: recognition writes it down, translation decides what to do with it, and speech synthesis reads it aloud. Each stage can break it in a different way, and each needs a different fix.
- Recognition
- The name was misheard and written as different words, such as a surname rendered as a common word
- Translation
- The name was translated, transliterated or inflected when it should have stayed as it was
- Synthesis
- The text was correct, but the voice guessed the wrong sounds for an unfamiliar spelling
- Expectation
- The voice used a valid pronunciation in the target language that differs from how the owner says it
The last row is worth taking seriously. Many brands and names are pronounced differently in different countries, and a local pronunciation is not automatically an error.
Why text to speech guesses
A speech synthesis model turns letters into sounds using patterns learned from training data. Common words have reliable patterns; invented brand names, rare surnames and names from other languages do not. Faced with an unfamiliar spelling, the model applies the spelling-to-sound habits of the language it is speaking, so an English brand read by a German or Spanish voice may follow German or Spanish rules. How models handle this across languages is covered in multilingual text to speech.
Acronyms add another layer of ambiguity. The same letters can be read as a word or spelled out, letter names differ between languages, and capitalization alone does not always tell a model which reading is meant. Product codes mixing letters and digits, such as X2R-40, are similarly open to interpretation.
How recognition and translation errors carry through
The voice only reads what it receives. If recognition writes a surname as an ordinary word, the translation will translate that word and the voice will pronounce the translation perfectly, producing something that sounds fluent and is completely wrong. Long recordings are recognized in chunks, so the same name can even be spelled two ways in one video.
Translation can also change a correct name: by translating a descriptive brand name, adapting it to a local script, or adding grammatical endings in languages that inflect nouns. Deciding which names should stay untouched is a translation question, covered in depth in translating names and technical terms; this article focuses on what you hear.
Diagnose the stage in this order
- Note every timestamp where a name, brand or acronym sounds wrong in the dub.
- Find the same moments in the source-language transcript. If the name is misspelled there, recognition is the cause.
- If the source transcript is right, check the translated transcript or translated subtitles. If the name changed there, translation is the cause.
- If both transcripts show the name correctly, the voice guessed, and the fix happens at the audio stage.
- Check whether the error repeats at every mention or only in some, which tells you whether a single replacement clip can be reused.
Fixes for each cause
Recognition errors. Improve how the name is heard. Have the speaker say it clearly and unhurriedly at its first mention, reduce background sound, and keep the microphone close; the guide to improving transcription accuracy lists the main levers. In a script you control, a speaker can also spell a critical name once, though that adds words viewers will hear.
Translation errors. If the translated subtitles are what viewers rely on, correct the name in the SRT file before publishing. For the voice, the line has to be replaced in the audio, as below.
Synthesis guesses. When you control the text input, as in a standalone text-to-speech tool, you can steer pronunciation: respell the name phonetically for the target language, hyphenate syllables, spell out an acronym with spaces or periods, or write numbers as words. Some engines support SSML, a W3C markup standard whose phoneme element specifies a pronunciation directly. Support varies by engine, so check its documentation.
Replacing a line in an editor. Download the translated audio, import it into a video or audio editor under the picture, and cover each wrong mention with a corrected clip: one generated by a text-to-speech tool where you can respell the name, or recorded by a fluent speaker. Match loudness, keep the clip within the original line's timing, and export. Editing dubbed audio in a video editor covers mixing, levels and export.
A worked example: a brand name in a Spanish dub
Suppose a 6-minute product tutorial introduces a fictional app called Lyvanta, presented by a founder named Siobhan Kearney. In the Spanish dub, the app name sounds roughly right, but the founder's first name is read letter by letter in Spanish fashion, and in one chunk the app appears as Le Vanta. The English transcript shows both spellings, so the Le Vanta error came from recognition, while Siobhan was spelled correctly and the voice simply guessed. The team records a Spanish-speaking colleague saying the founder's name, replaces three mentions in an editor, and fixes the one misheard app name the same way. The dub itself cost 300 credits (30¢); the fixes took about twenty minutes.
Notice the two errors needed different diagnoses even though both sounded like pronunciation problems.
Limits and trade-offs of each fix
- Phonetic respelling works per engine and per voice; a respelling that helps one voice can make another worse.
- Spliced replacement lines can differ in tone from the surrounding voice, especially when recorded by a person rather than the same synthetic voice.
- A local pronunciation of a brand may be acceptable or even preferred by viewers in that market; ask a native speaker before replacing it.
- Fixing the audio does not fix the subtitles, and the reverse; check both deliverables.
- Rare names in noisy recordings may be misheard repeatedly, and replacing many lines by hand becomes slow for long videos.
What mydubly produces and what it does not
mydubly's dub is generated from the recognized and translated transcript: Whisper transcribes the speech, a translation engine translates the text, and speech synthesis reads it in one of 8 stock adult voices, with one chosen voice speaking the whole video. Short fragments are merged into sentences before synthesis. You receive the translated audio, the new voice over the original music and effects, as an M4A file and a translated video carrying that same audio, along with transcripts and subtitles you can use to diagnose the stage that went wrong.
There is no pronunciation dictionary, no way to edit the text before the voice is generated, and no voice cloning or lip-sync. mydubly cannot import a corrected SRT to re-voice it. Corrections therefore happen either upstream, in the recording, or downstream, in an editor with the downloaded audio. The AI dubbing quality checklist includes names and numbers in its review pass.
When another workflow fits better
If a video is dense with names that must be exactly right, such as a medical lecture or a legal briefing, a script-first workflow in a text-to-speech tool with pronunciation controls, or a human voice actor, will save time over repeated fixes. For a quick test of how your names come through, dub a short excerpt in the video translator before committing a whole series.
Next step: test the names first
Pick a two-minute excerpt with the most names and acronyms in your video, dub it with AI dubbing, and run the three-stage diagnosis above. Dubbing costs 50 credits per minute with a 2-minute minimum, so that test costs 100 credits (10¢).
Frequently asked questions
Why does the AI voice read an acronym as a word?
Speech models learn from text where some acronyms are pronounced as words and others are spelled out, and capital letters do not reliably signal which. When you control the input text, writing the letters with spaces or periods usually forces a letter-by-letter reading. In a speech-to-speech dub, replace the line in an editor if it matters.
Can I add a custom pronunciation to a mydubly dub?
No. mydubly has no pronunciation dictionary and does not let you edit text before the voice is generated. To fix a name, record or generate a corrected clip elsewhere and place it over the original line in an editor using the downloaded translated audio.
Is a different pronunciation in another language always wrong?
Not always. Many international brands and place names have established local pronunciations, and viewers may find the original pronunciation odd. Ask a native speaker of the target language which version audiences expect, and check whether the brand owner has a preference.
Why is the same name spelled two different ways in my transcript?
Long recordings are recognized in pieces, and each piece is processed without knowledge of how the name was spelled earlier. An unfamiliar name can therefore be heard differently from one section to the next. A clear first mention and cleaner audio make consistent spelling more likely.
Do I need to fix the subtitles as well as the voice?
Usually, yes. The translated subtitles and the voice come from the same translated text, so a misheard or mistranslated name typically appears in both. Correct the SRT or VTT file in a text or subtitle editor before publishing, and replace the audio separately.