Why audio drama resists automatic dubbing
Without a picture, the voice does everything. Listeners know who is speaking, how old they are, where they are and what they feel almost entirely from how a line sounds. Drama scripts also rarely use dialogue tags such as "she said", because the casting makes them unnecessary. Put every character in one voice and the conversation collapses into a monologue that listeners can't follow. The general trade-off between one voice and many is covered in single voice vs multi voice dubbing; in audio drama, it is rarely a trade-off at all.
The sound design matters just as much. Footsteps on gravel, a door closing in another room, rain against a window, a music sting on a reveal: these carry plot and mood, and in a finished mix they are tangled with the dialogue. mydubly's dub keeps it by separating the dialogue out, but in a dense drama mix that separation is imperfect: faint traces of the original voices can remain, the sound design can sound slightly thinner, and any sung material loses its vocals. Then there is performance. Whispered confessions, shouted arguments, crying, characters talking over each other: this is where the drama lives, and it is exactly where synthetic voices are weakest.
What you need from your own production files
How far you can go depends on what you kept from production. The most useful materials are:
- The final script, with character names, stage directions and sound cues.
- The session from your audio editor, or exported stems: one dialogue track per character, plus music and effects.
- A music-and-effects mix, often called an M&E track: everything except the dialogue. Dubbing studios use it as the bed under a new-language cast.
If you only have the finished episode, mydubly separates the dialogue from the music and effects automatically when it dubs it, but on a dense drama mix the result is rarely as clean as a real M&E track; keeping background music in translated video explains why. For episodes you haven't mixed yet, exporting an M&E track and per-character dialogue stems at the end of each mix costs minutes and keeps every translation option open.
Translate from the script, not from the audio
For a drama, the script is a far better source than any transcript. It says who speaks each line, which no speech recognition output will, and it includes the stage directions and sound cues a translator needs to understand a scene. It also preserves what makes characters distinct on the page: a dialect, a formal or casual register, a verbal tic, a name that carries meaning.
Audio drama adaptation has a freedom screen dubbing lacks. There are no lips to match, so a translator can restructure sentences freely. Lines still have to fit the scene's timing, though: a line that must end before a door slams, a pause before a scream, a joke that lands just as the music stops. Give the translator the audio as well as the script, and mark the timing-critical moments.
Performers ad-lib, and final mixes drift from the script. A transcript of the finished episode shows what was actually said; comparing it with the script catches the changes before the translator works from an outdated page. An AI translated transcript can also serve as a rough draft for the translator to compare against, or as a quick way to judge whether a language is worth commissioning, as long as everyone treats it as a draft. Transcripts from mydubly have no speaker labels, so the script remains the source for who says what.
Three ways to release a translated drama
- Full recast
- A translated and adapted script, a new voice actor per character, a director, and a fresh dialogue mix over the original M&E track. Closest to the original experience, and the most expensive and slowest route.
- Per-character stems with stock voices
- Each character's dialogue stem is dubbed separately with a different stock voice, then mixed over the M&E track. Characters stay distinct and the sound design survives, but performances are flat and the cast is limited to the voices available.
- Translated transcripts or subtitles
- The original audio stays untouched and listeners read along with a translated transcript, or watch a video version with subtitles. Every performance is kept, at the cost of asking listeners to read.
Narrator-led shows are a fourth case. Anthologies read by one narrator, or fiction told in the first person by a single voice, behave much more like audiobooks, and one well-chosen stock voice can work; translating an audiobook covers that format.
Dubbing a small cast from character stems
This route keeps characters distinct without a recast. It takes careful editing, and it has a trade-off worth understanding before you start: each stem is translated on its own, so the translation sees one side of every conversation. Replies like "yes, that one", pronouns that refer to the other speaker, and the formality characters use with each other all lose their context. Plan for a fluent review.
- From your session, export one dialogue stem per main character. Each stem should run the full length of the episode from the same start point, with silence wherever that character isn't speaking. Export the M&E track the same way.
- Group minor roles onto shared stems by voice type, so the number of stems fits the voices available.
- Run each stem through the translator as an audio file, choosing the target language and a different stock voice for each character.
- Read each translated transcript. Long silences are the kind of input where speech recognition models occasionally invent text, so look for lines nobody spoke; Whisper hallucinations explains what to look for.
- Import the translated voice tracks into your audio editor, align them at the start, and mix them over the M&E track.
- Re-create the space. Synthetic voices arrive clean and dry, without the room, phone or radio effects you applied to the original dialogue, so add reverb and EQ per scene.
- Listen through the whole episode against the sound cues, and fix any line that now collides with an effect or music sting.
Non-verbal sounds deserve their own pass. Gasps, laughter, screams and breaths are usually dropped by speech recognition, so they disappear from the translated stems. They are also largely language-neutral, so you can often lift them from the original dialogue stems and place them back in the translated mix.
A mystery series, worked through
Suppose a six-episode mystery series runs about 25 minutes per episode, with a detective who also narrates, three other main characters and a handful of minor roles. The producer has stems and an M&E track and wants to test a Spanish version with one episode before committing to more.
The producer exports five stems: the detective, the suspect, the inspector, the landlady and one stem for the minor roles. Each gets its own voice, for example Male narrator for the detective, Female expressive for the suspect, Male calm for the inspector, Female for the landlady and Female energetic for the minor roles. Dubbing is billed on each file's length, silence included, so five 25-minute stems come to 5 × 25 × 50 = 6,250 credits ($6.25).
A Spanish-speaking listener reviews the mix. The story is easy to follow and the characters are distinct, but the landlady switches between the formal and informal "you" when addressing the detective, because her lines were translated without his. The climactic argument also sounds flat. The producer fixes the formality in a few patched lines, releases the episode clearly labeled as an AI-voiced translation, and publishes the Spanish translated transcript alongside it. If Spanish listeners respond, the plan is to commission a translated script and recast the series properly.
Limits and trade-offs of AI voices in drama
- Performance: stock voices read lines; they don't act them. Screams, whispers, sobs and comic timing come out flat, and climactic scenes suffer most.
- Cast size: mydubly has 8 voices, five female and three male, so larger casts have to double up roles.
- Age: all the voices are adult, so child characters can't be voiced convincingly.
- Character through accent: dialects and accents that define a character don't carry over.
- Songs: sung material should usually stay as it is, perhaps with translated text, rather than be read by a speaking voice.
- Rights: translating a script you licensed may need separate translation rights, and performer contracts may restrict derivative versions or AI voices. This is not legal advice; check your agreements and ask a qualified professional if unsure.
- Disclosure: tell listeners when a version uses synthetic voices, in the episode description and ideally in the audio.
Where mydubly fits for audio drama
The audio translator accepts MP3, WAV, M4A, AAC, OGG and FLAC files up to 2 hours long. For each file it returns a transcript and a timestamped transcript, SRT and VTT subtitles, and optionally a translated transcript; with AI dubbing, it also returns translated audio as an M4A file, timed to the original lines, with the original speech removed and any music and effects in the file kept under the new voice. One chosen voice speaks the whole file. It doesn't label speakers, clone voices or import scripts.
That makes it a fit for specific jobs: translated transcripts for listeners and accessibility, a rough draft to hand a translator, a scratch dub to test interest in a market, and the per-character stem workflow above. It is not a substitute for a directed recast of a performance-driven drama, and for a flagship release a human cast will serve the story better.
Where to start
Check what production files you have, and export an M&E track and character stems for at least one episode. Then try the cheapest test first: a translated transcript of one episode for a few fluent listeners, followed by a stem-based pilot through the audio translator if the story holds up. When you're ready to publish, running a multilingual podcast covers feeds and release planning, and the podcast translation use case covers interview-style shows.
Frequently asked questions
Can mydubly give each character a different voice automatically?
No. mydubly uses one chosen voice for each file and doesn't identify speakers. The workaround is to export each character's dialogue as a separate stem from your session, run each stem with a different voice, and mix the results in your audio editor. That only works if you have the stems; a finished mix can't be split by character this way.
Is a single-voice version of an audio drama ever worth releasing?
Sometimes, if you rework the script into a narrated adaptation. Adding dialogue tags and description turns a drama into something closer to an audiobook, which one voice can carry. Simply dubbing the original drama with one voice usually leaves listeners unable to tell characters apart, because the script relies on casting to show who is speaking.
How does the cost of a full recast compare with an AI version?
They are priced in completely different ways. A recast involves a translator or adapter, voice actors for each role, a director, studio time and a new mix, and rates vary widely by country and union status, so get quotes. An AI dub is priced per minute of each file: 50 credits per minute, so a 25-minute stem costs 1,250 credits ($1.25).
Should translated episodes go in the same podcast feed?
Usually not. Listeners subscribe to a feed expecting one language, and mixed-language feeds confuse both people and podcast apps. A separate feed per language, with its own title and description in that language, is the common approach. Running a multilingual podcast covers naming, hosting and cadence for language-specific feeds.
Can a translated transcript make my drama accessible to more people?
Yes. A transcript helps deaf and hard-of-hearing listeners and people who prefer to read, and a translated transcript extends that to other languages. mydubly's transcripts don't label speakers, so add character names and important sound cues from your script before publishing, and have a fluent reader check the translation.