Audio translator

Audio Translator

Translate a recording into another language — as a new spoken track you can publish, or as a translated transcript you can read. Works with MP3, WAV, M4A, AAC, OGG, FLAC and WebM.

01 — Upload

Your audio

Audio + text outputText only — timestamped transcript with no audio in results.

The audio file stays in this browser. Language is required: pick the spoken language for a plain transcript, or another language to translate it.

02 — Result

Your transcript

Your timestamped transcript will appear here.

Waiting for upload

25 languages available

Listen or read?

Translated voice
An M4A in the target language, from 50 credits per minute. A 30-minute episode costs 1,500 credits ($1.50).
Translated transcript
The text in both languages, at 1 credit per minute. The same episode costs 30 credits (3¢).

If you only need to understand a recording — a voicemail, an interview, a lecture — the translated transcript is usually enough and far cheaper.

Common uses

  • Podcasts launching a feed in a new language.
  • Recorded sermons and talks for multilingual communities.
  • Interviews and voice messages in a language you don't speak.

How it works

  1. Drop in an audio file. It's read in your browser and sent in small chunks.
  2. The audio is processed in 30-second chunks: speech is transcribed and the spoken language is detected automatically.
  3. Each segment is translated into your chosen language.
  4. The translation is read aloud in the voice you chose, and each line is fitted to the original timing — sped up by at most 1.15× with natural pitch.
  5. AI vocal separation removes the original speech, and the translated voice is mixed over the remaining music and effects, which dip automatically while the new voice speaks.
  6. Download the results. Audio and text on our servers are deleted after delivery, within 30 minutes at most.

What you get

Translated audio
M4A with the same mix: translated voice over your original music and effects
Transcript
Plain text (.txt) of what was said
Timestamped transcript
Every line prefixed with its time, e.g. [12:04]
Subtitles
SRT and VTT, in the language you select

Supported files

  • Audio: MP3, WAV, M4A, AAC, OGG, FLAC, WebM.
  • Length: up to 2 hours per file.
  • Languages: speech in 21 languages is recognised automatically; translation and voices cover the same 21.

Privacy

Audio is sent over HTTPS in small chunks and deleted after delivery, within 30 minutes at most, together with transcripts and voice files. Nothing is used to train models. Your account keeps only the file name, length, languages and credits used. Details in the privacy policy.

Good to know

  • One of 8 stock voices reads the whole translation — no voice cloning and no lip-sync.
  • Music and sound effects are kept by AI vocal separation, which isn't perfect: in dense mixes faint traces of the original voice can remain, and sung lyrics are removed along with the speech.
  • No speaker labels — transcripts don't say who is speaking.
  • Files only: no links from YouTube or social apps, and no live audio.
  • Subtitles come as SRT/VTT files rather than burned into the picture.

Three reasons people translate a recording

  • To understand it. A voice message from a relative, an interview with a source, a lecture in another language: transcript mode with a target language gives a readable translation for 1 credit per minute, and nothing needs to be spoken.
  • To publish it. A podcast, course audio or narration for listeners in another language needs dub mode and its translated audio file.
  • To analyse it. Researchers with interviews in a language they don't read use the translated transcript as a working copy and have key passages checked by a fluent speaker. Translation in qualitative research covers the method decisions.

What the translated audio file is, and isn't

Dub mode on an audio file gives you an M4A: one stock voice reading the translation, with each line placed where the original was spoken. Because of that placement, the translated file follows your original's timeline, which makes it easy to line up with your own stems in an editor if you dubbed a speech-only export and want to build your own mix.

The original speech is removed by AI vocal separation, and the intro music, jingles, ad stings and ambient sound that remain are mixed under the translated voice, lowered while it speaks. The file is mono. Separation isn't perfect: in dense or loud mixes faint traces of the original voice can remain and the background can sound slightly thinner, and a jingle with sung lyrics keeps its instruments but loses its vocals. Every speaker becomes the same voice. For a solo show or a narrated course that's rarely a problem. For a two-host conversation, listeners lose the cue of who is talking, so consider stronger verbal hand-offs or keeping that format as translated text. Fiction with several characters is hit hardest; translating audio dramas explains when a full recast is the honest answer.

Preparing the recording before you translate it

Translate the cleanest version of the recording you have. Recognition, and the separation of voice from background, are more reliable when music sits low under speech. If you upload the published mix, its music bed is kept under the translated voice; if you upload a speech-only export instead, recognition improves but the translated file has no music, so you'd add your intro and bed back in your editor. Cut ad reads and sponsor segments that won't run in the new language, so you don't pay to translate them. If hosts were recorded on separate tracks, export one mixed track, since each upload is processed as a single file. And trim long silences or pre-show chatter at the start; billing is per minute of the file you upload.

If you can't read the target language yourself, ask a fluent listener to check the first episode, or use back translation as a rough spot check before anyone else hears it.

Launching a podcast feed in another language

  1. Dub one representative episode first and ask a fluent listener for notes on meaning, names and pacing.
  2. Decide on a separate feed per language, which is the common setup; publishing a podcast in several languages covers naming and hosting.
  3. Pick one voice and keep it for every episode so the show sounds consistent.
  4. Listen to the translated M4A with its music bed in place. If you want stereo music or your own levels, dub a speech-only export instead and mix the translated voice with the clean intro and music stems in your editor (optional).
  5. Translate the show notes. The dub-mode subtitle file holds the translated script as spoken; strip the cue numbers and timings to get plain text to edit from.
  6. Publish on a regular schedule. The podcast translation use case covers the rest of the workflow.

Costed scenarios

Weekly 40-minute interview podcast, one extra language

2,000 credits ($2.00) per episode for the translated voice, subtitles and transcript.

6-minute voice message from a relative

Transcript mode with your language as the target: 6 credits (0.6¢).

70-minute audiobook chapter

3,500 credits ($3.50). Files must be two hours or less, so split long books by chapter. Check that you hold the rights to translate the text and recording first; [translating an audiobook](/blog/translating-an-audiobook) covers both.

When a machine translation of a recording isn't enough

  • Legal and official use: courts, immigration and evidence usually require a certified or sworn translation, which an AI transcript is not. A machine draft can still help you prepare; see certified translation and AI.
  • Health and safety instructions: meaning errors carry real risk, so treat the output as a draft for a qualified translator.
  • Conversations as they happen: mydubly works on recorded files only, with no live interpreting.

For text without a voice in the same language, the audio to text converter is the cheaper route. Target-language pages such as Portuguese and Chinese cover what changes in that language, and MP3 files covers the most common podcast format.

Translate audio into…

By file format

Learn more

Articles that go deeper on the technology and workflows behind this tool.

Frequently asked questions

Can I translate a voice message?

Yes, if you can save it as an audio file. Short files are billed at the per-file minimum.

Will both podcast hosts get their own voice?

No. One stock voice reads the whole translated track. If telling hosts apart matters, offer translated subtitles or a translated transcript alongside the audio.

Can I keep my intro music in the translated episode?

Yes. The original speech is separated out and your intro, outro and music bed stay under the translated voice, lowered while it speaks. Sung lyrics are removed along with the speech, so a sung jingle keeps only its instruments. If you have your music stems, dubbing a speech-only export and mixing the translated voice with them in an editor gives the cleanest result.

Can the recording be in any of the 21 languages?

Yes. The spoken language is detected from the audio, and any of the 21 languages can be the target, so a Korean recording can become Spanish without passing through a step you have to manage.