MP3 files

MP3 to Text & MP3 Translator

Drop in a MP3 file as it is — no converting first. Get the words back as text and subtitles, or translate the speech into another language.

01 — Upload

Your audio

Audio + text outputText only — timestamped transcript with no audio in results.

The audio file stays in this browser. Language is required: pick the spoken language for a plain transcript, or another language to translate it.

02 — Result

Your transcript

Your timestamped transcript will appear here.

Waiting for upload

25 languages available

About MP3 files

MP3 is the most widely used audio format — podcasts, recorders, voice apps and music all use it.

Common sources: podcast episodes, digital voice recorders, audio exported from editors, call recordings.

Every browser reads MP3.

MP3 to text

Low-bitrate MP3s (phone calls, old recorders) still transcribe, but clean audio at 64 kbps or higher gives noticeably better results.

52-minute MP3

A podcaster turns a 52-minute MP3 episode into a transcript for show notes and an accessible text version.

Cost: 52 credits (5.2¢).

Translating a MP3

Choose audio-and-text output for a translated voice track; choose text-only for a translated transcript at a much lower price.

18-minute MP3

A language teacher translates an 18-minute English MP3 interview into Spanish audio and a Spanish transcript for a class.

Cost: 900 credits (90¢) with a translated voice, or 18 credits (1.8¢) for translated text only. Switch Full translation output on in the tool above for a voice track.

What the bitrate tells you before you start

MP3 discards sound it predicts you won't notice, and the bitrate decides how much goes. A 128 kb/s stereo podcast keeps plenty of detail for speech. A 32 kb/s mono file from an old dictaphone or a call recorder keeps far less, and consonants such as s, f and th smear together, which are exactly the sounds a recognizer uses to tell words apart. The bitrate is shown in the file's info panel (Get Info on a Mac, Properties → Details on Windows). Re-encoding a low-bitrate MP3 at a higher bitrate, or converting it to WAV, restores nothing; the detail is already gone. Audio bitrate for speech recognition explains why.

Variable-bitrate MP3s are normal and fine. One quirk: a VBR file missing its header can show the wrong duration in some players. Since length is what's billed and limited, re-save a file whose length looks wrong with an audio editor or an MP3 repair tool before uploading.

Podcast episodes: intros, music beds and inserted ads

Podcast files bring their own clutter. A music intro with no speech can leave a stray line at the top of the transcript, because recognition models sometimes guess words over music (why that happens); delete it or trim the intro. Speech over a loud music bed transcribes less cleanly than dry speech, so if you produce the show, export a voice-only stem for the transcript and keep the full mix for listeners.

An episode downloaded from a public feed, rather than exported from your editor or host, may contain dynamically inserted ads that change from listener to listener. They'll end up in the transcript. Use the master file you uploaded to your host instead.

Host and guest names won't appear, because transcripts don't label speakers. Most podcasters add names during the clean-up pass, using the timestamped transcript to jump to any line where the speaker is unclear.

Mono, stereo and call recordings

Stereo interview with host left and guest right
Recognizers typically work on a mono mix, so a voice that exists on one channel only ends up quieter. Check neither side is far below the other.
Mono call recording at 32 kb/s
Usually transcribes, but names, numbers and addresses need extra review time.
Recorder MP3 at 128 kb/s or more
Bitrate is rarely the limit here; distance from the speaker and room echo matter more.
MP3 extracted from a video
Works, but if you still have the video, upload that and skip one round of lossy compression.

Call recordings deserve a note of their own. A traditional phone line carries only a narrow band of frequencies, so even a high-bitrate MP3 of a call can't contain the detail a studio recording has; the bitrate of the file and the quality of the line are separate limits. Phone call recording transcription covers what recognition does with narrowband audio and the consent questions to settle before recording.

Which downloads matter for audio

Transcript mode returns a timestamped .txt, a plain transcript, and SRT and VTT subtitles. The SRT is more useful for audio than it sounds: many audiogram tools and video editors import it to caption a short clip from an episode. With a target language picked, the plain transcript and subtitles come back translated. Dub mode turns an MP3 into translated audio delivered as M4A rather than MP3, with the original speech removed and your music beds kept under the new voice, plus the original-language transcript and subtitles that match what the new voice says. If your podcast host only takes MP3, convert the M4A once in any audio editor.

Budgeting a podcast season

A 10-episode season

Ten 45-minute episodes transcribed for show notes come to 450 credits (45¢), inside a first $1 top-up. A Spanish dub of the 3-minute trailer adds 150 credits (15¢); dubbing a full episode would be 2,250 credits ($2.25) each.

Podcast transcription covers the full workflow, and show notes from a transcript turns the text into chapters and quotes. Still choosing a recording format? MP3 vs WAV for transcription and the WAV page explain when the uncompressed file is worth keeping.

Using the tool on this page

  • The spoken language is detected automatically; you select the output language — the spoken language itself for an untranslated transcript, or another language for a translation. Files can be up to 2 hours long.
  • Audio is sent over HTTPS in small chunks and deleted with its transcripts within 30 minutes of delivery (how files are protected).
  • You can download a timestamped transcript, subtitles (SRT and VTT) and plain text. Every output, step and limit is explained on the audio to text page.

Frequently asked questions

Is there a size limit for MP3 files?

The limit is length, not size: up to 2 hours per file.

What format is the translated audio?

The translated speech is delivered as an M4A (AAC) audio file, plus the transcript and subtitles as text files.

Can I get song lyrics from a music MP3?

mydubly is built for speech. Singing over instruments is hard for speech recognition, so expect gaps and wrong words in sung sections. Spoken-word recordings such as podcasts, talks, interviews and calls are what it's meant for.

My MP3 runs longer than 2 hours. What should I do?

Split it at a natural pause into parts under 2 hours. A lossless MP3 cutter avoids re-encoding. Each part is billed separately, and the second part's timestamps begin from zero, so add the cut point when you merge them.

Can I transcribe an episode straight from a podcast app or a link?

No. The file has to be on your device, and there's no link import. Download or export the MP3 first, ideally the master from your host for the inserted-ad reason above.