Audio to text

Audio to Text Converter

Drop in a recording and get the words back as text — plain, timestamped, or as subtitles. Works with MP3, WAV, M4A, AAC, OGG, FLAC and WebM, in 21 languages.

01 — Upload

Your audio

Audio + text outputText only — timestamped transcript with no audio in results.

The audio file stays in this browser. Language is required: pick the spoken language for a plain transcript, or another language to translate it.

02 — Result

Your transcript

Your timestamped transcript will appear here.

Waiting for upload

25 languages available

Recordings it handles well

Speech to text that's built for files

This is speech-to-text for recordings, not live dictation. That trade-off means you can process up to two hours at once, get timestamps on every line, and pay a fraction of a cent per minute.

Example

A 52-minute podcast episode costs 52 credits (5.2¢); a 3-minute voice memo is billed at the 5-credit minimum.

How it works

  1. Drop in an audio file. It's read in your browser and sent in small chunks.
  2. The audio is processed in 30-second chunks: speech is transcribed and the spoken language is detected automatically.
  3. Download the results. Audio and text on our servers are deleted after delivery, within 30 minutes at most.

What you get

Transcript
Plain text (.txt) of what was said
Timestamped transcript
Every line prefixed with its time, e.g. [12:04]
Subtitles
SRT and VTT, in the language you select (select the spoken language for untranslated subtitles)

Supported files

  • Audio: MP3, WAV, M4A, AAC, OGG, FLAC, WebM.
  • Length: up to 2 hours per file.
  • Languages: speech in 21 languages is recognised automatically; translation and voices cover the same 21.

Privacy

Audio is sent over HTTPS in small chunks and deleted after delivery, within 30 minutes at most, together with transcripts and voice files. Nothing is used to train models. Your account keeps only the file name, length, languages and credits used. Details in the privacy policy.

Good to know

  • No speaker labels — transcripts don't say who is speaking.
  • Files only: no links from YouTube or social apps, and no live audio.
  • Subtitles come as SRT/VTT files rather than burned into the picture.

Where recordings come from, and what to send

There's no need to convert a recording before uploading. MP3, WAV, M4A, AAC, OGG, FLAC and WebM audio are all accepted, and converting a compressed file to a larger format doesn't restore anything the compression removed.

iPhone Voice Memos and many Android recorders
Usually M4A; send it as is. M4A files
Field recorders and audio software
Often WAV: large but fine to upload. WAV files
Published podcast episodes
Usually MP3. MP3 files
Video calls and screen recordings
Upload the video itself on the video to text page at the same price; there's no need to extract the audio first

MP3 or WAV for transcription? covers what to keep for archives.

A five-minute check that saves a re-run

  • Listen to the first minute and one from the middle on headphones. If you struggle to follow the speech, the transcript will struggle too.
  • Check for one-sided audio. A microphone recorded to only one channel sounds fine on speakers but odd on headphones; fix it before upload if you're unsure how it will be mixed.
  • If each person was recorded on a separate microphone file, mix them into one track first. Each upload is transcribed as a single file.
  • Phone-line recordings carry a narrow band of frequencies, so expect more errors on names and numbers and budget more review time.

The five-minute audio check and mixing several microphones go into the fixes.

From transcript to finished interview document

  1. Run transcript mode and download the timestamped transcript.
  2. Add speaker labels by hand. Transcripts don't say who is speaking; how diarization works explains why this is hard and the practical workarounds.
  3. Verify every quote you plan to use by jumping to its timestamp and listening.
  4. Pick a transcription style before editing. Verbatim or clean verbatim explains when fillers and false starts matter.
  5. Replace names of participants with codes before the file leaves your project folder, if your consent terms require it.

For podcasts the route is shorter: the plain transcript becomes show notes, the timestamped transcript gives you chapter times, and the SRT or VTT file covers video versions of the episode and web players that display captions.

Costs for typical recordings

75-minute research interview

75 credits (7.5¢).

A week of voice memos: twelve 2-minute files

Uploaded one by one, each is billed at the 5-credit minimum: 60 credits (6¢). Joined into one 24-minute file first: 24 credits (2.4¢).

Two-hour board meeting

120 credits (12¢), the longest single file accepted.

File transcription, dictation or a human transcriber

Three routes turn speech into text, and they suit different work. Dictation software types as you speak, which is ideal for drafting an email and useless for a recording you already have. File transcription, which is what this page does, takes a finished recording and returns the whole text with timestamps. A human transcriber costs far more and takes longer, but handles heavy crosstalk, strong accents on poor audio, and strict verbatim formats such as legal transcripts more reliably.

A common middle path is to run the recording here first, then pay a person only to check and correct the passages that matter. Dictation or transcription? and AI vs human transcription compare the routes in more detail.

Recordings that need extra care

  • Crosstalk: when two people speak at once, recognition usually follows one voice and loses the other. Ask people to take turns if you're the one recording.
  • Language coverage: speech in 21 languages is recognised. Of Indian languages, only Hindi is covered, so a recording in another Indian language won't transcribe properly.
  • Very long sessions over two hours need splitting at a pause; transcribing long audio files covers where to cut.
  • Digitised tapes and old archive recordings: hiss and uneven speed make recognition harder. Capture at a healthy level and go easy on noise reduction, which can remove parts of the speech along with the hiss; from cassette to transcript covers the capture settings.

Job-specific workflows are in interview transcription, podcast transcription and voice memo to text. Spoken-language pages such as Spanish and German cover what to check for each language, and the audio transcription guide walks through a first run.

Transcribe speech in…

By file format

Learn more

Articles that go deeper on the technology and workflows behind this tool.

Frequently asked questions

Is this real-time speech to text?

No. It transcribes recorded files. For dictation as you speak, use your device's built-in dictation.

Can it tell who is speaking?

No. Transcripts don't include speaker labels. For interviews, add names by hand as you review; the timestamps make it quick to find where the voice changes.

Can I transcribe a recorded phone call?

Yes, if you have the recording as a file and the consent your jurisdiction requires. Phone audio is narrowband, so review names and numbers carefully. Transcribing phone calls covers exporting recordings from common apps.

Is a WAV file transcribed more accurately than an MP3?

Only if the WAV was recorded that way. Converting an MP3 to WAV makes the file bigger without adding detail, so upload whichever file you already have.