Guide

How to Transcribe Audio to Text

Transcribing audio comes down to three steps: get the recording into a file you can open, run it through automatic transcription, then clean up the few things machines get wrong.

01 — Upload

Your audio

Audio + text outputText only — timestamped transcript with no audio in results.

The audio file stays in this browser. Language is required: pick the spoken language for a plain transcript, or another language to translate it.

02 — Result

Your transcript

Your timestamped transcript will appear here.

Waiting for upload

25 languages available

Step 1: get the recording as a file

iPhone Voice Memos
Tap the memo → Share → Save to Files or AirDrop. You get an M4A.
Android recorder
Share the recording or find it in Files → Recordings; usually M4A or MP3.
Field recorder
Copy the WAV from the SD card. See WAV to text.
Podcast
Export the final mix as MP3 or WAV.
Zoom
Cloud recordings include an audio-only M4A.

Step 2: transcribe it

  1. Open the audio to text converter.
  2. Drop in the file — MP3, WAV, M4A, AAC, OGG, FLAC and WebM all work, up to 2 hours.
  3. Select the spoken language for an untranslated transcript, or another language for a translated transcript as well.
  4. Download the plain or timestamped transcript.
Cost

A 30-minute interview costs 30 credits (3¢). Files under 5 minutes are billed at the 5-credit minimum.

Step 3: clean it up

  • Search for names and specialist terms — these are the most common errors.
  • Add speaker names where it matters; automatic transcripts here don't label speakers.
  • Check every quote against the recording using the timestamps before you publish it.

What affects accuracy

A phone held close to the speaker beats an expensive microphone across the room. Background music, echo and crosstalk hurt most. Read how to improve transcription accuracy before your next recording.

Recordings that need extra preparation

Clean single-speaker audio needs nothing but an upload. These recordings benefit from a few minutes of work first:

Phone call recordings
Phone audio is narrow-band and loses detail. Transcribe the call recorder's own file rather than a copy played through a speaker and re-recorded; see phone call recording transcription.
Longer than 2 hours
Split at a natural pause into parts under 2 hours, and note each part's start time so you can map timestamps back to the full recording.
Speech on one stereo channel
Listen in mono. If the other channel holds only noise or hum, export a mono file from the voice channel before uploading.
Very quiet recordings
Normalise the level in an audio editor first; whispered or distant speech is where words go missing. More in transcribing quiet audio.
Digitised tapes
Trim leader hiss and long blank stretches at the start and end, which can produce stray text and add billed minutes.
Dozens of short memos
Each file is billed at least 5 credits, so join short memos into one file with a second of silence between them.

Costing a batch before you start

One 1-minute voice memo
5 credits (0.5¢), because of the per-file minimum
Ten 1-minute memos joined into one file
10 credits (1¢) instead of 50 credits (5¢) as ten separate files
A 45-minute research interview
45 credits (4.5¢)
A 2.5-hour workshop split into two 75-minute parts
150 credits (15¢) in total
The same workshop with subtitles in a second language
Another 150 credits (15¢) for a second pass with that language as the target

A single run with a target language already gives both texts (Download transcript and Download original transcript). The second pass is only needed for subtitles, because each run produces one subtitle language.

Shaping the transcript for what comes next

The raw transcript is a starting point; how you shape it depends on where the text is going.

  • Interview article: add "Q:" and the interviewee's name at each turn, cut false starts, and keep the timestamped copy so every quote can be checked.
  • Research coding: keep the timestamped version, add participant IDs such as P1 and P2, and leave hesitations in if your method treats them as data.
  • Podcast show notes: scan the timestamps for topic changes and turn them into a chapter list.
  • HR or legal record: keep it verbatim and mark inaudible passages with their timestamp instead of guessing at the words.
  • Audiogram or waveform video for social media: build it in an editor from the audio, then import the SRT from the same run as its captions so the words line up with the speech.

How to format a transcript has templates for each of these. For recurring interview work, the interview transcription page describes a full workflow.

Adding speaker names without replaying everything

Transcripts don't label speakers, so attribution is a manual pass. For a two-person interview, open the timestamped transcript next to the audio, play at 1.5× in your media player, and type initials at each change of voice; questions usually make the interviewer's turns obvious. For group recordings, the work gets easier if someone notes during the session who is talking at a few moments. Even one or two anchor times per person make similar voices easier to tell apart later. If the transcript will be shared beyond the research team, label turns by role (Interviewer, P3) rather than by name from the start, so nothing needs anonymising afterwards.

When the text has gaps or odd output

  • Words lost during laughter or crosstalk: mark the spot as [crosstalk] with its timestamp, and fill it in from the audio only if the content matters.
  • Sentences running together without full stops: see transcript punctuation errors for quick repair patterns.
  • Numbers written inconsistently, sometimes as figures and sometimes as words: standardise them with your style guide during the proofread.
  • Repeated phrases over long silences or noise: speech recognition can produce stray text where nobody is talking, so trim long dead stretches before transcribing.

For iPhone recordings specifically, the M4A to text page covers what's inside those files.

Frequently asked questions

Which audio format is best for transcription?

Any format works; recording quality matters more. WAV and FLAC avoid compression artifacts, but a clear MP3 transcribes just as well.

Is my recording kept?

No. Audio is deleted after delivery, within 30 minutes at most, and never used for training.

Does it work for a recording that mixes two languages?

The spoken language is detected automatically, and a recording that switches languages can be handled inconsistently. If the switch happens in long blocks, split the file at each change and transcribe the parts separately.

Can I get the transcript in another language?

Yes. Pick a target language before starting and the plain transcript and subtitle downloads come back translated. Select the language that was spoken if you need the text untranslated.

How do I join several short recordings into one file?

Any audio editor can do it: place the clips one after another with a short silence between them and export a single MP3, M4A or WAV. Note each clip's start time so you can find it in the transcript.