Recordings it handles well
- Interviews and phone calls you have consent to record.
- Podcast episodes for show notes and accessible transcripts.
- Voice memos from iPhone and Android.
- Research interviews and focus groups.
Speech to text that's built for files
This is speech-to-text for recordings, not live dictation. That trade-off means you can process up to two hours at once, get timestamps on every line, and pay a fraction of a cent per minute.
A 52-minute podcast episode costs 52 credits (5.2¢); a 3-minute voice memo is billed at the 5-credit minimum.
How it works
- Drop in an audio file. It's read in your browser and sent in small chunks.
- The audio is processed in 30-second chunks: speech is transcribed and the spoken language is detected automatically.
- Download the results. Audio and text on our servers are deleted after delivery, within 30 minutes at most.
What you get
- Transcript
- Plain text (.txt) of what was said
- Timestamped transcript
- Every line prefixed with its time, e.g. [12:04]
- Subtitles
- SRT and VTT, in the language you select (select the spoken language for untranslated subtitles)
Supported files
- Audio: MP3, WAV, M4A, AAC, OGG, FLAC, WebM.
- Length: up to 2 hours per file.
- Languages: speech in 21 languages is recognised automatically; translation and voices cover the same 21.
Privacy
Audio is sent over HTTPS in small chunks and deleted after delivery, within 30 minutes at most, together with transcripts and voice files. Nothing is used to train models. Your account keeps only the file name, length, languages and credits used. Details in the privacy policy.
Good to know
- No speaker labels — transcripts don't say who is speaking.
- Files only: no links from YouTube or social apps, and no live audio.
- Subtitles come as SRT/VTT files rather than burned into the picture.
Where recordings come from, and what to send
There's no need to convert a recording before uploading. MP3, WAV, M4A, AAC, OGG, FLAC and WebM audio are all accepted, and converting a compressed file to a larger format doesn't restore anything the compression removed.
- iPhone Voice Memos and many Android recorders
- Usually M4A; send it as is. M4A files
- Field recorders and audio software
- Often WAV: large but fine to upload. WAV files
- Published podcast episodes
- Usually MP3. MP3 files
- Video calls and screen recordings
- Upload the video itself on the video to text page at the same price; there's no need to extract the audio first
MP3 or WAV for transcription? covers what to keep for archives.
A five-minute check that saves a re-run
- Listen to the first minute and one from the middle on headphones. If you struggle to follow the speech, the transcript will struggle too.
- Check for one-sided audio. A microphone recorded to only one channel sounds fine on speakers but odd on headphones; fix it before upload if you're unsure how it will be mixed.
- If each person was recorded on a separate microphone file, mix them into one track first. Each upload is transcribed as a single file.
- Phone-line recordings carry a narrow band of frequencies, so expect more errors on names and numbers and budget more review time.
The five-minute audio check and mixing several microphones go into the fixes.
From transcript to finished interview document
- Run transcript mode and download the timestamped transcript.
- Add speaker labels by hand. Transcripts don't say who is speaking; how diarization works explains why this is hard and the practical workarounds.
- Verify every quote you plan to use by jumping to its timestamp and listening.
- Pick a transcription style before editing. Verbatim or clean verbatim explains when fillers and false starts matter.
- Replace names of participants with codes before the file leaves your project folder, if your consent terms require it.
For podcasts the route is shorter: the plain transcript becomes show notes, the timestamped transcript gives you chapter times, and the SRT or VTT file covers video versions of the episode and web players that display captions.
Costs for typical recordings
75 credits (7.5¢).
Uploaded one by one, each is billed at the 5-credit minimum: 60 credits (6¢). Joined into one 24-minute file first: 24 credits (2.4¢).
120 credits (12¢), the longest single file accepted.
File transcription, dictation or a human transcriber
Three routes turn speech into text, and they suit different work. Dictation software types as you speak, which is ideal for drafting an email and useless for a recording you already have. File transcription, which is what this page does, takes a finished recording and returns the whole text with timestamps. A human transcriber costs far more and takes longer, but handles heavy crosstalk, strong accents on poor audio, and strict verbatim formats such as legal transcripts more reliably.
A common middle path is to run the recording here first, then pay a person only to check and correct the passages that matter. Dictation or transcription? and AI vs human transcription compare the routes in more detail.
Recordings that need extra care
- Crosstalk: when two people speak at once, recognition usually follows one voice and loses the other. Ask people to take turns if you're the one recording.
- Language coverage: speech in 21 languages is recognised. Of Indian languages, only Hindi is covered, so a recording in another Indian language won't transcribe properly.
- Very long sessions over two hours need splitting at a pause; transcribing long audio files covers where to cut.
- Digitised tapes and old archive recordings: hiss and uneven speed make recognition harder. Capture at a healthy level and go easy on noise reduction, which can remove parts of the speech along with the hiss; from cassette to transcript covers the capture settings.
Job-specific workflows are in interview transcription, podcast transcription and voice memo to text. Spoken-language pages such as Spanish and German cover what to check for each language, and the audio transcription guide walks through a first run.
Transcribe speech in…
By file format
Learn more
Articles that go deeper on the technology and workflows behind this tool.
- How speech to text works, from waveform to punctuated sentence
- End-to-end Whisper versus hybrid HMM/DNN speech recognition
- Transcribing long audio files: chunking, boundaries and the 2-hour limit
- How background noise, echo and music affect speech recognition
- Speaker Diarization: How Software Works Out Who Spoke When
- True Verbatim, Clean Verbatim or Edited: Choosing a Transcription Style
- Writing Show Notes, Chapters and Summaries from an Episode Transcript
- How Students Can Use AI Transcription Without Cutting Corners
- Transcribing Oral History Interviews for the Archive
- Getting Usable Transcripts from Focus Group Recordings
- Machine or Person? Comparing AI and Human Transcription
- Recording With Several Microphones: How to Mix for Transcription
- Transcribing Fast Talkers, Commentators and Sped-Up Recordings
- Getting a Usable Transcript From Quiet or Distant Audio
- Transcribing Recorded Phone Calls: Narrowband Audio, Consent and Exports
- From Cassette to Transcript: Digitizing Old Tapes for Transcription
- Clipped Audio: What It Does to Speech and How Far It Can Be Repaired
- How to Extract Audio from a Video, and When You Don't Need To
- Recorder Settings for Speech: What to Set and What to Leave Alone
- How to Record Audio on Your Phone So It Transcribes Well
- Recording a Remote Interview over a Video Call, Cleanly and with a Backup
- A Five-Minute Audio Quality Check Before You Transcribe
- Choosing a Microphone for Speech and Transcription
- From Customer Interview to Written Case Study
- Editing an Interview Transcript into a Q&A or Feature Article
- How to Learn a Language with Podcasts, from Learner Shows to Native Ones
- Dictation Practice: Write Down What You Hear and Learn from Every Mistake
- Extensive Listening: How to Build Hours of Easy Listening into Your Week
- Audio Codecs Explained: How Sound Is Compressed and Played Back
- Audio Sample Rates From 8 kHz to 48 kHz, and Which Ones Matter for Speech
- How Much Bitrate Does Speech Need Before Transcripts Suffer?
- Mono or Stereo for Voice Recordings: What Changes When Channels Are Combined
- Why a Transcript Skips Words, Sentences or Whole Minutes
- Muffled Speech in a Recording: Causes, Checks and Realistic Fixes
- Hum and Buzz in a Recording: Tracking Down Mains Noise and Cleaning It Up
- MP3 or WAV for Transcription? It Depends on What You Do Next
- Audio Bit Depth: What 16, 24 and 32-Bit Float Really Change
- Compressing Voice Audio: Settings, Limiters and When to Leave It Alone
- Treating a Room for Voice: Reflections, Echo and Where Absorption Goes
- How Acoustic Echo Cancellation Works, and What It Does to Recorded Calls
- Voice Activity Detection: How Software Decides When Someone Is Speaking
- What Makes Speech Intelligible, and Why Loud Is Not the Same as Clear
- How to EQ a Voice Recording: Frequencies, Cuts and Sensible Boosts
- Plosive Pops in Voice Recordings: Prevention and Repair
- Harsh S Sounds in Voice Recordings: Causes, De-Essers and Technique
- dBFS and Audio Levels Explained: Meters, Headroom and Recording Targets
- Speaking for the Transcript: Delivery Habits That Make Recordings Easy to Transcribe
- How to Organize a Growing Collection of Audio Files
- Searching AI Transcripts When the Words May Be Wrong
- How to Review a Long Recording Without Listening to All of It
- How to Summarize a Transcript Accurately
- From Interview Recordings to Themes: Preparing Transcripts for Analysis
- Choosing and Writing Transcription Conventions for Research Data
- Recording Interviews in the Field Without Losing Data
- Returning Transcripts to Participants: A Practical Guide to Member Checking
- How to Build a Corpus of Spoken Language
- Doing Content Analysis on Interview, Media and Meeting Transcripts
- How Long It Really Takes to Transcribe an Interview
- Why AI Transcripts Get Punctuation Wrong, and How to Clean It Up
- Twenty-Five or 25? Making Numbers Consistent in AI Transcripts
- Gaps, Clicks and Stutters: Tracking Down Audio Dropouts
- Voice Only in the Left Ear? Fixing One-Sided Audio
- Chipmunk or Slow-Motion Voice? Fixing Audio at the Wrong Speed
- AAC or MP3? How Two Lossy Audio Codecs Differ in Practice
- Dictation or Transcription? Speaking Text Live vs Converting a Recording
- Speech Recognition vs Voice Recognition: What Was Said, Who Said It, and When
Frequently asked questions
Is this real-time speech to text?
No. It transcribes recorded files. For dictation as you speak, use your device's built-in dictation.
Can it tell who is speaking?
No. Transcripts don't include speaker labels. For interviews, add names by hand as you review; the timestamps make it quick to find where the voice changes.
Can I transcribe a recorded phone call?
Yes, if you have the recording as a file and the consent your jurisdiction requires. Phone audio is narrowband, so review names and numbers carefully. Transcribing phone calls covers exporting recordings from common apps.
Is a WAV file transcribed more accurately than an MP3?
Only if the WAV was recorded that way. Converting an MP3 to WAV makes the file bigger without adding detail, so upload whichever file you already have.