About M4A files
M4A is AAC audio in an MP4 container — the format of iPhone Voice Memos and many Apple and Android recorders.
Common sources: iPhone Voice Memos, Android voice recorders, GarageBand exports, audio downloaded from your own videos.
All modern browsers read M4A.
M4A to text
Share a Voice Memo to Files (or AirDrop it to your Mac) to get the M4A, then drop it in.
A founder records a 24-minute idea session in Voice Memos and turns the M4A into text to clean up into a plan.
Cost: 24 credits (2.4¢).
Translating a M4A
Short memos are billed with a per-file minimum, so combine several short memos into one file when you can.
A sales rep translates a 9-minute M4A voice briefing into Japanese for a colleague in Osaka.
Cost: 450 credits (45¢) with a translated voice, or 9 credits (0.9¢) for translated text only. Switch Full translation output on in the tool above for a voice track.
One extension, two kinds of audio
An .m4a file is an MP4 container holding sound only, and that sound is nearly always AAC. The exception matters for iPhone users. Voice Memos has an Audio Quality setting (under Settings → Apps → Voice Memos on recent iOS), and Lossless stores Apple Lossless (ALAC) under the same .m4a extension. ALAC is supported less widely than AAC, so if a lossless memo won't load in your browser, try Safari or convert a copy to AAC. Compressed, the default, is plenty for speech. The same container also sits behind .m4b audiobooks and .m4r ringtones, but only .m4a is on the accepted list. AAC vs MP3 explains why AAC holds up at the low bitrates memos use.
Getting recordings out of the app
- Mac with iCloud sync: drag a memo from the Voice Memos app to the desktop and you have the M4A, no phone needed.
- Before exporting, try Voice Memos' Enhance Recording (the wand in the edit view) on a noisy memo; it reduces background noise and room echo. Compare before and after, and leave it off on clean recordings.
- Android: Google's Recorder and Samsung Voice Recorder both share files, but Recorder offers several share options. Pick the audio file, not a link or the app's own text.
- Zoom: local recordings save an audio-only M4A next to the MP4, smaller to move around than the video and all you need when only the words matter.
Zoom's per-participant files
Zoom's local recording settings include an option to save a separate audio file for each participant, so you get one M4A per person, each carrying mostly one voice. The setting applies to local recordings and has to be switched on before the meeting starts; it can't be produced afterwards from an ordinary mixed recording, so turn it on for interviews and panels you know you'll want attributed. That helps because transcripts don't label speakers: transcribe each file on its own and you know who said every line, then merge the timestamped transcripts by time if you want one document. Two caveats. Each file holds bleed from others and long silences while they talk, and recognition models occasionally drop stray words into long silences, so skim those stretches. And cost scales with the number of files, not the meeting length, as the table shows. Why telling speakers apart automatically is hard covers other workarounds, and remote interview setup covers recording each side locally.
Short memos and per-file minimums
- One 50-minute meeting as a single mixed M4A
- 50 credits (5¢)
- The same meeting as four per-participant files
- 200 credits (20¢)
- A 90-second idea memo transcribed on its own
- 5 credits (0.5¢), the per-file minimum
- Twelve 90-second memos joined into one 18-minute file
- 18 credits (1.8¢) instead of 60 credits (6¢)
- A 1-minute memo dubbed into Japanese
- 100 credits (10¢), because a dub bills at least 2 minutes
Joining memos takes a minute in GarageBand or Audacity: place them on one track in order, leave a second or two of silence between them so the joins are easy to find in the timestamped transcript, and export as M4A or WAV.
Which output to reach for
Transcripts and subtitles always come back as text files. In dub mode the translated voice is delivered as M4A too, so an M4A memo goes in and an M4A comes out. For someone dictating ideas, the plain transcript is usually the useful file; for a recorded talk you plan to post as an audiogram, it's the SRT. A memo meant for a colleague who reads another language usually needs only a translated transcript, which costs the same per minute as the original-language one; a translated voice track is worth the dub price when the listener will hear it on the move rather than read it. Voice memo to text covers the dictation workflow end to end, and recording on a phone covers getting a cleaner memo next time. If your recording is actually a video, the MP4 page applies.
Using the tool on this page
- The spoken language is detected automatically; you select the output language — the spoken language itself for an untranslated transcript, or another language for a translation. Files can be up to 2 hours long.
- Audio is sent over HTTPS in small chunks and deleted with its transcripts within 30 minutes of delivery (how files are protected).
- You can download a timestamped transcript, subtitles (SRT and VTT) and plain text. Every output, step and limit is explained on the audio to text page.
Frequently asked questions
How do I get a Voice Memo off my iPhone?
In Voice Memos, tap the recording, then Share → Save to Files or AirDrop. The file is an M4A you can drop straight in.
Is the translated audio also M4A?
Yes, the translated voice is delivered as M4A.
My Voice Memo is three hours long. How do I fit the 2-hour limit?
Duplicate the memo in Voice Memos, then trim one copy to the first half and the other to the second, cutting at a pause. Upload them as two files; the second file's times start from zero, so note where it began.
Should I convert M4A to MP3 before uploading?
No. M4A is read directly, and converting adds a second round of lossy compression for no benefit.
My memos have long pauses while I think. Does that matter?
Pauses are billed like any other minute, because cost follows length. Voice Memos' Skip Silence only changes playback, not the file, so trim long gaps in the editor if you want a shorter file. Long silences are also where stray words occasionally appear, so glance at the lines around them.