Why phone audio sounds the way it does
Traditional telephone networks were designed to carry intelligible speech using as little capacity as possible. Classic landline audio is sampled at 8 kHz, which can represent frequencies only up to 4 kHz, and in practice the passband is roughly 300 to 3,400 Hz. That is plenty for a listener to follow a conversation, but it cuts off the low warmth of the voice and, more importantly for transcription, the high-frequency energy of consonants.
The sounds that live mostly above that ceiling include s, f, th and sh. On a narrowband call, "fifty" and "sixty" or "f" and "s" can sound nearly identical, which is exactly why people spell names over the phone with words like "S as in Sam". A recognizer has the same problem with less context than a human listener.
Many mobile and internet calls are better than this. Wideband "HD Voice" on mobile networks and many internet calling apps carry speech up to about 7 kHz or more when both ends and the network support it. But calls often fall back to narrowband when one side is on an older network or a landline, and a single recording can switch quality partway through.
What internet calls add: compression and packet loss
Calls made through internet apps use compressed audio codecs that adapt to the connection. On a good network they can sound better than a landline. On a weak one the app lowers quality, and lost data packets cause brief dropouts, robotic warbling or clipped syllables. The codec behavior behind this is described in the Opus audio codec.
For transcription, the dropouts matter more than the reduced quality. A missing 200-millisecond fragment can remove a whole short word, and the recognizer will often write a fluent sentence around the gap. Calls also produce more crosstalk than in-person conversations, because the delay on the line makes people start talking at the same time and then both stop.
What recognition does with narrowband audio
Speech recognition models such as Whisper work on 16 kHz audio. A phone recording sampled at 8 kHz is converted up to that rate, but conversion can't create the frequencies the phone network discarded; the upper half of the spectrum stays empty. Audio sample rate explained covers why higher rates can't restore missing content.
Most modern recognizers are trained largely on wideband audio, so narrowband speech is somewhat out of their comfort zone. In practice, common words in clear sentences come through well. The errors cluster in places where the missing frequencies mattered: unfamiliar names, spelled-out email addresses, reference numbers, and words distinguished only by an s or f sound. Audio enhancement tools that claim to restore bandwidth produce mixed results and can introduce artifacts, so compare a transcript of the enhanced file against the original before relying on one.
Consent: check before you record
Recording laws differ by country and, in the United States, by state. Some places allow a call to be recorded with the consent of one participant, which can be the person recording; others, including California, require all parties to agree. When participants are in different places, the stricter rule may apply. Workplace rules, data protection laws and professional codes can add further requirements, and businesses that record customer calls usually need to tell callers at the start.
This is general information, not legal advice. If the recording matters, for research, journalism, employment or evidence, check the rules that apply to every participant's location or ask a qualified professional. Telling everyone at the start that the call is recorded, and that the recording will be transcribed by an online service, is the simplest way to avoid most problems. Cloud speech processing and privacy covers what to tell people about where the audio goes.
Getting the recording off the phone or app
Where the call is recorded makes a large difference to quality.
- Built-in phone call recording
- Recent versions of iOS and some Android phone apps offer call recording in some regions, usually announcing it to all participants. Availability depends on device, software version and country, so check current official documentation. The result is usually a clean recording of both sides.
- Call or meeting app recording
- Internet calling and meeting apps often record locally or to the cloud and export M4A, MP4 or WAV. This captures the audio before playback through a speaker, which is ideal.
- Business phone systems
- Call-center and VoIP systems usually export WAV or MP3, sometimes with each party on a separate stereo channel.
- Speakerphone into a second device
- Works in a pinch, but adds room echo and loudspeaker distortion on top of the phone audio. Use it only when nothing else is available.
- Voicemail
- Many phones let you share a voicemail as an audio file; some carriers' voicemail services only allow playback.
Some Android recorder apps save AMR or 3GP files, and some messaging apps save voice notes as .opus files. If a file isn't in a common format such as M4A, MP3, WAV, AAC, OGG or FLAC, convert it to WAV or M4A in an audio editor before transcribing; a lossless conversion to WAV adds no further quality loss.
A worked example: phone interviews for a study
A hypothetical research team conducts twelve phone interviews of about 25 minutes each through a call-recording service, with participants' recorded consent. The service exports stereo WAV files with the interviewer on the left channel and the participant on the right. Each interview costs 25 credits (2.5¢) to transcribe. The team uploads the files as they are, so both sides are combined into one transcript, then adds speaker initials during review. For two interviews where exact attribution is critical, they split the channels in an audio editor and transcribe each side separately, then merge the two timestamped transcripts.
The stereo split is worth knowing about. When a system records each party on its own channel, you have a clean way to attribute speech that an ordinary mixed recording doesn't offer, at the cost of transcribing twice and merging. Mono vs stereo for speech explains what happens when such channels are combined.
Steps from recorded call to checked transcript
- Confirm that recording was lawful and that participants know the recording will be transcribed.
- Record at the source, using the phone's own recorder, the call app or the phone system, rather than a speakerphone.
- Export the original file without re-compressing it, and note who was on the call and when.
- Convert only if the format isn't supported, preferably to WAV.
- Trim hold music, menu prompts and long waits from the start and end; recognizers can produce stray text from them.
- Transcribe the file, with a translated transcript if participants spoke another language.
- Review every number, date, amount, email address and spelled-out name against the audio; these are where narrowband errors concentrate.
- Add speaker names by hand and mark dropouts as [inaudible] with a timestamp rather than filling them in.
Limits of transcribing calls
- Frequencies discarded by the phone network are gone, and no setting or enhancement reliably brings them back.
- Packet-loss dropouts remove words, and fluent-looking text can hide the gap.
- Crosstalk caused by line delay is common and hard for any recognizer.
- Hold music and automated menus can produce stray text in the transcript.
- A transcript of a call is not a certified record; for disputes or legal use, have a qualified transcriber or the relevant professional review the original recording.
Using mydubly for recorded calls
mydubly works on recordings after the call, not during it. It doesn't join, record or translate live calls, and it doesn't import from links, so download or export the recording first. Upload an MP3, WAV, M4A, AAC, OGG or FLAC file to audio to text, or a call recorded as video to video to text. The browser converts the audio to 16 kHz mono, so stereo per-party channels are combined, and sends compressed chunks over HTTPS for transcription with Whisper.
You get a plain transcript, a timestamped transcript and SRT and VTT subtitles. If the call was in a language you don't speak, choose a target language for a translated transcript as well; the audio translator covers that use. There are no speaker labels. Audio and results are deleted within 30 minutes of the job finishing, and the privacy policy says user data isn't used for training, but participants should still be told that an online service will process the recording. Transcription costs 1 credit per minute with a 5-credit minimum, so a 3-minute voicemail costs 5 credits (0.5¢).
Call recordings are dense with names, numbers and dates, so review those first using the approach in how to proofread an AI transcript.
Next step: one call, checked properly
Take one recorded call you already have, trim the hold music, and transcribe it. Then check every number and name against the audio. That review shows you how much your particular phone setup costs in accuracy and whether it's worth changing how you record. For broader tips on recording speech, see how to improve transcription accuracy.
Frequently asked questions
Can I transcribe a call while it's happening?
Not with mydubly, which works only on finished recordings and has no live or real-time mode. Some calling and meeting apps offer their own live captions, and dedicated live-transcription services exist. For accuracy, a recording transcribed afterward and then reviewed usually gives a more reliable record than live captions, which have to commit to each word immediately. See real-time speech translation for how live systems differ.
Does it help to ask the other person to use headphones or a headset?
Yes, when the call is recorded on your side. A headset keeps their voice close to their microphone and prevents your voice from echoing back through their speaker into their microphone. It doesn't widen the phone network's bandwidth, but it reduces room echo and crosstalk, which are among the most damaging problems on recorded calls.
Why are email addresses and reference numbers so often wrong?
They combine the two hardest things for narrowband transcription: letters distinguished only by high-frequency sounds, such as s and f or b and d, and strings with no sentence context to guide the recognizer. Ask callers to read codes using a spelling alphabet, repeat them back during the call, and always verify them against the audio rather than trusting the transcript.
Can I separate the two speakers on a call recording?
Only if they were recorded on separate channels, which some business phone systems and call-recording services do. In that case, split the stereo file into two mono files in an audio editor and transcribe each one, then merge by timestamp. If both voices are mixed into one channel, add speaker names by hand during review; mydubly doesn't label speakers.
Is it worth transcribing a call that was recorded on speakerphone?
Usually, yes, as long as you can follow the conversation yourself. Expect more errors than a direct recording, because the room echo and the loudspeaker's distortion sit on top of the phone audio. Run a short test on the most important passage and decide from the result whether the transcript is usable as a draft or only as an index to the recording.