Audio & video formats

Audio Codecs Explained: How Sound Is Compressed and Played Back

An audio codec is the method used to turn sound samples into a compact stream of bits and back again; the name is short for coder-decoder. Lossless codecs such as FLAC rebuild every original sample exactly, while lossy codecs such as AAC, MP3 and Opus discard detail that listeners are unlikely to notice in exchange for much smaller files. Most video from phones and editors carries AAC audio, and the file extension you see names the container, which is not always the same thing as the codec inside.

9 min read · Updated

What an audio codec actually does

A microphone signal becomes digital audio as a long series of numbers called samples. A codec is the agreed recipe for storing those numbers efficiently. The encoder side compresses them into a bitstream; the decoder side expands that bitstream into samples a sound card can play. Encoder and decoder must follow the same specification, which is why a file can refuse to play on a device that simply lacks the right decoder.

A codec is not a file type. The file type is the container: a wrapper that stores one or more encoded streams along with timing and metadata. An .m4a file is an MP4-family container that usually holds AAC, a .wav file is a container that usually holds uncompressed samples, and an .ogg file can hold Vorbis, Opus or FLAC. How containers carry streams, and how a stream can be moved between them without re-encoding, is covered in browser video muxing.

That distinction explains a common surprise. Renaming a file from .ogg to .mp3 changes nothing inside it; the bytes are still Vorbis or Opus, and a player that trusts the extension may refuse to open it.

Uncompressed PCM: the reference point

Pulse-code modulation, or PCM, is audio with no compression at all: every sample stored as a plain number. It is the format editors work in internally and the format every other codec is ultimately decoded back into. Its data rate follows directly from three settings:

Sample rate
How many samples per second, for example 44,100 or 48,000
Bit depth
How many bits per sample, commonly 16 or 24
Channels
One for mono, two for stereo
Data rate
Sample rate × bit depth × channels, in bits per second
CD audio
44.1 kHz, 16-bit, stereo: 1,411.2 kb/s
Typical video sound
48 kHz, 16-bit, stereo: 1,536 kb/s

WAV and AIFF files almost always contain PCM, and professional cameras and recorders often write PCM into MOV or MXF files. It is simple, fast to edit and free of artefacts, but large. What the sample rate itself controls is explained in audio sample rate explained.

Lossless compression: smaller files, identical samples

Lossless codecs look for redundancy, mostly by predicting each sample from the previous ones and storing only the small prediction error. When the file is decoded, the output matches the original PCM bit for bit. FLAC is the open, widely supported example; Apple Lossless (ALAC), usually stored in .m4a files, does the same job in Apple's ecosystem.

How much smaller a lossless file gets depends on the material. Quiet, sparse recordings such as a single voice in a calm room compress more than dense, loud music. Because nothing is thrown away, lossless files can be converted to any other format later without stacking up damage, which makes them the natural choice for archives and masters.

Lossy compression: discarding what listeners are unlikely to hear

Lossy codecs go further by removing information. They transform short blocks of audio into frequency components and use a psychoacoustic model, a set of rules about human hearing, to decide which components can be dropped or stored coarsely. Loud sounds mask quieter sounds at nearby frequencies, very high frequencies matter less to intelligibility, and the codec spends its limited bits where the ear is most sensitive.

The trade-off is set mainly by the bitrate, the number of kilobits per second the encoder is allowed to use. At generous bitrates, good lossy codecs are very hard to tell apart from the original. At low bitrates, artefacts appear: a swirly or metallic texture, smeared consonants and a dull top end. The practical effect of bitrate on speech, and on transcripts, is the subject of audio bitrate for speech recognition.

The other cost is generation loss. Each time lossy audio is decoded and encoded again with a lossy codec, a new round of approximation is applied on top of the last one. One extra encode at a sensible bitrate is usually harmless; a chain of them is not.

The codecs you will meet most often

PCM
Uncompressed. Found in WAV, AIFF and many professional camera files. Lossless and large
FLAC
Open lossless codec. Common for archives, music libraries and high-quality recorder exports
ALAC
Apple Lossless, usually in .m4a files. Lossless, mostly used within Apple software
MP3
MPEG-1 Audio Layer III, the 1990s format that made digital music portable. Plays almost everywhere; its key patents have expired
AAC
Advanced Audio Coding, MP3's successor within the MPEG standards. Generally more efficient than MP3 at the same bitrate and the default soundtrack codec in MP4 video
Vorbis
An open codec from the Xiph.Org Foundation, used in Ogg and WebM files. Largely replaced by Opus for new work
Opus
An open IETF standard designed for both speech and music, especially efficient at low bitrates. Common in WebM, Ogg and voice apps
AC-3 and E-AC-3
Dolby Digital formats used for surround sound in broadcast, discs and some streaming

Opus deserves a special mention for anyone working with speech. It was designed to handle both voice calls and music in a single codec, it stays intelligible at bitrates where older codecs fall apart, and modern browsers can encode it natively. The full story, including its speech and music modes and its containers, is in our article on the Opus audio codec.

Which audio codecs live inside video files

The audio codec in a video depends far more on the device or app that made it than on the file extension. As a rough guide to what is commonly found:

  • Phone cameras, iPhone and Android alike, usually record AAC in MP4 or MOV.
  • Video editors and meeting apps usually export MP4 with AAC unless you choose otherwise.
  • Browser-based screen recorders often produce WebM with Opus audio.
  • Game-capture and streaming tools commonly record AAC, sometimes into MKV, and may add extra audio tracks.
  • Professional cameras and field recorders frequently write uncompressed PCM, often with several channels or tracks.
  • Broadcast masters and disc rips may carry AC-3, E-AC-3 or other surround formats.

To find out for certain, open the file in a media inspector such as MediaInfo, or look at the codec information panel in a player like VLC. They list each stream with its codec, sample rate, channel count and bitrate.

Example: one 45-minute interview, four ways

Stored as 48 kHz, 16-bit stereo PCM, the interview takes about 518 MB. As FLAC it shrinks noticeably with no change in sound. As stereo AAC at 128 kb/s it is about 43 MB, and as mono Opus at 32 kb/s about 11 MB. All four versions transcribe well if the original recording was clean, but only the PCM and FLAC copies can be re-edited and re-encoded later without adding new loss.

Choosing a codec for speech recordings

  1. Keep the original file exactly as the recorder or camera wrote it. It is the highest-quality copy you will ever have.
  2. For editing and archiving, use PCM in WAV or FLAC. Storage is cheap compared with re-recording an interview.
  3. For sharing speech, AAC or Opus at a moderate bitrate keeps files small and voices clear. MP3 is fine where maximum compatibility matters.
  4. Avoid converting from one lossy codec to another unless a platform demands it, and never do it repeatedly.
  5. Before converting anything, check what you already have with a media inspector, because the right answer is often to leave the file alone.

Common codec mistakes and their limits

  • Converting MP3 to WAV or FLAC to improve quality. The file gets bigger but the lost detail does not come back; you have only stored the existing damage losslessly.
  • Treating the extension as the codec. An .mp4 can hold several audio codecs and an .ogg can hold three; check the stream itself.
  • Compressing speech at very low bitrates to save space, then finding names and numbers hard to make out later.
  • Exporting from one lossy format to another on every step of an editing chain, so artefacts accumulate.
  • Assuming a lossless file is automatically a good recording. FLAC preserves a noisy, distant or clipped recording perfectly, faults included.

What mydubly accepts and what it returns

For practical purposes, the codec of your file matters less than whether your browser can open it. mydubly accepts video as MP4, MOV, WebM, MKV and M4V, and audio as MP3, WAV, M4A, AAC, OGG and FLAC, up to 2 hours per file. The page first checks that the browser can read the file. It then decodes the soundtrack locally to 16 kHz mono and compresses about 30 seconds at a time to Opus at around 32 kb/s for upload; the video itself never leaves your device. Whatever codec you started with, the speech recognizer receives the same kind of input.

For a transcript you get plain text, a timestamped transcript and SRT and VTT subtitles. For a dubbed version, the translated audio, the new voice mixed over the original music and effects, arrives as an M4A file, which is mono AAC audio in an MP4-family container, and the translated video keeps your original picture with that same audio in place of the original soundtrack. If your file has several audio tracks, only one is used, so export with the main speech as the only or main track.

How speech to text works shows what happens to the decoded samples after the codec has done its job.

Next step: check one of your own files

Open a recording you use often in a media inspector and note its audio codec, sample rate, channels and bitrate. If it is clean speech in any of the formats above, you can transcribe it as it is with mydubly's audio to text tool; the M4A to text page covers the most common phone recording format, and a translated voice track is available through AI dubbing. A 10-minute test transcript costs 10 credits (1¢).

Frequently asked questions

Is MP3 a codec or a file format?

Both, which is why it causes confusion. MP3 names the compression method, MPEG-1 Audio Layer III, and the .mp3 file that holds a stream encoded with it. AAC is similar: a raw .aac file contains just the stream, while an .m4a or .mp4 file wraps the same AAC audio in a container that can also hold metadata, chapters or video.

Is AAC better than MP3?

At the same bitrate, AAC generally sounds better than MP3, especially at lower bitrates, because it is a newer design with more efficient tools. At high bitrates both can be very hard to tell apart from the original. MP3 still has the edge in compatibility with very old players and car stereos, so the choice depends on where the file will be played.

Does converting a lossy file to FLAC or WAV improve quality?

No. Converting decodes the lossy audio and stores the result without further loss, so it stops additional damage but cannot restore what the original encoder removed. The new file is larger and sounds the same. It can still be useful as an editing format, because further exports from it will not stack another lossy generation on the source.

Which audio codec should I export for transcription?

If you control the export, a lossless format such as WAV or FLAC is the safest choice, and AAC or MP3 at a moderate bitrate also works well for clear speech. The recording conditions matter much more than the codec: a close microphone in a quiet room beats any export setting. See the improve transcription accuracy guide for what to fix at the source.

Why does a file play on my computer but not on my phone?

Usually because the phone lacks a decoder for the codec inside, even when it recognizes the container. AC-3 surround audio, some Opus and Vorbis files on older Apple devices, and uncommon PCM variants are typical culprits. A media inspector tells you which codec is involved, and re-encoding just the audio to AAC usually solves it while leaving the video stream untouched.