Audio & video formats

How Video Codecs Shrink Moving Pictures, and Why It Matters for Your Files

A video codec is the method that compresses a sequence of pictures into a manageable stream and decodes it again for playback. It works by squeezing each frame on its own and, far more powerfully, by storing only what changes from one frame to the next. H.264 remains the most compatible codec, HEVC and AV1 reach similar quality at lower bitrates but need newer decoders, and ProRes trades size for easy editing.

9 min read · Updated

The problem a video codec solves

Uncompressed video is enormous. A single 1080p frame has about 2.07 million pixels, and with the 8-bit, 4:2:0 color layout used by most consumer video, each frame takes roughly 3.1 MB. At 30 frames per second that is about 93 MB every second, or roughly 750 megabits per second. A one-hour recording would fill hundreds of gigabytes.

Consumer 1080p video usually travels at a few megabits per second instead, so a codec typically shrinks the picture by a factor in the hundreds. It can only do that by being lossy: the decoded picture is a close approximation of what the camera captured, not an exact copy. How close depends on the codec, the encoder's settings and the bitrate it was allowed.

A codec is also separate from the file that holds it. MP4, MOV, MKV and WebM are containers that carry a video stream next to audio and metadata; the article on browser video muxing explains how streams move between containers. This article is about what is inside the video stream itself.

Compressing a single frame

The first tool is intra-frame compression, which treats each frame much like a still photo. The encoder divides the picture into blocks, converts each block from pixel values into frequency components with a mathematical transform, and then quantizes those components, rounding fine detail coarsely while keeping broad shapes and edges. Entropy coding then packs the result into as few bits as possible.

Before any of that, most consumer video discards some color information through chroma subsampling. Human vision is much more sensitive to brightness than to color detail, so 4:2:0 video stores color at a quarter of the brightness resolution. Professional formats often keep 4:2:2, which holds more color detail for grading and green-screen work.

Compressing what changes between frames

The second tool, inter-frame compression, is where most of the savings come from. Consecutive frames are usually very similar: a presenter moves their hands, but the wall behind them stays put. Instead of storing each frame in full, the encoder searches the previous or following frames for matching blocks, records a motion vector saying where each block moved, and stores only the small residual difference.

That produces three kinds of frames:

I-frame
Stored in full using only intra-frame compression. Decodable on its own
P-frame
Predicted from earlier frames. Much smaller, but depends on what came before
B-frame
Predicted from both earlier and later frames. Smallest, with the most dependencies
GOP
Group of pictures: the run of frames from one I-frame to the next
Keyframe
An I-frame where decoding can start cleanly, also called a sync point or IDR frame

Keyframes have practical consequences you meet every day. A player can only start decoding at a keyframe, so seeking jumps to the nearest one and decodes forward. A tool that cuts video without re-encoding can only cut cleanly at a keyframe, which is why lossless cuts snap to slightly different points than you chose; splitting a video file covers working with that. Screen recordings and phone videos can have keyframes several seconds apart, which makes them compact but less precise to cut.

The codecs you will meet

H.264 (AVC)
Standardized in 2003. Decoded in hardware by almost every phone, computer and TV made in the last decade or more. The default choice for compatibility
HEVC (H.265)
Standardized in 2013. Designed to reach similar quality to H.264 at substantially lower bitrates. Recorded by iPhones and many cameras, but playback support depends on the device and browser
VP9
Google's royalty-free codec, used heavily on the web and in WebM files. Widely supported in browsers
AV1
Royalty-free codec from the Alliance for Open Media, finalized in 2018. Efficient, but slow to encode and hardware-decoded only by relatively recent chips
ProRes
Apple's family of editing codecs. Every frame is stored independently, so files are very large but quick to scrub and edit
Older codecs
MPEG-2 on DVDs and broadcast, VP8 in older WebM files, and various camera-specific formats

Efficiency comparisons between codecs are always approximate. They depend on the content, the resolution and above all on the particular encoder and settings, so treat any single number you read with caution.

Profiles, levels and bit depth

Saying a file is H.264 or HEVC does not fully describe it. Each codec defines profiles, which are sets of features an encoder may use, and levels, which cap resolution, frame rate and bitrate so a decoder knows what it must handle.

  • H.264 has Baseline, Main and High profiles for everyday video, plus High 10 and High 4:2:2 profiles used by some cameras. Hardware decoders in consumer devices generally handle the everyday profiles but often not the 10-bit or 4:2:2 ones.
  • HEVC Main is 8-bit and Main 10 is 10-bit. HDR video from phones uses 10-bit, which older decoders may not support.
  • A level that is too high for a device, such as 4K at a high frame rate on an older phone, can stop playback even though the codec is technically supported.

This is why two files that both say H.264 can behave differently: one plays everywhere, the other only on a desktop media player.

Why one device plays a file and another doesn't

Video decoding is demanding, so devices lean on dedicated hardware decoders, and browsers usually rely on whatever the operating system and graphics chip provide. A codec with patent licensing, such as HEVC, may only be available where the device maker has licensed it or where hardware support exists. Newer codecs such as AV1 play smoothly only on chips that decode them in hardware; elsewhere, software decoding may stutter or be unavailable. Support changes with every operating system and browser release, so check current documentation when a file matters. The specific case of phone footage is covered in HEVC compatibility.

The cost of every re-encode

Changing the video codec, the resolution or the bitrate means decoding every frame and encoding it again. That takes time, keeps the processor busy and, because the codec is lossy, applies a fresh round of approximation on top of the original. One careful re-encode at a generous setting is usually invisible; repeated ones show up as softness, blockiness and banding.

Re-encoding is necessary when you need a different codec for compatibility, a smaller file, burned-in text or any change to the picture. It is not necessary to change the container, replace or remove an audio track, or cut at keyframes. Those jobs only copy compressed packets, as the guide to converting video formats explains in more detail.

How mydubly treats your video stream

mydubly only needs the sound to translate a video, so it leaves the picture alone. Your browser first opens the file to read its duration, then decodes just the audio locally and uploads compressed audio chunks; the video file itself never leaves your device. When you choose a translated voice, the translated audio is produced and the browser swaps it in by remuxing: the compressed video packets are copied unchanged next to the new audio, so the picture is never re-encoded.

The output is MP4 when the video codec can go into MP4, which covers the codecs above in normal use, and MKV when it can't. Because the video stream is a copy, the translated file keeps the original codec, profile and resolution, and its playback compatibility is the same as the original's. Only the audio changes: the original speech is removed with AI vocal separation, and the original music and effects are kept under the translated voice.

Example: a 20-minute 4K iPhone clip

A presenter films a 20-minute product walkthrough in 4K HEVC and translates it into German with a voice track. The browser reads the file, sends only compressed audio, and later copies every HEVC frame into a new MP4 next to the German audio. No frame is decoded or re-encoded, so the picture is identical to the original and the file is about the same size. The voice costs 1,000 credits ($1.00). If the audience uses older devices without HEVC support, the translated file needs the same conversion the original would have needed.

Reading the codec details of your own file

  1. Open the file in MediaInfo, a free inspector for Windows, macOS and Linux, and read the Video section for format, profile, bit depth, frame rate and bitrate.
  2. In VLC, open the file and choose Tools, then Codec Information, or Window, then Media Information on a Mac.
  3. On a Mac, QuickTime Player's Movie Inspector (Window, then Show Movie Inspector) lists the codec and resolution.
  4. From a terminal, ffprobe prints every stream, for example: ffprobe -v error -show_entries stream=codec_name,profile,pix_fmt,width,height -of compact input.mp4
  5. Note the profile and bit depth as well as the codec name, since those explain most playback failures.

Limits of what a codec name tells you

  • Quality is not set by the codec alone. A well-encoded H.264 file at a healthy bitrate can look better than a starved AV1 file.
  • Bitrate comparisons only hold within one codec and similar content. A talking head and a confetti-filled concert need very different bitrates for the same quality.
  • The extension says nothing certain about the codec. An .mp4 can contain H.264, HEVC, AV1 or others.
  • Frame timing is a separate property. Phones and screen recorders often produce variable frame rate, which causes its own editing problems, explained in variable frame rate explained.
  • Hardware support moves quickly, so advice about which devices play which codec ages fast.

Next step

Check the codec, profile and bit depth of the files you publish most often, and keep an untouched original of anything important. When you want a version in another language without re-encoding the picture, the video translator handles the audio and leaves your frames exactly as they are; the MOV to text page covers iPhone footage specifically, and the video translation glossary defines the related terms.

Frequently asked questions

Is MP4 a video codec?

No. MP4 is a container, a file format that holds a video stream, one or more audio streams and metadata. The video inside an MP4 is usually H.264, but it can also be HEVC, AV1 or another codec. That is why two MP4 files can behave differently on the same device: the container is the same, but the codec, profile or bit depth inside differs.

Does a higher bitrate always mean better quality?

Within the same codec, encoder and content, more bitrate generally means fewer artefacts, but the gains shrink as the bitrate rises. Across codecs the comparison breaks down, because newer codecs need fewer bits for similar quality. Above a certain point, extra bitrate mainly makes the file larger. Judge by watching difficult scenes, such as fast motion or fine detail, rather than by the number.

Why are ProRes files so large?

ProRes stores every frame independently instead of predicting frames from their neighbors, and it uses mild compression to keep detail for grading. That makes it easy for editing software to jump to any frame and decode it quickly, at the cost of files many times larger than H.264 or HEVC delivery copies. It is designed for editing and mastering, not for uploading or sharing.

Should I export my videos in H.264 or HEVC?

Choose H.264 when the file must play on as many devices and apps as possible, or when a platform's upload guidance recommends it. Choose HEVC when smaller files at similar quality matter and you know the audience's devices support it, for example within Apple's ecosystem. Many platforms re-encode uploads anyway, so check their current upload recommendations rather than assuming one is always preferred.

Can I change a video's codec without losing quality?

Not with lossy codecs. Changing the codec means decoding and encoding every frame, which applies new compression. A careful encode at a generous bitrate keeps the loss very small, and encoding to a lossless or near-lossless editing format avoids visible loss at the price of huge files. Changing only the container, by contrast, copies the stream and loses nothing.