Preparing your files

How to Extract Audio from a Video, and When You Don't Need To

To extract audio from a video, first check which audio codec the file contains, then either copy that stream into a matching audio container (lossless and instant) or re-encode it to a format you choose. QuickTime, VLC and Audacity do this through menus; ffmpeg does it in one command. If your goal is a transcript or translation, you may not need to extract anything, because many tools, mydubly included, accept the video directly.

9 min read · Updated

Decide whether extraction is the right step

Extraction is worth doing when you actually need an audio file: to edit the sound in an audio editor, to archive an interview without gigabytes of picture, to send a recording to someone who only needs to listen, or to feed a tool that only takes audio. It is also the fix when a video container isn't accepted by the tool you want to use, or when the speech sits on a secondary audio track.

It is not worth doing just to get a transcript. Most transcription tools read the audio out of the video themselves, and doing it by hand adds a step where you can pick the wrong track, re-encode twice or trim off the start. If the audio to text or video-to-text tool you plan to use accepts your video format, upload the video and skip the rest of this article.

Look inside the file before you touch it

A video file is a container holding separate streams: usually one video stream and one or more audio streams. The audio stream has its own codec, such as AAC, Opus, MP3, AC-3 or uncompressed PCM. Knowing which one you have decides everything else, because a lossless copy only works if the destination container can hold that codec.

The quickest way to see the streams is ffprobe, which ships with ffmpeg. Run ffprobe -hide_banner input.mp4 and look for lines such as "Stream #0:1(eng): Audio: aac (LC), 48000 Hz, stereo". That tells you the codec (AAC), the sample rate, the channel layout and, if the file is tagged, the language. A file with several audio lines has several tracks, and the order matters when you choose one.

Without a terminal, VLC shows the same information under Tools, then Codec Information, and MediaInfo is a free graphical app that lists every stream with its codec and duration. The audio codec explainer covers what the codec names mean.

Stream copy versus re-encoding

There are two fundamentally different ways to get audio out of a video.

A stream copy takes the compressed audio data exactly as it is and writes it into a new audio-only file. Nothing is decoded or compressed again, so there is no quality loss and it finishes in seconds even for a two-hour file. The catch is that the new file's extension has to suit the codec.

Re-encoding decodes the audio and compresses it again with a codec you choose. It takes longer and, if the target is lossy (MP3, AAC, Opus), it discards a little more information on top of whatever the original compression already removed. In return you get exactly the format, channel count and bitrate you want.

AAC in MP4, MOV or M4V
copy to .m4a
Opus in WebM or MKV
copy to .ogg or .opus
Vorbis in WebM
copy to .ogg
MP3 in any container
copy to .mp3
FLAC in MKV
copy to .flac
PCM in MOV
copy to .wav
AC-3, E-AC-3 or DTS
usually re-encode for wide compatibility

For speech, a single lossy re-encode at a sensible bitrate is not audible in a transcript. Repeated re-encoding is where problems accumulate, so copy when you can and re-encode once when you must.

Extracting with QuickTime, VLC or Audacity

On a Mac, QuickTime Player can export the sound from most MOV and MP4 files: open the video, choose File, then Export As, then Audio Only. You get an M4A file containing AAC audio. It is quick, but you can't choose the track, the bitrate or a different format.

VLC is free on Windows, macOS and Linux. Use Media, then Convert/Save (on macOS, File, then Convert/Stream), add the video, and pick an audio profile such as "Audio - MP3" or "Audio - FLAC". VLC always re-encodes in this dialog, so pick FLAC if you want to avoid a second lossy step and plan to edit further.

Audacity opens the audio from video files once its optional FFmpeg library is installed. Import the video, and the audio appears as a waveform you can trim, clean and export as WAV, MP3, FLAC or other formats. Audacity is the right choice when you want to edit the sound anyway, because extracting and editing happen in one place. Exact menu names in all three apps change between versions, so check the current documentation if a menu isn't where you expect.

Extracting with ffmpeg

ffmpeg gives you full control and is the only common tool here that makes true stream copies easy. These steps cover the cases people hit most often.

  1. Inspect the file with ffprobe -hide_banner input.mp4 and note the audio codec and how many audio streams there are.
  2. For AAC audio in an MP4, MOV or M4V, copy it without re-encoding: ffmpeg -i input.mp4 -vn -c:a copy audio.m4a (the -vn option drops the video, and -c:a copy copies the audio as is).
  3. For a WebM screen recording with Opus audio, copy into Ogg instead: ffmpeg -i input.webm -vn -c:a copy audio.ogg
  4. To choose a specific track, add a map option. Audio streams are counted from zero, so the second audio track is 0:a:1, as in ffmpeg -i input.mkv -map 0:a:1 -c:a aac -b:a 128k track2.m4a
  5. To produce a small speech-friendly file in one re-encode, downmix to mono and use a moderate bitrate: ffmpeg -i input.mov -vn -ac 1 -c:a libmp3lame -b:a 96k audio.mp3
  6. To get an uncompressed file for editing, write 16-bit WAV: ffmpeg -i input.mov -vn -c:a pcm_s16le audio.wav
  7. To keep only part of the recording, add a start and end time before the output name, for example -ss 00:02:10 -to 00:47:30. Audio frames are short, so cutting a copied audio stream lands very close to the requested time.
  8. Play the result from start to finish at a few points and compare the duration with the original video.

If ffmpeg reports that a codec isn't supported in the output container, the extension doesn't match the codec. Either choose a matching extension from the table above or switch from -c:a copy to an explicit encoder.

Worked example: a recording with three audio tracks

Conference panel recorded on a camera

A 70-minute MKV from a panel discussion has three audio tracks: track one is the room microphone, track two is a mix of the panel lavaliers, and track three is the interpreter booth. The organizer wants a clean English transcript of the panel and a copy of the interpretation for the archive.

Running ffprobe shows three Opus streams labeled 0:a:0, 0:a:1 and 0:a:2. Listening briefly to each (VLC lets you switch tracks under Audio, then Audio Track) confirms that the lavalier mix is the clearest for speech. The organizer copies it out with ffmpeg -i panel.mkv -map 0:a:1 -c:a copy panel-lav.ogg, which takes a few seconds, and repeats with 0:a:2 for the interpreter track.

The lavalier file goes to transcription. The room-mic track is discarded: it has more echo and audience noise, and a transcription tool that only reads one track would probably have picked the first one, which is the worst of the three. That is the real reason extraction helped here, not file size.

Choosing a format for the extracted file

Pick the format by what happens next, not by habit.

  • Editing or restoration: WAV or FLAC, so later edits and exports start from the cleanest possible copy.
  • Transcription or translation: a stream copy in whatever codec the video already used, or a single re-encode to MP3, M4A or Ogg. Speech recognizers typically downsample audio internally, so an enormous WAV adds upload time without adding accuracy. The trade-offs are covered in MP3 vs WAV for transcription.
  • Archiving: keep the stream copy, or FLAC if you had to decode anyway.
  • Sharing for listening: M4A or MP3 plays almost everywhere.

Mistakes that cost quality or sync

  • Converting MP3 or AAC to WAV and expecting better sound. Decoding a lossy file into an uncompressed container makes the file bigger, not better.
  • Re-encoding several times while passing a file between people. Agree on one master copy.
  • Extracting the wrong track. A music-only, commentary or interpreter track transcribes as nonsense or in the wrong language.
  • Trimming the start of the audio when you plan to put it back under the video later. Keep the full length so it lines up; trim only copies meant for listening or transcription.
  • Downmixing 5.1 film audio without listening first. Dialogue usually sits in the center channel, and a careless downmix can leave voices buried under music and effects.
  • Ignoring a one-sided recording. Some cameras record a lavalier on the left channel only; a mono mix or a single-channel extract fixes it, as explained in mono vs stereo for speech.

Where mydubly fits, and where extraction is still useful

mydubly accepts MP4, MOV, WebM, MKV and M4V video, and MP3, WAV, M4A, AAC, OGG and FLAC audio, up to 2 hours per file. When you upload a video, the browser decodes the audio locally to 16 kHz mono and sends only compressed audio chunks over HTTPS, so the video file itself never leaves your device. Extracting audio yourself first therefore doesn't make the upload more private and barely makes it smaller; the upload size breakdown shows why.

Extraction is still useful in a few cases. If your video is in a container outside that list, such as AVI or WMV, extract or convert the audio first. If the file has several audio tracks, mydubly uses only one, so put the speech on its own file or make it the main track. And if you want to clean or trim the audio before transcribing, extract it into an editor first.

Keep in mind what you lose by uploading audio instead of video. From an audio file you get a transcript, a timestamped transcript, SRT and VTT subtitles, an optional translated transcript, or a translated voice track, but no translated video, since there is no picture to put the new voice under. A 45-minute talk costs 45 credits (4.5¢) for a transcript either way.

Recordings longer than two hours usually need splitting as well as extraction; translating long videos explains the limit and where to cut.

Next steps

If you only need text, try uploading the original video to video to text before extracting anything. If you have an audio file already, the guide to transcribing audio walks through the rest, and the audio to text tool takes the file from there.

Frequently asked questions

Does extracting audio from a video reduce quality?

Not if you use a stream copy. Copying the existing audio stream into a matching container, such as AAC into M4A, keeps every bit of the original. Quality only drops when you re-encode to a lossy format, and even then a single re-encode at a moderate bitrate is inaudible for speech. The real damage comes from repeated lossy conversions, so keep one master copy and convert from it.

Why does my extracted audio play on only one side?

The original recording probably put a single microphone on one channel, which is common with camera inputs and some lavalier receivers. The video may have sounded fine in an editor that summed the channels, but a standalone file exposes it. Fix it by extracting to mono, for example with the -ac 1 option in ffmpeg, or by keeping only the channel that contains the voice.

Can I extract audio from a video on a phone?

Many mobile video editors can export a sound-only file, but menus vary by app and operating system version, so check your editor's help pages. If the goal is a transcript or a translation, you usually don't need to: browser-based tools can read the audio directly from the video on the phone. The phone translation guide covers that route.

Is it legal to extract audio from someone else's video?

Extracting audio is technically simple, but the recording still belongs to whoever made it. Personal study use is treated differently from republishing in many places, and rules vary by country and by platform terms. This is general information, not legal advice; for anything you plan to publish, read translating someone else's video legally and get permission where needed.

Which audio format should I pick if I'm not sure?

If you will edit the audio, choose WAV or FLAC. If you only need to listen or transcribe, choose a stream copy when ffmpeg can make one, or M4A otherwise, since it plays almost everywhere and stays small. Avoid exotic codecs and multichannel layouts for speech; a mono or stereo file at a normal sample rate is all a transcription tool needs.