Preparing your files

Screen Recording with Audio: Microphone, System Sound and Clean Narration

A screen recording can capture two kinds of sound: your microphone and the computer's own audio. For tutorials and demos, record the microphone for narration and add system audio only when the viewer must hear it, ideally on a separate track. Built-in recorders on macOS and Windows handle the microphone easily; capturing system sound cleanly usually needs a tool such as OBS Studio.

8 min read · Updated

Microphone audio, system audio, or both

Microphone audio is your voice and everything else the microphone hears: the room, the keyboard, the fan. System audio, also called desktop or internal audio, is what the computer plays: app alerts, a video you are demonstrating, the other people on a call.

Most tutorials and product demos need only the microphone. System audio matters when the sound is part of what you are teaching, such as an audio editing app, a video player feature or a game. It is also where most problems start: notification pings, music left playing in another tab, and the doubled voice that appears when the microphone is captured twice.

Decide before you record which sounds the viewer actually needs. Leaving system audio off is a legitimate choice, and the safest one for narration-led videos.

Recording audio on macOS

The built-in recorder opens with Shift-Command-5 (or from QuickTime Player's New Screen Recording). Its Options menu lists microphones; choose your external or built-in microphone there, because the default can be None, which produces a silent video.

The built-in recorder has historically not captured the sound the Mac itself plays, and that has been true through recent versions. To include system audio, use an app built on Apple's screen capture framework, such as OBS Studio's macOS screen capture source, or route sound through a virtual audio device. Apple changes these features between releases, so check the current macOS documentation for your version.

The first time any app records the screen or microphone, macOS asks for permission under Privacy and Security in System Settings. If a recording is silent or the app can't see your microphone, that permission is the first thing to check.

Recording audio on Windows

Windows has two built-in options. The Xbox Game Bar (Windows key plus G) records a single app window with its sound and can include the microphone, but it can't record the desktop or File Explorer. Newer versions of the Snipping Tool on Windows 11 can record a screen region and offer microphone and system audio toggles. Both are updated through the Microsoft Store, so the options you see depend on your version.

In Settings, under System and Sound, check which input device is the default and test its level, because recorders often pick the default microphone without asking. A headset plugged in last week may still be the default.

Keeping voice and system sound on separate tracks

OBS Studio, which is free on Windows, macOS and Linux, treats each sound source separately: Desktop Audio and Mic/Aux appear in its audio mixer with their own volume faders and filters. Its advanced output settings can also write several audio tracks into one recording, and the Advanced Audio Properties window decides which source goes on which track.

A practical layout is track one for a mix of microphone and system audio, track two for the microphone alone, and track three for system audio alone. Track one is what plays by default. Tracks two and three let you rebalance, fix a loud alert or remove a section of system sound in an editor afterwards. Choose MKV or MOV as the recording format for multiple tracks, and remux to MP4 later if needed; OBS has a remux tool in its File menu. The multiple audio tracks article explains how players choose between tracks.

How double audio and echo creep in

Doubled or echoing narration has a few repeatable causes:

  • Monitoring turned on. If OBS's audio monitoring or a system setting plays your microphone back through the speakers or headphones that the desktop capture records, your voice is captured twice, a few milliseconds apart, and sounds hollow.
  • Speakers instead of headphones on a call. The microphone hears the other person through your speakers while the system capture records them directly, so their voice appears twice.
  • Two captures of the same app. Capturing an application's audio and the whole desktop audio at once records that app twice.
  • A virtual audio device routed in a loop, so the output feeds back into the input.

Wear headphones whenever system audio is recorded, keep monitoring off unless you need it, and capture each sound only once. A ten-second test recording reveals every one of these problems before you lose a full take.

Levels and settings that keep narration clean

  • Set the microphone so your normal speaking voice peaks around minus 12 to minus 6 dBFS on the recorder's meter, with loud moments never touching zero.
  • Keep system audio well under the voice. If a demo sound plays while you talk, the voice must win.
  • Use a noise suppression filter sparingly. OBS includes one, and it helps with fans, but heavy settings make speech sound watery, which also hurts recognition.
  • Switch on Do Not Disturb or Focus so notification sounds and banners don't land in the recording.
  • Keep the microphone away from the keyboard, or use a quieter keyboard; typing transcribes as nothing but drowns out quiet words.
  • Leave sample rate at the recorder default, usually 48 kHz, and audio bitrate at the default or higher. Speech doesn't need more; the speech recording settings article explains why.

Microphone choice matters more than any of these settings; choosing a microphone for speech covers it.

Example: a support walkthrough recorded for translation

Help-center tutorial on a Mac

A support engineer records a 9-minute setup walkthrough in OBS, with a USB microphone and the app's alert sounds she wants viewers to hear. She plans to publish English captions and a Spanish dubbed version.

She sets up three tracks: the mix on track one, microphone on track two, system sound on track three, with headphones on and monitoring off. A short test shows an alert sound peaking louder than her voice, so she lowers Desktop Audio until her voice clearly dominates.

The English transcript and captions come from the default mixed track, for 9 credits (0.9¢). For the Spanish version, she knows the translated video keeps the alert sounds under the new voice but lowers them while the voice speaks. She dubs it for 450 credits (45¢), then lays the system-audio track from her original recording under the Spanish voice in her editor, so the alerts play at their full, original level.

A checklist before you press record

  1. Choose the input microphone in the recorder itself, not just in system settings.
  2. Decide whether system audio is needed at all; if not, turn it off.
  3. If it is needed, put on headphones and confirm that monitoring is off.
  4. In OBS or similar, assign microphone and system sound to separate tracks and keep a mixed track first.
  5. Close music, video tabs and chat apps, and turn on Do Not Disturb.
  6. Record ten seconds of talking with one system sound, then play it back on headphones.
  7. Check that the voice is clear, nothing is doubled, the level isn't clipping and the alert sits below the voice.
  8. Record the real take, and leave a second of silence at the start and end.

What you can't fix after recording

Some problems are cheap to prevent and expensive or impossible to repair. Echo from doubled capture can't be cleanly separated from a single mixed track. Music or alerts recorded at the same level as your voice can't be pulled apart perfectly, even by the automatic voice separation mydubly applies to dubs, which can leave faint traces of your voice or a thinner-sounding background, and the background music article explains why voice and music are hard to separate. Clipped peaks lose information for good.

Separate tracks are the main safeguard: they turn most mistakes into a volume adjustment in an editor instead of a re-record. For videos you know will be translated, the broader production advice in making videos translation-ready applies too, such as speaking in complete sentences and leaving small pauses between steps.

How mydubly uses the audio in a screen recording

mydubly reads the audio from MP4, MOV, WebM, MKV and M4V files in the browser, so a screen recording can be uploaded as it is, up to 2 hours long. If the file has several audio tracks, only one is used, so make the track with your narration the first or only track before uploading.

Speech recognition works from that one track. Clear narration with system sounds kept low gives a cleaner transcript and SRT or VTT subtitles. Alert sounds and music aren't transcribed, but loud ones can mask words.

For a translated video, your narration is removed by AI vocal separation and the new voice is mixed over what remains, so app sounds and music recorded with the screen are kept, lowered while the new voice speaks. The picture is not re-encoded, so screen text stays as sharp as you recorded it, but any text on the screen stays in its original language. If you want those sounds at full level or perfectly clean and recorded system audio on its own track, dub a narration-only export and mix the translated voice with your system-audio track in an editor; adding that track to the audio of a normal dub would double the sounds. Translating finished recordings, including UI naming tips, is covered in the screen recording translation use case.

Next step

Run the ten-second test on your current setup today, before your next real recording. When you have a clean take, upload it to video to text for a transcript and captions, or to AI dubbing for a narrated version in another language.

Frequently asked questions

Why does my Mac screen recording have no sound?

Two causes cover most cases. The built-in recorder's Options menu may have no microphone selected, so nothing is recorded. Or you expected the sound the Mac was playing, which the built-in recorder has historically not captured. Pick a microphone in Options, check that the recording app has microphone permission in System Settings, and use a tool with system audio capture if you need the computer's sound.

Can I record system audio without recording my microphone?

Yes. In OBS, mute or remove the Mic/Aux source and keep Desktop Audio. In Windows recorders with separate toggles, switch the microphone off. This is useful for capturing a webinar replay you have rights to use, or a demo where the app's sound is the point. Keep notification sounds off, because they will be recorded at the same level as everything else.

Should I record narration live or add a voice-over afterwards?

Live narration is quicker and sounds natural, but every stumble means redoing the screen action too. Recording the screen silently and narrating afterwards in an audio editor gives cleaner audio, lets you script the wording, and makes retakes easy. For videos that will be transcribed or translated, a scripted voice-over usually produces clearer sentences and fewer errors.

What frame rate should a screen recording use?

For tutorials and software demos, 30 frames per second is plenty and keeps files smaller. Use 60 when you show fast motion, such as scrolling, animation or gameplay. Frame rate has no effect on audio quality or on transcription. Some recorders use variable frame rate to save space, which can cause sync issues in certain editors, so check that setting if you edit heavily.

How do I stop keyboard noise in a screen recording?

Move the microphone away from the keyboard and closer to your mouth, use a directional microphone pointed away from the desk, or put a soft mat under the keyboard. A noise suppression filter reduces steady sounds better than sharp clicks. If you record the voice-over separately after the screen capture, there is no typing in the narration at all.