What an M&E track contains
M&E stands for music and effects. In film and television it is a standard deliverable, prepared so that distributors in other countries can record a new dialogue track and mix it with the original soundtrack. The same idea works for any video you might translate.
A complete M&E contains everything a viewer hears except the words that will be translated:
- Music
- Score, library tracks, stings and transitions.
- Sound effects
- Whooshes, impacts, interface sounds, anything added in the edit.
- Ambience
- Room tone, street noise, nature sounds and other backgrounds.
- Foley and production effects
- Footsteps, door closes, objects handled on camera.
- Non-verbal human sounds
- Crowd noise, laughter or applause that should stay in every language.
What it leaves out is dialogue, narration and voice-over. Song lyrics are a judgment call: if they are part of the music and should not be translated, they stay in the M&E; if a song carries meaning you want to subtitle or adapt, note it separately.
Stems, and why the dialogue must stay separate
Professional mixes are often delivered as stems: a dialogue stem, a music stem and an effects stem, which together add up to the final mix. An M&E is simply the music and effects stems combined, or exported together, with dialogue muted.
For creators, the practical version is two exports per finished video: a dialogue-only stem and an M&E stem, each the full length of the final edit, starting at the first frame. Played together at their original levels, they should sound like your published mix. If they do not, something is on the wrong track.
The dialogue stem matters as much as the M&E. Speech recognition hears speech more clearly without music competing with it, and a clean dialogue file is the right input for translation. A clear voice in the original recording is the foundation of both; recording with multiple microphones covers keeping voices separate from the start.
Why export one from every project
It is tempting to export an M&E only for videos you plan to translate. Exporting it every time is cheaper in the long run, for several reasons:
- You often decide to translate a video months after publishing, when it turns out to perform well. By then the project may be archived, the music library account cancelled or the plug-ins updated.
- It lets you re-cut a video, replace a sponsor read or fix a line of narration without remixing the whole soundtrack.
- It protects against losing the project file. A finished mix cannot be cleanly separated; an M&E stem plus a dialogue stem can always be recombined.
- It costs almost nothing: a few minutes of export time and a file the same length as the video.
Store the M&E next to the master and the dialogue stem, with the same date and slug in its name, and back it up as carefully as the master.
How to export a clean M&E
The work happens during the edit. A few habits make the export trivial:
- Keep every voice on dedicated dialogue tracks in your timeline, and nothing else on them.
- Put music, effects and ambience on their own tracks, even short stings and whooshes.
- When the edit is locked, mute the dialogue tracks and export the rest as a WAV the full length of the timeline. Use the same sample rate as your project.
- Solo the dialogue tracks and export them the same way.
- Play both files together against the master to confirm nothing is missing.
Keep any processing that belongs to the whole mix, such as a final limiter, in mind. If it sits on the master bus, it will act differently on each stem than on the full mix, which matters little for dubbing but explains small level differences.
The filled M&E problem
The hard part of an M&E is sound that was recorded on the dialogue microphone. When you mute dialogue, everything that mic captured disappears with it: the room ambience behind the speaker, a door closing on camera, a laugh from someone off screen, the sound of the product being unboxed.
In professional work, an M&E where these gaps have been replaced is called fully filled. Editors rebuild them with room tone recorded on set, effects from a library and new foley. For creator videos, a lighter approach is usually enough:
- Record 30 seconds of silence in the room at every shoot and lay it under the M&E to fill the holes.
- Recreate the handful of on-camera sounds viewers would miss, such as a click or a pour, from a library.
- Accept small gaps in talking-head videos, where music usually covers them.
If on-camera sounds are central to the video, such as cooking, ASMR or product demos, plan for this at the recording stage with a separate microphone for the action.
When a project has no M&E
Older videos often exist only as a final mix. Your options, roughly in order of quality:
- Reopen the project and export stems, if the project and all its media still exist.
- Rebuild the soundtrack, if you still have the raw voice recording: dub that, then put the original music track from your library and a few key effects under the translated voice. Often quicker than it sounds for talking-head videos. Don't add music under a dub of the final mix, which already carries the original background.
- Dub the final mix as it is. mydubly separates the original speech from the background automatically and keeps the music and effects under the new voice. This works well on simple mixes; on dense ones faint traces of the original voice can remain and the background can sound thinner, so listen closely before publishing.
- Choose subtitles for that video instead of a dub, which keeps the original mix untouched.
What to listen for in the kept background, and the step-by-step dialogue-only route with stems, are in keeping background music in translated videos, so it is not repeated here.
Example: two videos, two outcomes
A travel creator decides to dub two videos into Spanish. The newer one, 18 minutes long, has a dialogue stem and an M&E exported at the end of the edit. The dialogue stem is dubbed for 900 credits (90¢), and the Spanish voice track is laid over the M&E in the editor in one pass. The older one exists only as a final mix with music under every line, so it is dubbed as it is. The kept background sounds thinner than the original in busy passages. With no dialogue-only file, adding the library music on top would only double it, so the creator accepts the dub for the main sections and chooses subtitles for a market-scene sequence where the crowd sound matters.
Limits and mistakes to avoid
- Leaving a voice in the M&E. A guest's line on a music track, or a narration take on an effects track, ends up untranslated under the dub. Play the M&E alone before using it.
- Exporting stems of different lengths or start points. They must line up with the master frame for frame.
- Ducking baked into the music. If you automated music dips under speech, the M&E has holes where the original voice was. For dubbing, an undipped music stem is easier, since the dub's timing differs slightly; keep the automation in the project and redo it for each language.
- Forgetting licenses. Exporting a stem does not change what your music license allows; check that it covers the translated version you plan to publish.
An M&E cannot fix everything. If music was recorded live with the speaker, as in a performance or a vlog with a street musician, it is on the dialogue mic and cannot be separated by exporting stems.
How the M&E fits a mydubly dub
mydubly removes the original speech from whatever file you upload and keeps the rest, mixing the music and effects under the dubbed voice and lowering them while it speaks. Uploading the full mix therefore gives you a dub with its background without any M&E, within the limits of separation: in dense mixes faint traces of the original voice can remain, the background can sound slightly thinner, sung vocals are removed along with the speech, and the audio is mono. The M&E is the route to the cleanest, fully controllable result. For that, upload the dialogue stem, which gives cleaner recognition than the full mix and comes back with little or nothing under the voice, and use the translated audio mydubly returns as an audio file. Because each dubbed line is timed to where the original line was, allowing it to start up to 0.3 seconds early or end up to 0.6 seconds late, the translated audio lines up with your timeline when placed at zero, alongside the M&E.
mydubly does not take an M&E file as an input and does not mix your stems for you. That remix happens in your video editor; editing dubbed audio in a video editor covers the editing side. If you only have the final mix, mydubly can still transcribe and dub it, and the background is kept automatically.
Next step
Add the two stem exports to the end of your editing routine this week, and keep the textless picture alongside them as described in textless versions. When a video is ready to translate, dub the dialogue stem with AI dubbing, then follow how to dub a video with AI for the full job.
Frequently asked questions
What does M&E stand for in video?
M&E stands for music and effects. It is the soundtrack of a video with the dialogue removed, so a new language can be recorded or generated and mixed with the original music and sound design. In film and television it is a standard delivery item.
Is an M&E the same as an instrumental?
Not quite. An instrumental is a music track without vocals. An M&E contains all the music plus sound effects, ambience and on-camera sounds, everything in the video except the spoken dialogue.
Should narration be in the M&E?
No. Narration and voice-over are speech that will be translated, so they belong on the dialogue stem. Leaving narration in the M&E means it plays untranslated under the dubbed voice.
What file format should an M&E be?
An uncompressed WAV at your project's sample rate is a safe choice, exported the full length of the final edit from the first frame. Use the same length and start point for the dialogue stem so both line up with the master.
Can I create an M&E from a finished video?
Only approximately. Source separation can split a mix into voice and background, with varying artifacts depending on how dense the mix is; mydubly does this automatically when you dub a finished video, so you may not need a separate M&E at all. If you still have the raw voice recording, dubbing that and rebuilding the soundtrack from your original music files is often cleaner for simple videos.
Do I need an M&E if I only use subtitles?
Not for the subtitled version, since the original mix stays as it is. It is still worth exporting, because it costs little and keeps the option of dubbing, re-cutting or replacing a line later.