AI video translation

The A-to-Z of Video Translation Terms

This video translation glossary defines the terms you will meet when translating, dubbing or subtitling video, grouped by the stage of the work they belong to: speech recognition, translation, voice, subtitles, media files and project workflow. Each definition is short and says why the term matters in practice. Terms within each group are in alphabetical order, and the last sections show how the vocabulary fits together in a real job and which pairs of terms people most often mix up.

10 min read · Updated

Speech recognition terms

Every translation of spoken video starts by turning speech into text. These are the words you will meet at that stage.

ASR (automatic speech recognition)
Software that converts spoken audio into text. Also called speech-to-text. Its output quality sets the ceiling for everything downstream, so most translation errors trace back here.
CER (character error rate)
Like WER but counted in characters. Used for languages such as Chinese and Japanese, where text is not split into words by spaces.
Diarization
Working out who spoke when, by grouping stretches of audio by speaker. A separate task from recognition; many transcription tools do not do it.
Hallucination
Text a recognition model produces that was never spoken, most often during silence, music or noise. A reason to trim dead air before processing.
Language identification
Predicting which language is being spoken from the audio itself, so the user does not have to specify it. Less reliable on very short clips and mixed-language recordings.
Segment
A short run of recognized text with a start and end time, typically a phrase or sentence. Segments become subtitle cues and the units that get translated.
Timestamp
The start or end time attached to a segment or word. Timestamps let subtitles appear at the right moment and let reviewers jump to the audio behind a line.
VAD (voice activity detection)
A step that finds where speech is present and where it is not. Used to skip silence or to choose sensible places to cut long audio.
WER (word error rate)
The standard accuracy measure for ASR: substitutions plus deletions plus insertions, divided by the number of words in a correct reference transcript. Lower is better. A full walkthrough is in word error rate explained.
Whisper
An open-source family of multilingual speech recognition models released by OpenAI, trained on a large amount of web audio. It handles transcription, language identification and timestamps.

Translation terms

Back-translation
Translating a translation back into the original language to spot meaning changes. In model training, the same name refers to generating synthetic training data this way.
Formality and register
The level of politeness or familiarity a language marks, such as tú versus usted in Spanish. Speech often leaves it implicit, so machine translation has to guess.
Glossary or termbase
A list of names and terms with approved spellings and translations. The simplest way to keep a product name consistent across a video series.
MT (machine translation)
Automatic translation of text from one language to another by software.
NMT (neural machine translation)
The current mainstream form of MT, using neural networks, usually transformer models, trained on large collections of translated sentence pairs. Explained further in neural machine translation.
Post-editing
A person correcting machine translation output. Light post-editing fixes errors of meaning; full post-editing also polishes style.
Source and target language
The language spoken in the original video and the language you are translating into.
Text expansion
The tendency of a translation to be longer or shorter than the original. It affects how long subtitles stay readable and how fast a dubbed voice must speak.
Transcreation
Rewriting a message for another culture rather than translating it, common for slogans, humor and marketing. A creative task that machine translation does not perform.

Voice and dubbing terms

Dubbing
Replacing the original speech with speech in another language. In film, usually implies matching lip movements; in AI tools, often means a replacement voice track timed to the original lines.
Lip-sync
Matching translated speech to the visible mouth movements of the speaker, either by careful script adaptation or by AI models that redraw the mouth area.
M&E track (music and effects)
The soundtrack minus dialogue, delivered so a new language can be laid on top. Keeping one makes dubbed versions sound like the original mix.
Loudness normalization
Adjusting audio to a target perceived loudness, usually measured in LUFS, so different clips and videos play at a consistent level.
Prosody
The rhythm, stress and intonation of speech. Synthetic voices are judged heavily on it, and translated scripts can lose the original emphasis.
Tempo adjustment or time-stretching
Changing the speed of speech without changing its pitch. Dubbing tools use it to fit longer translations into the original timing.
TTS (text-to-speech)
Software that turns written text into spoken audio. Modern neural TTS can sound natural and speak many languages from one model.
Voice cloning
Generating speech in a specific person's voice from samples of that voice. Raises consent questions and is distinct from choosing a stock voice.
Voice-over
A translated voice laid over the original audio, which remains faintly audible underneath. Common in news and documentaries; compared in dubbing vs voice-over.

Subtitle and caption terms

Captions
Text of the speech in the same language as the audio, often with sound cues, mainly for viewers who are deaf or hard of hearing. In American usage the word often covers subtitles too.
Closed and open captions
Closed captions can be switched on and off by the viewer; open captions are burned into the picture and always visible.
Cue
One subtitle event: the text plus the start and end time when it appears on screen.
Forced narrative
Subtitles that only translate on-screen text or brief foreign-language moments, shown even to viewers who have subtitles turned off.
Reading speed
How fast viewers must read a subtitle, often expressed in characters per second. Long translations in short cues push it too high.
SDH (subtitles for the deaf and hard of hearing)
Subtitles that include speaker identification and sound descriptions, such as music or a door slamming.
SRT
SubRip Text, the most widely accepted subtitle file format: numbered cues, timestamps and plain text. The difference from VTT is covered in SRT vs VTT.
Subtitles
Text translating the speech into another language for viewers who do not understand the audio.
VTT (WebVTT)
The subtitle format used by HTML5 web video players. Similar to SRT, with a header line and support for basic styling and positioning.

File and media terms

AAC
A widely supported compressed audio codec, the usual audio inside MP4 files.
Bitrate
The amount of data used per second of audio or video, usually in kilobits per second. Speech stays intelligible at much lower bitrates than music needs.
Codec
The method used to compress and decompress audio or video, such as H.264 for video or AAC and Opus for audio. Different from the container.
Container
The file format that holds encoded streams together, such as MP4, MOV, MKV or WebM. A container can hold several audio tracks and subtitle streams.
Keyframe
A video frame stored in full, from which following frames are reconstructed. Cutting video without re-encoding can only start cleanly at a keyframe.
Mux and demux
Muxing combines already-encoded audio and video streams into one container; demuxing separates them. Neither changes the encoded data itself.
Opus
An efficient open audio codec, well suited to speech at low bitrates, usually stored in Ogg or WebM files.
PCM
Uncompressed digital audio: a stream of sample values. Large, but the format recognition models ultimately consume.
Re-encoding or transcoding
Decoding a stream and compressing it again. It takes time and loses some quality, which is why swapping only the audio track is preferable.
Sample rate
How many audio samples are stored per second, such as 16 kHz or 48 kHz. Speech recognition commonly works at 16 kHz.

Workflow and project terms

Chunk
A piece of a long recording processed on its own, typically tens of seconds long. Chunking keeps memory use manageable and lets pieces be processed in parallel.
Internationalization
Preparing content so it can be localized easily later, such as keeping on-screen text editable and recording clean dialogue.
Localization
Adapting content for a market, including language, cultural references, formats and sometimes compliance. Broader than translation; see video translation vs localization.
Textless master
A final export of the video with all titles and captions removed, so translated text can be added for each language.
Transcription
Writing down speech in the same language it was spoken. The first deliverable in most translation workflows, and useful on its own.

Why the vocabulary helps: one job, start to finish

Knowing the terms pays off when you read a quote, brief a reviewer or troubleshoot a result, because you can say exactly which stage went wrong. "The WER is fine but the register is off" points to translation, not recognition. "The cue timing drifts" points to timestamps, not the voice.

Worked example

Suppose a 10-minute tutorial is translated into French. The audio is demuxed and decoded to 24 kHz PCM (a transcript-only job would use 16 kHz), then split into about 20 chunks of roughly 30 seconds. ASR returns segments with timestamps; MT turns them into French; TTS voices each sentence; tempo adjustment fits lines that ran long; the clips are loudness-normalized and joined into one AAC track; and that track is muxed with the original video stream, which is never re-encoded. Alongside the video come SRT and VTT files whose cues are the translated segments. In mydubly this costs 10 × 50 = 500 credits, or 50¢.

  1. When a result disappoints, first decide whether the problem is in the text or the sound.
  2. If it is in the text, compare the source transcript with the audio to separate ASR errors from MT errors.
  3. If it is in the sound, check whether the issue is the voice itself (prosody), the fit (tempo adjustment) or the mix (loudness, background balance).
  4. Fix the earliest stage first, since later stages inherit its mistakes.

Commonly confused terms and mistakes

  • Captions and subtitles are used interchangeably in everyday American English, but in localization briefs the distinction matters: same language versus another language.
  • Dubbing and voice-over differ in whether the original voice remains audible. Asking for one and receiving the other is a common mismatch.
  • Container and codec are often conflated. "MP4" names a container; whether a player can open it depends on the codecs inside.
  • Muxing is not the same as encoding. Replacing an audio track by muxing keeps the video quality exactly as it was.
  • Translation, localization and transcreation sit on a scale from literal to creative. A machine translation output is not a transcreation, however fluent it reads.
  • A low WER does not guarantee a good translation, and a fluent translation can hide recognition errors. Each stage needs its own check.

How these terms map to mydubly

ASR
Whisper, with large-v3-turbo as the default deployment; returns timestamped segments per chunk
Language identification
Automatic; you choose only the target language
MT
A neural machine translation engine translates each chunk's segments
TTS
Chatterbox Multilingual by default, one of 8 stock voices per video, with no voice cloning
Chunks
About 30 seconds each, cut at the quietest 50 ms frame within the last 6 seconds before the mark
Upload codec
Ogg Opus at 32 kb/s, about 120 KB per 30 seconds, with WAV as a fallback
Timing fit
Elastic placement plus small tempo changes, then loudness normalization and a gapless AAC track
Mux
Done in your browser without re-encoding the video; MKV output if the video codec cannot go in MP4
Subtitles
SRT and VTT files, never burned in; transcripts do not label speakers

Things on the glossary list that mydubly does not do include diarization, lip-sync, voice cloning, open captions, mixing in a separately supplied M&E track and transcreation. The video file stays on your device; only audio chunks are uploaded, as explained on private video translation.

Where to go next

With the vocabulary in hand, the quickest way to make it concrete is to run a short clip through the video translator and look at each output: the transcripts, the SRT and VTT cues and the translated audio. For a stage-by-stage explanation of the pipeline these terms describe, read how AI video translation works.

Frequently asked questions

What is the difference between ASR, MT and TTS?

They are the three stages of most AI video translation. ASR turns speech into text, MT translates that text into another language, and TTS turns the translated text back into speech. Errors usually start in the earliest stage and carry forward.

What does muxing mean in video translation?

Muxing means packing already-encoded streams, such as the original video and a new audio track, into one file without compressing them again. It is how a translated voice can be added to a video without any loss of picture quality.

Is transcreation the same as localization?

No. Localization is the whole process of adapting content for a market, while transcreation is one creative technique within it, used when a line must be rewritten rather than translated, such as a slogan or a joke.

What is an M&E track and do I need one?

It is the full soundtrack with the dialogue removed. Traditional dubbing needs one to keep the original music and sound effects. mydubly separates the background automatically, so you don't need one, but if you have it, dubbing a dialogue-only export and mixing the translated voice with your M&E track gives the cleanest result.

Why does subtitle reading speed matter for translation?

Translations can be longer than the original, but the cue stays on screen for the same time. If there are too many characters per second, viewers cannot finish reading, so long lines may need shortening or splitting.