What makes a translation system neural
Older statistical systems were pipelines of separately trained components: word alignment, a phrase table, a language model, a reordering model, and a search procedure that combined their scores. In NMT a single network maps the source sentence to a probability distribution over target sentences, and every parameter is trained together against one objective. There are no hand-built tables to inspect; the knowledge lives in weights.
The architecture that won is the transformer encoder-decoder from 2017. Earlier neural systems used recurrent networks, which read a sentence one token at a time and had to squeeze its meaning into a fixed-size state. Transformers process all tokens in parallel and use attention to connect any token to any other, which both improved quality and made training on large datasets practical. If you want the gentler, historical version of this story, read how AI translation works first.
The encoder-decoder transformer, layer by layer
The encoder is a stack of identical layers; the original paper used six. Each layer has a self-attention block, in which every source token gathers information from every other source token, followed by a small feed-forward network applied to each token separately. Residual connections and layer normalization keep training stable. The output is one context-rich vector per source token.
The decoder is a similar stack with two differences. Its self-attention is masked, so when predicting the fifth target token it can only see tokens one to four. And each decoder layer adds a cross-attention block, where target positions query the encoder's output. Cross-attention is where alignment happens: when the decoder is about to write the French "valise", its cross-attention weights concentrate on "suitcase" in the source. A final linear layer and softmax turn the top decoder state into a probability for each token in the vocabulary.
Training on parallel corpora
An NMT model learns from parallel text: sentences paired with their human translations. Public sources include European Parliament proceedings, United Nations documents, subtitles, software localization files and the large OPUS collection, plus web-mined corpora built by finding pages that exist in two languages and aligning their sentences automatically.
Web-mined data is noisy, so cleaning is a large part of the work. Typical filters remove pairs where one side is in the wrong language, pairs with implausible length ratios, exact duplicates, and pairs whose two sides don't look like translations of each other according to a cross-lingual similarity model.
Training uses teacher forcing: the decoder is fed the correct previous target tokens and asked to predict the next one, and the loss is the cross-entropy between its predicted distribution and the true token. Label smoothing, which moves a little probability mass away from the correct answer, is commonly used and tends to make models better calibrated for beam search. Strong systems for high-resource pairs train on tens or hundreds of millions of sentence pairs.
Subword vocabularies: BPE and SentencePiece
Word-level vocabularies leave too many words unknown, and character-level models make sequences long and slow. The standard compromise is a subword vocabulary learned with byte-pair encoding (BPE) or the unigram method, usually through the SentencePiece toolkit. BPE starts from characters and repeatedly merges the most frequent adjacent pair until the vocabulary reaches a chosen size, often in the tens of thousands.
Most systems share one vocabulary between source and target. That lets the model copy names and numbers through unchanged, because "Kubernetes" or "4K" is tokenized the same way on both sides. The trade-off is fairness across scripts: a vocabulary dominated by English and other Latin-script text splits Hindi, Thai or Amharic into more pieces per word, making those sequences longer and harder to model. Many vocabularies also include a byte fallback so that no input character is ever completely unknown.
Beam search and other decoding choices
At inference the model must search for a high-probability output. Greedy decoding commits to the top token at every step and cannot recover from an early bad choice. Beam search keeps the k highest-scoring partial translations (a beam width of 4 to 10 is common), extends each by every possible next token, keeps the top k again, and stops when the candidates have ended.
Source (French): "Il a pris la porte." The literal reading is "He took the door", and after the first two tokens "He took" a greedy decoder is locked in. With a beam of five, a candidate beginning "He left" survives alongside "He took", and once the model scores the complete sentences, "He left." or "He was shown the door." can outrank the literal version, provided the idiom appeared often enough in training. If it didn't, every beam converges on "He took the door" and the output is fluent and wrong.
Raw beam search favors short outputs, because every extra token multiplies in another probability below one. Engines therefore apply length normalization, dividing the score by a function of length. Too little normalization causes truncated translations; too much encourages padding. Random sampling, popular for creative text generation, is rarely used for translation because faithfulness matters more than variety.
Back-translation and other ways to stretch data
Parallel text is scarce for most language pairs, but single-language text is plentiful. Back-translation exploits that. To improve an English-to-Czech model, you train a Czech-to-English model, run it over a large pile of genuine Czech text, and add the resulting synthetic English-Czech pairs to the training data. The target side is real, fluent Czech, which is what the model learns to produce, so even imperfect synthetic sources help. Marking synthetic pairs with a tag so the model knows they are machine-made tends to work better than mixing them in silently.
Two related techniques are common. Knowledge distillation trains a smaller, faster student model on the translations of a large teacher. Multilingual training puts many language pairs into one model, so a pair with little data can borrow structure learned from related pairs; Ukrainian benefits from Russian and Polish data, for instance.
Where NMT is strong
- Fluency: output usually reads as natural text in the target language.
- Reordering across clauses, which statistical systems handled poorly.
- Morphology: agreement of case, gender and number within a sentence.
- Speed and cost: a mid-sized NMT model translates many sentences per second on modest hardware, which makes it practical for long transcripts.
- Predictability: with beam search, the same input to the same model gives the same output.
Weaknesses and failure modes, and where they come from
Every NMT weakness traces back to something in the design:
- Hallucination under domain shift. When input looks unlike the training data (garbled recognition output, unusual formatting, a single repeated word), the decoder can generate fluent text unrelated to the source, because it is falling back on what is probable in the target language.
- Under-translation. Long or complex sentences sometimes lose a clause, and beam search's preference for short outputs makes it worse. A dropped "not" is the most dangerous version.
- Over-translation and repetition, where a phrase is emitted twice.
- Copying the source through untranslated, a side effect of shared vocabularies and noisy training pairs.
- Default gender and formality. If the source doesn't specify, the model picks whatever was most frequent in training, which reproduces bias ("the nurse" becomes feminine, "the engineer" masculine).
- Sentence isolation. Most engines translate sentence by sentence or in short passages, so pronouns and terms lose their antecedents across boundaries.
Neural machine translation in mydubly
mydubly translates speech, not edited prose, and that changes which failure modes matter. The input to the translation step is the output of speech recognition: Whisper large-v3-turbo returns timestamped segments for each window of roughly 30 seconds, and each chunk's segments are translated by a neural machine translation engine. The specific engine is configurable in the deployment, so this article describes the technique rather than a particular product.
Because recognition comes first, an error such as a misheard name or a missing word reaches the translator as clean-looking text, and NMT will translate it fluently. The practical defenses are a clear recording (see how to improve transcription accuracy) and a review of the source-language transcript, which mydubly delivers alongside the translated one, the SRT and VTT subtitles and the AI-dubbed track. Errors that only appear in the translated transcript belong to the translation step; errors already present in the source transcript came from recognition.
Going further
To see how these mechanics play out on your own material, translate a short recording with the video translator and compare both transcripts line by line. When you want to measure quality more formally, the guide to machine translation quality evaluation explains BLEU, chrF, COMET and human review frameworks, and common machine translation errors gives you a checklist to read against.
Frequently asked questions
How much parallel data does a neural translation model need?
Strong engines for widely used pairs are trained on tens of millions of sentence pairs or more. Usable models can be built from far less when back-translation and multilingual training supplement the data, but quality in low-resource pairs is noticeably less reliable, especially for specialist vocabulary.
Is a transformer the same thing as neural machine translation?
No. NMT is the task of translating with a neural network trained end to end; the transformer is the architecture most NMT systems now use. Earlier NMT systems used recurrent networks, and the transformer itself is used for many tasks besides translation, including speech recognition and text generation.
Why does an NMT model sometimes repeat a phrase or produce text unrelated to the source?
Repetition and hallucination usually appear when the input is unlike the training data or very long. The decoder keeps predicting probable target-language text, and without a strong signal from the source it drifts. Cleaner, sentence-shaped input reduces both problems.
Can a neural machine translation model learn my industry's vocabulary?
Yes, through fine-tuning on in-domain parallel text, or through terminology features in some engines that force particular target terms. Without either, expect general-language choices for specialist terms and plan to review them against a glossary.
Is beam search still used with modern translation models?
Dedicated NMT engines generally still use beam search or a close variant because it improves faithfulness over greedy decoding. Large language models used for translation more often decode greedily or with low-temperature sampling, mainly for speed and because their serving stacks are built for chat.