AI translation

AI translation, explained: how a machine turns one language into another

AI translation works by turning a sentence into numbers, letting a neural network work out how each piece of the sentence relates to every other piece, and then generating the translation one small unit at a time, with every choice conditioned on the source and on what has already been written. That is the short answer. The longer one runs through roughly seventy years of attempts, from hand-written grammar rules to statistical phrase tables to today's transformer models, and that history explains a lot about why modern systems succeed and fail the way they do.

8 min read · Updated

Rule-based translation: grammars written by hand

The first public demonstration of machine translation, the Georgetown–IBM experiment in 1954, converted a few dozen carefully chosen Russian sentences into English using a small vocabulary and a handful of rules. For the next several decades the dominant approach stayed the same in spirit: linguists wrote bilingual dictionaries and grammar rules, and the software applied them. A typical rule-based system parsed the source sentence into a structure (subject, verb, object, modifiers), transferred that structure into the target language's conventions, then generated the output word by word.

Rule systems were predictable: if a translation was wrong you could trace the rule that produced it. The problem was scale. Every exception had to be written by a person, and natural language is mostly exceptions. "Time flies like an arrow" can be parsed as a statement about time, an instruction to time some flies, or a claim about insects called "time flies". Humans resolve that instantly from world knowledge; a rule system needs another rule, and then another.

Statistical machine translation: letting data decide

Around 1990, researchers at IBM tried learning translation probabilities from text that already existed in two languages, starting with the Canadian parliamentary record, published in English and French. By the 2000s this had matured into phrase-based statistical translation: a phrase table of source phrases with candidate translations and probabilities, a language model scoring fluency, and a search that picked the highest-scoring combination. Google Translate ran on this kind of system for about a decade before moving to neural models in 2016.

Statistical systems were far more robust than rule-based ones, but they worked on short local phrases. They struggled when meaning depended on words far apart, such as a German verb at the end of a long clause, and often produced output that was fluent in three-word stretches and incoherent as a whole.

Neural machine translation: one network for the whole sentence

Around 2014, sequence-to-sequence neural networks showed that a single model could read an entire source sentence and generate a translation directly, with no phrase table at all. A year later, the attention mechanism let the model look back at any part of the source while writing each output word. In 2017 the transformer architecture, introduced in the paper "Attention Is All You Need", made attention the core of the whole network and trained much faster on modern hardware. Almost every translation system in use today, dedicated or general-purpose, descends from that design.

If you want the engineering detail (encoder and decoder layers, training on parallel text, beam search, back-translation), the companion article on neural machine translation goes deeper. The rest of this piece follows a single sentence through the process.

Tokens and embeddings: turning words into numbers

A neural network cannot read letters. The first step is tokenization: splitting the text into units from a fixed vocabulary, usually a few tens of thousands of entries. Modern systems use subword tokens, so common words are a single token while rarer words are split into pieces. A word like "untranslatable" might become "un", "translat" and "able", depending on the vocabulary. This lets the model handle words it has never seen whole, including new product names, by composing them from familiar parts.

Each token is then mapped to an embedding: a list of several hundred or several thousand numbers learned during training. Tokens used in similar ways end up with similar embeddings, so "car" sits near "vehicle" and the German "Auto". Because a transformer processes all tokens in parallel, each token's position is added to its embedding, leaving a grid of numbers, one row per token.

Attention: how the model decides what each word depends on

Attention is the step that lets context flow between words. For every token, the model computes how relevant every other token is and blends in information from the relevant ones. After several layers of this, the representation of each token reflects the whole sentence around it.

Consider "The trophy didn't fit in the suitcase because it was too big." To translate "it" into French, the model has to decide between "il" (agreeing with the masculine "trophée") and "elle" (agreeing with the feminine "valise"). If the sentence ends "too big", the trophy is the problem; if it ends "too small", the suitcase is. Attention gives the representation of "it" access to "trophy", "suitcase" and "big" at once, which is what makes that decision learnable. When the network writes the output, a second form of attention, cross-attention, lets each output token consult the encoded source sentence directly.

Producing the translation word by word

The output is not produced all at once. The decoder generates one token, appends it to what it has written, and repeats until it emits a special end-of-sentence token. At each step it produces a probability for every token in the vocabulary.

Worked example: German to English

Source: "Ich habe den Zug verpasst." Step 1: the model reads the whole German sentence. Step 2: the most probable first token is "I". Step 3: for the second token, cross-attention focuses on "verpasst", the verb at the very end of the German clause, and "missed" scores far above "have". Step 4: "the", then "train", then a full stop, then the end token. Result: "I missed the train." A word-for-word system would have produced "I have the train missed."

That verb jump in step 3 is exactly what statistical systems found hard and what attention handles naturally. Choosing the single most probable token at each step is called greedy decoding; dedicated translation engines usually run beam search instead, keeping several candidate translations alive and choosing among complete sentences at the end.

Where large language models fit in

Large language models behind chat assistants are also transformers that generate token by token. The difference is training: a dedicated translation model learns mainly from sentence pairs and does one job, while an LLM learns from enormous amounts of mostly single-language text and is told what to do through a prompt. That makes it good at instructions like "use formal German" and at reading a whole document, but more prone to adding things nobody asked for. See LLM translation vs machine translation for the trade-offs.

Strengths and limits of modern AI translation

Modern systems are very good at the things that used to be hardest:

  • Fluent, grammatical output in widely used language pairs.
  • Long-distance reordering, such as German verb-final clauses or Japanese subject-object-verb order.
  • Frequent idioms and fixed expressions that appeared often in training data.
  • Agreement within a sentence: gender, number and case on articles and adjectives.

The weaknesses follow from the mechanics described above:

  • The model predicts plausible text, so a mistranslation usually reads smoothly, which makes it easy to miss.
  • Rare names and specialist terms may be split into odd subword pieces and translated as if they were ordinary words.
  • Most systems translate a sentence or a short passage at a time, so information from earlier paragraphs is often lost. The article on context in machine translation explains what that does to pronouns and formality.
  • Garbled or ungrammatical input produces confident nonsense rather than an error message.

A fuller catalogue with examples is in common machine translation errors.

How AI translation fits into mydubly's pipeline

In a video or audio translator, text translation is the middle step of three. In mydubly, the browser extracts the audio and splits it into windows of about 30 seconds, cut at quiet moments. Speech recognition runs on OpenAI's open-source Whisper model (large-v3-turbo by default) and returns timestamped segments, which a neural machine translation engine then translates chunk by chunk; the engine is configurable in the deployment. The translated text becomes subtitles or is voiced by Chatterbox Multilingual and fitted back to the original timing.

The spoken language is detected automatically and you choose one of 21 target languages. Because transcripts in both languages come with the translated video, the SRT and VTT files and the translated audio, you can see exactly what the translation step produced. If you only need text, the subtitle generator and audio translator use the same recognition and translation steps.

Try it on a short clip

Run a two- or three-minute clip you know well through the video translator and read the two transcripts side by side. A three-minute dub costs 150 credits (15 cents), and the paired transcripts show where the model handled reordering and idioms gracefully and where a name or pronoun went astray.

Frequently asked questions

Does AI translation understand meaning, or does it just match patterns?

It learns statistical patterns, but at a scale where those patterns capture a lot of meaning: which words refer to each other, which senses fit which contexts, how sentences are structured. It has no access to the world beyond its training text, so it fails when the right translation depends on knowledge that isn't in the sentence, such as who 'she' is in a video you've only partly translated.

Why does AI translation sometimes produce sentences that sound perfect but are wrong?

The decoder picks tokens that are probable given the source and its own previous output, and fluent text is always probable. When the model misreads the source, it still writes a smooth, grammatical sentence around the misreading. That is why reviewers should compare against the source rather than judging the translation on how natural it sounds.

Is Google Translate rule-based, statistical or neural?

Google Translate used phrase-based statistical translation for roughly its first decade and announced its switch to neural machine translation in 2016. The underlying models have changed several times since, so check Google's own documentation for current details.

Why do AI translators break words into pieces?

A vocabulary of whole words would either be enormous or leave many words unknown. Subword pieces keep the vocabulary to a manageable size while still letting the model spell out any word, including names and newly coined terms, from smaller units it has seen before.

Can AI translate between two languages it never saw paired together?

Multilingual models trained on many language pairs can sometimes translate between two languages that never appeared together in training, by routing through shared internal representations. Quality in these zero-shot directions is usually lower than for pairs with plenty of direct training data.