Sentence-level and document-level translation
Most machine translation systems were built and evaluated sentence by sentence. Training data was aligned sentence to sentence, and the standard test sets scored sentences independently. A sentence-level engine translates "It was fixed yesterday" without knowing what "it" is.
Document-level translation feeds the model surrounding sentences, or a whole document, so it can resolve those references. Research systems do this with extended input windows; large language models do it naturally because they accept long prompts. More context is not free: inputs get longer, slower and costlier, and very long inputs bring their own risk of skipped passages. In practice, most production pipelines settle on a window somewhere between one sentence and a full document.
Pronouns that need an antecedent
English pronouns carry little grammatical information, but their translations often must agree with the noun they refer to.
- English "it" becomes German "er", "sie" or "es" depending on the grammatical gender of the noun: der Server (er), die Datei (sie), das Fenster (es).
- English "they" for a group of women becomes French "elles", for a mixed group "ils".
- Going the other way, Spanish, Italian, Japanese and Chinese routinely drop subjects. "Llegó tarde" can mean he, she, or you (formal) arrived late. Japanese "Kinō ikimashita" means "went yesterday" with no subject at all, and English requires one: I, we, he or they.
When the antecedent is in the same sentence or the one before, a context-aware engine usually gets this right. When it is several sentences back, the engine guesses, and the guess is typically the most frequent form in its training data.
Grammatical gender of speakers and listeners
Some languages mark the gender of the speaker or the listener, not only of nouns. A woman saying "I'm ready" in French says "je suis prête"; a man says "je suis prêt". Russian past-tense verbs agree with the speaker: "ya ustal" (man) and "ya ustala" (woman) for "I got tired". In Hebrew and Arabic, "you" and its verb forms differ for a man or a woman addressed.
The source text frequently contains no clue. A text translator does not hear the voice and does not see the speaker, so it applies a default, usually masculine. If a woman narrates a whole video, a sentence-level engine may switch between forms whenever a line happens to contain a gendered word, and default back elsewhere.
Formality and politeness across a whole text
Choosing between tú and usted, du and Sie, or a Japanese polite or plain style is a decision about the relationship, and it should stay fixed throughout a piece. A sentence-level engine decides afresh for each line. Short, neutral lines like "Click here" or "Let's begin" give it nothing to go on, so a translated tutorial can address the viewer as "Sie" in one sentence and "du" in the next. Readers notice this immediately; it is one of the most common complaints about machine-translated video.
Ambiguous words and the topic of the text
Many words have senses that only the topic resolves. "Pitch" is a sales presentation, a sports field, a musical frequency or the slope of a roof. "Court" is a legal institution or a tennis surface. "Release" is a software version, a press statement or a legal waiver. "Charge" can be electrical, financial or criminal.
Within one sentence, nearby words usually settle it: "the pitch deck" is unambiguous. The trouble is short sentences that rely on the topic established earlier. In a talk about startup fundraising, "The pitch went badly" is a presentation; translated in isolation, it could easily become the sports-field sense in some languages.
What mydubly's roughly 30-second chunks mean for context
mydubly processes speech in windows of about 30 seconds. Each cut is placed at the quietest 50-millisecond moment within the last six seconds before the 30-second mark, so windows usually run between 24 and 30 seconds and end in a pause rather than mid-word. Whisper transcribes each window into timestamped segments, and a neural machine translation engine translates each chunk's segments as a unit.
That has two consequences. Inside a chunk, the translator works with several sentences of spoken context, typically enough to resolve a pronoun whose noun was mentioned a few seconds earlier and to keep formality consistent within that stretch. Across chunks, there is no guarantee that what was said in one window informs the next. A referent introduced a minute ago, a speaker's gender stated at the start, or a formality choice made in the first chunk may not carry over.
Chunk 3 ends: "Last week we moved the server into the basement." Chunk 4 begins: "It runs much cooler there now." Translated into German with chunk 3 visible, "it" should be "er", agreeing with "der Server". With only chunk 4, the engine has no antecedent and is likely to choose the neutral "Es läuft dort jetzt viel kühler", which German speakers will read as "Things are running much cooler there" rather than a statement about the server. The sentence is grammatical and plausible, which is exactly why it slips past a quick review.
Cutting at pauses helps: chunk boundaries tend to fall between sentences, so individual sentences are never split. What it cannot do is give a later chunk the memory of an earlier one.
When more context helps, and where it falls short
More context reliably fixes pronoun agreement, keeps terminology and formality consistent, and disambiguates topic-dependent words. It has limits, though:
- It cannot recover information that was never stated. If nobody says the narrator's gender, even a whole-document model has to guess.
- Text translators never see the picture. "Put it over there" with someone pointing is unresolvable from words alone, and on-screen labels that explain a term are invisible.
- Very long inputs raise the risk of omissions and make each request slower and more expensive.
- Context can mislead. An early, unusual use of a term can bias later translations in the wrong direction.
Practical ways to give the translator more to work with
If you write or record the source, you can reduce context errors before they happen:
- After a pause or a topic change, repeat the noun rather than starting with "it", "this" or "they".
- Introduce people by name and role, and use names again when the conversation returns to them.
- Avoid one-word or verbless lines that only make sense with what came before.
- Mention the domain early and plainly, so ambiguous terms land in the right sense within the same stretch.
- When reviewing, read the translated transcript start to finish and note every switch in formality, speaker gender or the translation of a key term; these are the signature of missing context. The post-editing guide covers how to fix them efficiently.
Next step
To see how context affects your own material, run a few minutes of a real video through the video translator and read the target-language transcript continuously, not line by line. Pay attention to the first sentence after each pause. For a broader list of what to look for, see common machine translation errors, and for pairs where formality matters most, compare results on English to German or English to Japanese.
Frequently asked questions
Does document-level machine translation solve the pronoun problem?
It solves most cases where the antecedent appears in the text, because the model can see which noun a pronoun refers to. It cannot solve cases where the information was never stated or is only visible on screen, and long inputs introduce their own risk of skipped passages.
Why did the translation switch between formal and informal 'you'?
Each passage was translated with limited context, and short neutral lines don't signal which form is appropriate. The engine chose independently each time. Consistent formality usually requires a reviewer to standardize it, or an engine that supports a formality setting or document-level input.
Can a translation engine use what is shown on screen as context?
Text translation engines see only text. If a speaker points at a chart and says 'this one', or a label on screen defines a term, that information is lost unless it is also spoken. Narrating key visual information aloud makes translations more accurate.
Why are audio chunks cut at quiet moments rather than exactly every 30 seconds?
A fixed 30-second cut would regularly split words and sentences, which damages both recognition and translation. Placing each cut at the quietest moment shortly before the mark usually lands it in a pause between phrases, so every chunk contains whole sentences.
How can I tell which gender a translation assumed for the speaker?
Look for words that must agree with the speaker in the target language: past-tense verbs in Russian or Polish, adjectives after 'I am' in French or Spanish, verb endings in Hindi. Search the translated transcript for a few of these and check they match the actual speaker.