How to run the checklist
Review in passes, each with a single focus. Reading is faster than listening, so start with text, then use timestamps to jump to the moments that need ears. The reviewer for the meaning pass must be fluent in the target language; the timing and sound passes can be done by anyone attentive.
Have these files open side by side: the original video, the dubbed video, the transcript in the original language, the translated transcript and the target-language subtitle file. Keep a log with a timestamp, the problem and a severity: blocker, fix if cheap, or acceptable.
- Read both transcripts side by side for meaning, names and numbers.
- Watch the on-camera and on-screen-action moments for timing.
- Listen to a dense passage, a calm passage and a sample from the second half for pacing and loudness.
- Check the first and last thirty seconds closely.
- Verify captions, file format and the publishing package.
- Meaning
- Negations, numbers, instructions, omissions, invented text
- Names and terms
- Spelling, accidental translation, pronunciation, consistency
- Timing
- Line starts near the speaker, actions match narration, no long gaps
- Pacing
- No rushed stretches, no dragging lines, emphasis on key lines
- Loudness and sound
- Even level, no glitches, a decision about music
- Openings and endings
- First line clean, last word complete, intro music checked
- Captions and files
- Cue timing, reading speed, container accepted, disclosure added
Meaning, names and numbers
This is the pass that catches the errors that embarrass you. Go segment by segment through the translated transcript against the original.
- Negations: "not", "never", "no longer" and their equivalents. A dropped negation reverses an instruction.
- Numbers: prices, dates, quantities, percentages, version numbers and units. Check the digits, then the conversion if the source used units your audience does not.
- Instructions and warnings: step order, safety notes and anything a viewer will act on.
- Omissions: translated segments that are suspiciously short for their source.
- Invented text: stretches of silence or music that produced a stock phrase in the source transcript, such as a sign-off nobody said. These are covered in Whisper hallucinations.
- Names and brands: search the transcripts for every person, product and place. Look for names that were translated as ordinary words.
- Terminology: the same term should be translated the same way throughout. Long videos are processed in chunks, so drift between chunks is worth a search.
Afterwards, listen at the timestamps where names occur. A correctly spelled name can still be pronounced in a way your audience will not recognize.
Timing and sync
Watch the sections where the speaker is on camera and where the narration refers to something happening on screen.
- Line starts: does the voice begin roughly when the speaker starts talking? Automated timing fit may intentionally start a line a fraction early or end it a fraction late; flag only offsets that a viewer would notice.
- Actions: when the narration says "click Save", is the cursor clicking Save at about that moment?
- Gaps: long silences while the speaker's mouth is moving usually mean a segment was missed or merged.
- Collisions: a line cut short or crowding into the next.
Remember that a dub without lip-sync will never match mouth shapes. Judge whether the timing is comfortable, not whether lips match.
Pacing, delivery and loudness
Sample three stretches rather than listening to everything: the densest passage, a calm passage and a few minutes from the second half.
- Rushed lines: fast originals plus longer translations push tempo up; note any stretch that is tiring to follow.
- Dragging lines: very short translations stretched to fill a slot.
- Emphasis: key claims, contrasts and questions should sound like what they are; see prosody in text to speech for why they sometimes do not.
- Voice fit: does the voice style suit the content, and does it stay consistent to the end?
- Loudness: the level should be even across the video and comparable to your other uploads. Many platforms normalize playback loudness; check their published guidance.
- Artifacts: clicks, buzzing or garbled syllables. Listen once on a phone speaker and once on earbuds, because each reveals different problems.
Openings and endings
The first and last thirty seconds get the most attention from viewers and the least from reviewers.
- First line: viewers decide in seconds whether to keep watching. It must be accurate and cleanly delivered.
- Music intros and outros: if the dub keeps the original background, check that title music still sounds right. Separation removes sung vocals along with speech, so a theme song can lose its vocals. Decide whether to accept it, trim it, or, if you have the project, dub a dialogue-only export and mix the translated voice with the original theme yourself.
- Last line: the final word must be complete, not clipped, and the audio should end with the picture.
- Calls to action: website addresses, handles and promo codes read aloud by a synthetic voice often come out mangled. Check them by ear, and put them in the description as text.
Suppose a team dubs a 25-minute compliance training video into French, at 1,250 credits ($1.25). The French-speaking reviewer reads both transcripts first and logs nine issues: two misspelled product names, one dropped negation in a safety step, and six minor wording choices. She then watches the four on-camera minutes, samples three stretches for pacing, and finds that the sung theme in the opening has kept its instruments but lost its vocals. The negation is a blocker; the opening is fixed by dubbing a dialogue-only export of the project and mixing the original intro music under the returned translated voice in an editor; the minor wording choices are accepted.
Captions and deliverables
- Subtitles: open the target-language SRT or VTT and spot-check that cues appear with the speech and are short enough to read comfortably.
- Captions as files: upload them as separate tracks; viewers who prefer reading can turn them on.
- On-screen text: slides, lower thirds and interface text stay in the original language. Decide whether a note or re-exported slides are needed; translating on-screen text in video covers the options.
- Metadata: titles, descriptions and thumbnails need separate translation.
- Disclosure: add a short note that the audio is AI-generated, in line with the platform's rules.
When a full review is worth the time
Not every video needs every pass at full depth. A thorough review pays off for content viewers will act on, such as training, safety, medical or financial explanations; for anything published under a person's name; for videos that will be watched for years, like course modules; and for the first video in a new language, where it reveals systematic issues you can then watch for in the rest of the series. For a casual social clip, the meaning pass and the first and last thirty seconds may be enough.
Common review mistakes and what a checklist cannot catch
- Reviewing only the first few minutes. Problems in long videos cluster later, where fatigue hides them.
- Using a reviewer who is not fluent. A non-speaker can judge timing and sound, not meaning.
- Listening only on studio headphones, which hide problems a phone speaker exposes, and the reverse.
- Blaming the voice for a translation error. If the transcript is wrong, changing voices will not help.
- Treating the checklist as cultural review. It catches errors, not whether a joke, example or reference works for the audience; that takes someone from the market.
QC with mydubly's outputs
A full translation with mydubly's AI dubbing returns everything this checklist needs: the translated MP4 with the AI voice over the original music and effects, the translated audio file, SRT and VTT subtitles in the target language, and timestamped transcripts in both languages. Vocal separation removes sung vocals along with speech and can leave faint traces of the original voice in dense mixes, so the music and intro checks above apply to every mydubly dub. Transcripts do not label speakers, so in multi-speaker videos you identify who said what from the video itself.
Two timing details help you calibrate the timing pass: a line may start up to 0.3 s early or run up to 0.6 s late by design, and tempo stays at or below 1.15× by default, with a little extra only when a chunk cannot otherwise fit. Offsets within that range are expected. The video is muxed with the new audio in your browser without re-encoding the picture; if the video codec cannot go in MP4 the output is MKV, which is worth confirming against your platform's accepted formats.
When a fix is needed, choose the cheapest route. Wording and timing problems in subtitles can be corrected directly in the SRT or VTT file. Problems in the spoken audio require a new run, which costs the same per-minute rate again, so batch your source fixes before re-running.
Where to go next
Pick your next video, dub it with mydubly's AI dubbing, and run these passes before publishing. If the same issues keep appearing across videos, look upstream: challenges of AI dubbing explains which problems come from the source audio and which come from the pipeline.
Frequently asked questions
Who should review an AI dub before it is published?
Someone fluent in the target language should do the meaning pass, ideally from the target market. Timing, loudness and endings can be checked by anyone attentive, including the original creator. For safety, legal or medical content, use a reviewer who knows the subject too.
Can I review a dub without speaking the target language?
Partly. You can check timing, pacing, loudness, endings, captions and files, and you can confirm that names and numbers appear in the translated transcript. You cannot judge meaning or register, so a dropped negation would slip through.
What should I fix first when a dub has problems?
Fix blockers that change meaning, such as negations, numbers and safety instructions, before anything cosmetic. Then fix problems in the first and last thirty seconds, because those are the parts most viewers hear. Minor wording preferences can usually wait.
How do I check loudness without audio tools?
Play the dub next to one of your existing videos on the same device at the same volume, and listen for a noticeable difference. Then skip through the dub to check for jumps between sections. For precise targets, an editor's loudness meter and your platform's guidance are more reliable.
Should I re-run the whole video to fix one bad line?
Only if it is a blocker in the spoken audio. Subtitle errors can be fixed in the SRT or VTT file, and minor wording choices are often acceptable. If you do re-run, collect all source fixes first, since a full translation costs 50 credits per minute each time.