Accessibility & inclusive video

Multilingual Accessibility: Keeping Every Language Version of a Video Usable

Multilingual accessibility means a deaf, blind or screen reader user gets the same access in every language version of a video, not just in the original. In practice that takes reviewed captions of equal quality in each language, correct language tags on caption tracks and pages, controls and switchers a screen reader can announce properly, right-to-left captions that display correctly, and translated transcripts and description wherever the original has them.

9 min read · Updated

What multilingual accessibility covers

Accessibility work usually starts in one language. The original video gets careful captions, a transcript and perhaps audio description, and then translation adds subtitles in other languages that nobody checks with the same care. The result is that a deaf viewer who reads Spanish, or a blind viewer who listens in Arabic, gets a weaker version than an English speaker.

A useful test is parity: for each language you publish, could a disabled viewer who speaks only that language get everything a disabled viewer gets in the original? That covers four layers:

  • Text alternatives for sound: captions and transcripts in each language.
  • Text alternatives for the picture: audio description or a descriptive transcript in each language.
  • Machine-readable language information, so players, browsers and assistive technology handle each language correctly.
  • Interfaces such as language switchers and caption menus that work with a keyboard and a screen reader.

Legal obligations to provide information in particular languages are a separate question, discussed for public bodies in translating public information videos. This article covers how to make each version usable once you have decided to publish it.

Caption quality parity across languages

Translated captions often receive less review than the original because the team cannot read them. That is exactly where errors survive.

  • Review every language with a fluent reader, ideally one who knows the subject. An AI draft is a starting point, not a finished caption file.
  • Translate non-speech information too. If the original captions say [door slams] or [upbeat music], the translated file needs the equivalent in the target language, written to that language's captioning conventions.
  • Check reading speed per language. Some languages need more characters to say the same thing, so a cue that reads comfortably in English may flash past in German or Finnish. Splitting or condensing is part of quality, not a shortcut.
  • Keep speaker identification consistent. If the original marks who is speaking, the translation should too, with names spelled as the audience expects.
  • Verify timing after translation. A translated cue should still appear while the matching speech is heard.

Language tags on caption tracks, files and pages

Software cannot guess reliably which language a piece of text is in. It relies on language tags, usually written as BCP 47 codes such as es, ar, pt-BR or zh-Hans.

Web page
The lang attribute on the html element declares the main language of the page
Mixed-language content
A lang attribute on a smaller element marks a passage in another language, such as a quoted phrase or a translated transcript embedded on an English page
HTML video captions
Each track element carries srclang for the language code and label for the name shown in the caption menu
Video file tracks
MP4 and MKV audio and subtitle tracks can carry a language code that players use to name and choose tracks
SRT files
No language field at all; the language has to be declared by the player, the upload form or the file name
Video platforms
Uploaded caption files are usually assigned a language in the upload dialog

WCAG 2 includes success criteria for this: Language of Page at level A and Language of Parts at level AA. Check the current text of WCAG for the exact wording and exceptions. The practical point is simple: if a Spanish transcript sits on an English page without a lang attribute, a screen reader will read the Spanish words with English pronunciation rules, which can make it close to unintelligible.

Screen readers and language switching

Screen readers choose a speech voice and pronunciation rules based on the declared language. Many switch automatically when they meet a lang attribute on part of a page, provided the user has a voice installed for that language. Several things break this:

  • Missing or wrong tags. A page declared as English but written in French is read with English pronunciation.
  • Captions and transcripts injected by a player script without any language marking.
  • Labels in the wrong language, such as a caption menu that lists English names for every language while the rest of the interface is translated.
  • Text drawn as an image, which a screen reader cannot read in any language.

Test each language version with at least one screen reader, listening to the page title, player controls, caption menu and transcript. If the voice mispronounces most words, the language tag is usually the cause.

Language switchers people can find and use

A language switcher is only accessible if every user can operate it and understand the choices.

  • Name each language in its own language: Español, Deutsch, العربية, 日本語. A Japanese reader looking for their language should not have to recognize the English word Japanese.
  • Do not use flags alone. Flags represent countries, many languages span several countries, and an icon without a text label means nothing to a screen reader.
  • Make the switcher reachable and operable by keyboard, with a visible focus indicator and a clear accessible name such as Language.
  • Mark each option with its own lang attribute so a screen reader pronounces it correctly.
  • Keep the switcher in the same place on every page and every language version.

When separate language versions of a video exist as separate files rather than switchable tracks, link them clearly from each page. Combining several audio languages into one file is covered in making a video with two audio tracks.

Right-to-left captions

Arabic and Hebrew captions add display problems that a left-to-right reviewer may never notice. Punctuation can jump to the wrong end of a line, numbers and Latin brand names embedded in a sentence can reorder, and some players or burned-in workflows do not handle bidirectional text well. Test right-to-left captions in the actual player and on the actual platform, with a reader of the language, before publishing. Fonts that lack the script, or files saved in the wrong encoding, show boxes or garbled text instead. Layout and punctuation details are covered in right-to-left video localization.

Translated transcripts and description

A transcript helps deafblind users who read with a braille display, people who prefer reading, and anyone searching for a passage. It should exist in each language you publish, marked with the right language tag and placed where it is easy to find next to the video.

Audio description is often forgotten in translation. If the original version has a described track or a descriptive transcript, each language version needs its own, translated and then checked against the picture. Descriptions written for one language may not fit the pauses in the translated speech, because the translated dialogue takes a different amount of time. What audio description provides, compared with subtitles, is set out in audio description vs subtitles.

Hypothetical example: a library tutorial in three languages

A city library publishes a five-minute tutorial on reserving books in English, Spanish and Arabic. Reviewing the Arabic page with a screen reader, a volunteer finds that the embedded transcript has no lang attribute, so it is read with English rules, and that the caption menu labels the tracks English, Spanish and Arabic in English. The team adds lang attributes, relabels the tracks as English, Español and العربية, asks a fluent volunteer to fix two misplaced question marks in the right-to-left captions, and adds the missing sound cues to the Spanish file.

Common mistakes and limits

  • Treating translated subtitles as the accessible version. Hearing-oriented subtitles usually omit sound cues and speaker identification that deaf viewers need.
  • Reviewing only the original language. Errors in translated captions are just as harmful to the people who depend on them.
  • Relying on automatic platform captions in some languages. Quality varies by language and audio, and nobody may notice when it is poor.
  • Forgetting interface text: player buttons, caption menu labels and error messages also need translating and tagging.
  • Expecting tags to fix everything. A screen reader can only switch to a language it has a voice for, and a user's player may ignore track labels. Tagging is necessary, not sufficient.
  • Assuming spoken-language captions serve every deaf viewer. Many Deaf people use a sign language as their first language, and captions in a written language are not a replacement for it. Sign language is discussed in sign language and AI translation.

Where mydubly fits in multilingual accessibility

mydubly can produce the raw material for several of these layers, but not the finished accessible versions. It detects the spoken language automatically and works with 21 languages. Chinese output is Simplified, Arabic is Modern Standard Arabic and Norwegian is Bokmål. A transcript job can include a translated transcript, giving you the text in both languages, and each job returns one subtitle track, as both SRT and VTT files, in either the spoken language or the target language, saved as UTF-8 with a byte order mark. Plan one job per subtitle language you need. For a 10-minute video in three languages, that is three transcript jobs of 10 credits each, 30 credits (3¢) in total.

What mydubly does not do matters just as much here. Its subtitle files contain recognized speech only: no sound tags, no speaker labels and no styling. It does not create audio description or descriptive transcripts, does not produce sign language, does not translate on-screen text, and does not set language tags on your web page or player. A dubbed video uses one stock voice and contains no subtitle track; the subtitles are separate files. Every file should be reviewed by a fluent reader and completed with non-speech information before it is used for accessibility. The video localization page describes the outputs, and the VTT subtitle generator page covers the web caption format.

Next step: audit one language version end to end

Choose your least-reviewed language version and go through it as a user who depends on it: switch to that language with a keyboard, turn on captions, read a minute of them against the audio with a fluent colleague, open the transcript with a screen reader and listen to how it is pronounced. Write down each gap, fix the tags and labels first because they are quick, and then schedule review of the caption text.

Frequently asked questions

What does multilingual accessibility mean for video?

It means each language version of a video is as accessible as the original. Captions, transcripts, description and the player interface should all exist and work in every language you publish, so a disabled viewer who speaks only one of those languages is not left with a weaker version.

Do I need separate captions and subtitles for each language?

Often, yes. Subtitles for hearing viewers translate the dialogue, while captions for deaf and hard-of-hearing viewers also include sound cues and speaker identification. Some publishers provide one caption-style file per language that serves both groups; whichever you choose, it needs review by a fluent reader.

How do I set the language of a caption track?

On a web page, the track element takes a srclang attribute with a language code and a label with the language name that appears in the menu. On video platforms, you choose the language when you upload the file. SRT files have no internal language field, so the language must always be declared wherever the file is attached.

Why does my screen reader mispronounce a translated transcript?

Usually because the transcript is not marked with its language. Screen readers pick pronunciation rules from the lang attribute, so Spanish text inside an English page without its own tag is read as if it were English. Adding a lang attribute to the element that contains the transcript normally fixes it, as long as the user has a voice for that language installed.

Should language switchers use flags?

Not on their own. Flags stand for countries rather than languages, many languages are spoken in several countries, and an image without a text label tells a screen reader user nothing. Name each language in its own language and give the option a matching lang attribute.

Does audio description need translating too?

Yes, if the original version has it. A blind viewer of the Spanish version needs Spanish description, and the descriptions may need rewriting to fit the pauses in the translated dialogue. A descriptive transcript in each language is an alternative or a complement.

Can mydubly create accessible captions in several languages?

It can draft the speech text: one subtitle track per job, delivered as SRT and VTT, in the spoken language or a target language, plus transcripts in both languages for translated jobs. The files contain recognized speech only, so you add sound cues and speaker identification and have a fluent reader review them before publishing them as accessibility captions.