Industry & future

AI and Sign Language: What Machines Can and Can't Translate Yet

Automatic sign language translation is an active research field, not a solved problem. Sign languages are complete languages with their own grammar, expressed through hands, face, body and space, and recognizing or generating them reliably remains much harder than speech recognition. For Deaf viewers who prefer sign language, captions and voice dubbing are not equivalent, so qualified human interpreters remain the dependable option today.

7 min read · Updated

Sign languages are languages, not codes for speech

A sign language is a natural language with its own vocabulary, grammar and regional variation, developed within Deaf communities. It is not a signed version of the surrounding spoken language. American Sign Language and British Sign Language, for example, are distinct languages and largely unintelligible to each other, even though both countries speak English. ASL is historically related to French Sign Language instead.

Grammar in sign languages uses dimensions speech doesn't. Signers place people and objects in the space around them and refer back to those locations. Facial expressions and head movements carry grammatical meaning, such as marking questions or conditions, not just emotion. Handshape, movement, location and orientation all change meaning. Many countries legally recognize their national sign languages, and there are many distinct sign languages worldwide rather than one universal one.

Manually coded systems that follow spoken-language word order also exist, mainly in education. They are not the same as natural sign languages, and confusing the two leads to poor design decisions.

Why captions and dubbing don't serve every Deaf viewer

Captions turn speech into written text in the spoken language. For many Deaf people whose first language is a sign language, that written language is a second language, and reading fast captions in it is not the same as receiving information in their own language. Captions still matter enormously, and many deaf and hard-of-hearing people rely on them; the point is that they are not a universal replacement for sign language.

Voice dubbing, which translates speech into speech in another spoken language, does nothing for Deaf viewers at all. Neither does a translated transcript if the reader's strongest language is signed. The broader picture of which AI tools help which viewers is in AI and video accessibility, and the caption-specific distinctions are in captions vs subtitles.

Where automatic sign language recognition stands

Recognizing sign language from video, and translating it into written or spoken language, has been studied for decades and has improved with modern machine learning. It remains far behind speech recognition, for reasons that are well understood.

Data
Sign language video datasets are much smaller than speech datasets, and annotating them requires fluent signers.
Writing
Sign languages have no widely used written form, so systems often rely on glosses, rough word labels that lose grammar carried by face and space.
Visual detail
Handshapes, fingers, facial grammar and movement in three dimensions must all be captured from ordinary video, with occlusion and blur.
Variation
Signing varies by signer, region, age and context, as speech does, and models trained on a few signers generalize poorly.
Continuous signing
Isolated signs are easier to recognize than natural conversation, where signs blend and grammar spans space and time.

Established: research systems can recognize isolated signs and constrained phrases with reasonable accuracy in controlled conditions. Open: reliable translation of natural, continuous signing by many signers in everyday video. Treat claims of general-purpose sign language translation with skepticism and ask how they were evaluated, on which language and with which signers.

Signing avatars and generated sign

The reverse direction, generating sign language from text or speech, is often demonstrated with animated signing avatars. Avatars can display signs, and some projects have produced short, scripted announcements this way. Producing natural, grammatical signing with correct facial expressions, use of space and fluid transitions is much harder, and many avatars are criticized by Deaf viewers as stiff, unclear or missing the grammar that faces carry.

Newer research uses generative video models to synthesize signing from video of real signers. This is early-stage, and it raises the same consent and likeness questions as voice cloning, along with the risk that output looks fluent while being wrong.

What Deaf communities have raised

Organizations representing Deaf people have published statements urging caution about signing avatars and automated translation. Common themes include:

  • Deaf people should be involved in designing, testing and approving any sign language technology, not consulted afterward.
  • Automated systems should not replace qualified interpreters where accuracy and rights are at stake, such as in health, legal, emergency and education settings.
  • Poor-quality signing can be worse than none, because it signals access while delivering confusion.
  • Funding spent on unproven technology can come at the expense of interpreters and Deaf-led services.

These concerns are not hostility to technology. They reflect experience with tools that were built without the people they were meant to serve.

Adding sign language to a video today

The reliable way to include sign language in video is a qualified interpreter, ideally one with experience interpreting recorded media, and for some content a Deaf interpreter working alongside a hearing one.

Example: a 15-minute public information video

A city publishes a video about changes to waste collection. It books a qualified sign language interpreter, sends a transcript of the narration a few days ahead so the interpreter can prepare local street names and terms, and films the interpreter against a plain background. The interpreter appears in a large picture-in-picture box for the whole video, and the city also publishes captions and a text summary.

  1. Decide which sign language your audience uses; it follows country and community, not spoken language.
  2. Book a qualified interpreter early, and ask local Deaf organizations for recommendations.
  3. Share the script or transcript and any names, places and terms in advance.
  4. Film the interpreter with good lighting, a plain contrasting background and the full signing space, including the face, in frame.
  5. Place the interpreter large enough to read on a phone, and keep them on screen for the whole video.
  6. Add captions as well, since many deaf and hard-of-hearing viewers don't sign.
  7. Ask Deaf viewers for feedback before and after publishing.

Public-sector content raises its own considerations, discussed in public information video translation.

Mistakes hearing producers make

  • Assuming captions make a video fully accessible to all Deaf viewers.
  • Using an avatar or automated tool as the only sign language provision for important information.
  • Shrinking the interpreter to a tiny corner box, or cropping out the face.
  • Choosing the wrong sign language for the audience, such as ASL for a British audience.
  • Treating sign language as a word-for-word translation of the script, rather than letting the interpreter restructure meaning.
  • Not involving Deaf people in decisions about their own access.

What mydubly can and can't contribute

mydubly works with audio only. It recognizes speech, translates it and generates subtitles or a synthetic voice track; it does not recognize, translate or produce sign language, and it does not analyze the picture.

Its useful role is supporting the human workflow. The subtitle generator can produce the captions that should accompany an interpreted video, as SRT or VTT files, and the video to text tool can give an interpreter a timestamped transcript to prepare from. For the 15-minute city video, that transcript and subtitles cost 15 credits (1.5¢). Captions from mydubly have no speaker labels or sound cues, so they need review before they serve as captions for deaf viewers. For guidance on accessibility rules in schools and universities, see captions for educational video accessibility.

Next steps for your own videos

If your audience includes Deaf signers, plan for a qualified interpreter and captions, and talk to local Deaf organizations about what works for them. Follow sign language AI research with interest, but judge it by evaluations involving Deaf signers rather than demos. For the caption side of an accessible video, start with the subtitle generator.

Frequently asked questions

Is there one universal sign language?

No. Different countries and communities have their own sign languages, which developed independently and have their own grammar and vocabulary. Some international signing is used at global Deaf events, but it is not a replacement for national sign languages. Always confirm which sign language your audience uses before booking an interpreter.

Can my phone translate sign language into speech?

Not reliably for real conversations today. Some apps and research prototypes recognize individual signs or fingerspelling in controlled conditions, but natural, continuous signing by different people remains beyond current consumer technology. For important communication, use a qualified interpreter or written communication agreed with the Deaf person.

What is a Deaf interpreter?

A Deaf interpreter is a Deaf person, fluent in sign language, who works alongside or instead of a hearing interpreter. They may interpret from a written script or from a hearing interpreter's signing into a form that suits the audience, and they bring native fluency and cultural knowledge. Interpreter organizations can explain when one is recommended.

Should an interpreted video still have captions?

Yes. Many deaf and hard-of-hearing people don't use sign language, including many who lost hearing later in life, and captions also help hearing viewers watching without sound. Interpreting and captions serve overlapping but different audiences, so providing both makes a video usable by more people.

Are signing avatars acceptable for short announcements?

Opinions within Deaf communities vary, and acceptability depends on quality, context and whether Deaf people were involved. Many organizations caution against avatars for anything important or complex. If you consider one, test it with Deaf viewers from your audience and keep a human interpreter for content where accuracy matters.