Accessibility & inclusive video

WCAG Caption Requirements for Video and Audio, Explained

The Web Content Accessibility Guidelines (WCAG) have a group of success criteria for time-based media. In short: prerecorded video with sound needs captions (Level A), live video needs captions (Level AA), audio-only recordings need a text alternative such as a transcript (Level A), and video needs audio description or a full text alternative (Level A, with audio description required at AA). Automatic subtitles are a starting point, not a finished caption file.

4 min read · Updated

The success criteria that apply

WCAG 2.0 introduced these criteria, and WCAG 2.1 and 2.2 kept them:

1.2.1 Audio-only and Video-only (Prerecorded), Level A
Audio-only content needs a text alternative, such as a transcript. Video-only content needs a text or audio alternative describing it.
1.2.2 Captions (Prerecorded), Level A
Prerecorded video with audio needs captions, unless it's itself a clearly labelled media alternative for text.
1.2.3 Audio Description or Media Alternative (Prerecorded), Level A
Video needs audio description or a full text alternative of what's shown and heard.
1.2.4 Captions (Live), Level AA
Live video with audio needs captions.
1.2.5 Audio Description (Prerecorded), Level AA
Prerecorded video needs audio description.

Most organisations aim for Level AA, which includes everything at Level A. Higher Level AAA criteria add sign language, extended audio description and full media alternatives.

What counts as a caption

WCAG's understanding documents describe captions as covering not only dialogue but also who is speaking and meaningful non-speech sounds: [laughter], [phone rings], [music stops]. Captions must be synchronised with the media. This is the difference between captions and subtitles: subtitles assume the viewer can hear the sound and only need the words. Captions vs subtitles explains the distinction, and captioning sound effects and music covers how to describe sounds well.

Transcripts for audio-only content

Podcasts, recorded talks without video and audio messages on a website need a text alternative. A transcript that includes speaker names and relevant sounds meets the intent. Publish it near the audio, as text, not as an image. Publishing transcripts on your website covers placement, and descriptive transcripts covers the richer form that also describes visuals.

Audio description

Audio description narrates important visual information that isn't conveyed by the soundtrack: a chart on a slide, text on screen, an action. If the speaker already says what's shown ("as this chart shows, sales doubled"), less description is needed. Writing scripts that describe visuals as you go reduces the work later; how to write audio description and subtitles vs audio description go further.

Example: a university lecture library

An accessibility lead audits 200 lecture recordings. Automatic subtitles exist for all of them. She sets a process: each file is proofread, lecturer names are added at the start, sounds like [audience laughs] or [video plays without speech] are added, and slides with unread content get a short description in the transcript. The automatic files cut the work to a fraction, but the human pass is what makes them captions.

Do automatic captions meet WCAG?

WCAG doesn't ban automatic captions, but it expects the result to be accurate and complete. Automatic speech recognition makes errors with names, technical terms, accents and overlapping speech, and typically doesn't identify speakers or describe sounds. Uncorrected automatic captions therefore often fall short. The practical approach is to generate, then edit: correct errors, add speaker identification where it isn't obvious, add meaningful sounds, and check synchronisation.

Where mydubly fits

The subtitle generator produces the first draft: SRT and VTT files synchronised to the speech, plus timestamped and plain transcripts, for 1 credit per minute. That covers the words and timing. You'll need to add speaker names and sound descriptions, and proofread, to reach captions. mydubly doesn't provide audio description, sign language or live captions; it works on recorded files only.

Limits of this summary

  • This is an overview, not legal advice. Laws and procurement rules reference WCAG in different ways and versions; for example, the US Section 508 standards incorporate WCAG 2.0 Level A and AA, and the European standard EN 301 549 aligns with WCAG 2.1 AA. Check what applies to you.
  • Exceptions exist, such as media that is itself an alternative to text already on the page, clearly labelled.
  • Accessibility goes beyond captions: player controls, keyboard access and contrast matter too; an accessible video player and the video accessibility checklist cover them.

A practical captioning process

For teams with many videos, a repeatable process matters more than any single tool:

  1. Generate a draft from the final export with the SRT generator or another tool.
  2. Proofread against the audio: names, terms, numbers, then meaning. How to proofread an AI transcript gives a fast order.
  3. Add speaker identification where the speaker isn't obvious on screen, and meaningful sounds in square brackets.
  4. Check that no line runs too long or too fast to read.
  5. Publish as a caption track labelled captions, and publish a transcript where audio-only content needs one.
  6. Record who checked each file and when, so audits are quick.

For educational content, captions for educational videos adds subject-specific advice, such as reading out equations and code. Organisations translating video can apply the same process per language; translated captions need a fluent reviewer as well as the steps above.

Frequently asked questions

Does WCAG require captions?

Yes. Success criterion 1.2.2 requires captions for prerecorded video with audio at Level A, and 1.2.4 requires captions for live video at Level AA.

Do podcasts need transcripts under WCAG?

Audio-only content needs a text alternative under 1.2.1, which a transcript provides.

Is audio description required by WCAG?

At Level A, video needs audio description or a full text alternative (1.2.3). At Level AA, prerecorded video needs audio description (1.2.5).

Are automatic captions WCAG compliant?

Not automatically. Captions need to be accurate and include speaker identification and meaningful sounds, so automatic files usually need editing.

What's the difference between captions and subtitles for WCAG?

Captions include speaker identification and non-speech sounds for viewers who can't hear the audio. Subtitles cover speech only.

Which WCAG level should I aim for?

Most organisations and regulations target Level AA, which includes all Level A criteria.