Education & research

Captions and Transcripts That Make Course Videos Accessible

Educational video accessibility captions are synchronized text of everything spoken, plus meaningful sounds and speaker changes, so deaf and hard-of-hearing students get the same content as everyone else. Under WCAG 2, the guidelines most accessibility policies in education point to, prerecorded video with sound needs captions, prerecorded audio-only content needs a text alternative such as a transcript, and important visual information needs description. AI transcription gives you a fast first draft, but captions used for access need human review. This article explains the standard in plain terms and is not legal advice.

7 min read · Updated

What WCAG 2 says about prerecorded media

The Web Content Accessibility Guidelines are published by the W3C. Version 2.0 dates from 2008, 2.1 from 2018 and 2.2 from 2023, and the media requirements carried forward unchanged across those versions. Guideline 1.2, Time-based Media, holds the success criteria for audio and video. For prerecorded content, the ones that matter most in education are these:

1.2.1 Audio-only and Video-only (Prerecorded), Level A
Audio-only content, such as a recorded lecture podcast, needs a text alternative, which in practice means a transcript. Video with no sound needs a text alternative or an audio track describing it.
1.2.2 Captions (Prerecorded), Level A
Prerecorded audio in synchronized media, meaning video with sound, needs captions.
1.2.3 Audio Description or Media Alternative (Prerecorded), Level A
Important visual information needs either audio description or a full text alternative to the video.
1.2.5 Audio Description (Prerecorded), Level AA
Audio description is expected for prerecorded video; at this level a text alternative alone is no longer enough.
1.2.6 to 1.2.9, Level AAA
Sign language interpretation, extended audio description, a full media alternative and alternatives for live audio-only content.

Live captions, criterion 1.2.4 at Level AA, apply to live streams and sit outside a file-based workflow. Most institutional policies target Level AA, which includes all the Level A criteria above plus 1.2.4 and 1.2.5. The W3C's Understanding documents for each criterion explain the intent and the exceptions, such as media that is clearly labeled as an alternative to text already on the page, and are worth reading in full.

Why policies point to WCAG

WCAG is a technical standard, not a law, but laws and policies often reference it. In the United States, a 2024 Department of Justice rule under Title II of the Americans with Disabilities Act adopted WCAG 2.1 Level AA as the technical standard for the web content and mobile apps of state and local governments, which covers public schools, colleges and universities, with compliance dates phased by the size of the entity. In the European Union, the standard EN 301 549, used for public sector websites and apps, incorporates the WCAG 2.1 success criteria. Other countries have their own frameworks.

Rules, exceptions and deadlines change, and how they apply to a particular course or platform is a legal question. Check the current text with your institution's accessibility office or counsel rather than relying on any summary, including this one.

Captions are not the same as transcripts

A caption is synchronized: each piece of text appears while it is being spoken. A transcript is a document. WCAG treats them as alternatives for different kinds of media, so a transcript posted beneath a lecture video does not by itself satisfy the captions criterion for that video.

  • Captions carry all the speech, in the language spoken.
  • They identify the speaker when it is not obvious who is talking, for example an off-camera student question.
  • They include non-speech sounds that carry meaning, such as [laughter], [alarm sounds] or [music stops].
  • They are synchronized closely enough to follow and are readable, with lines short enough and on screen long enough.
  • They can be closed, so viewers switch them on, or open, burned into the picture. Either can serve access if the text is accurate.

A transcript for audio-only content should also identify speakers and include meaningful sounds. For video, a descriptive transcript that adds the important visual information is one way to provide the media alternative that criterion 1.2.3 allows at Level A.

How accurate do accessibility captions need to be?

WCAG does not set a numeric accuracy threshold. Its intent is that captions convey the same information as the audio. Other guidance fills in the practical detail: the FCC's caption quality standards for US television describe accuracy, synchronicity, completeness and placement, and the DCMP Captioning Key, widely used in education, gives detailed style rules. The working test is simple: a deaf student reading the captions should get the same content as a hearing student listening.

Raw automatic captions often fall short of that test. The problem is rarely that most words are wrong; it is that the errors land on the words that carry the most meaning in teaching: technical terms, names, numbers and negations.

One digit, a different answer

Spoken in a statistics lecture: "Reject the null hypothesis when p is below 0.05." Suppose the draft reads: "Reject the null hypothesis when p is below 0.5." The line is fluent, grammatical and wrong by a factor of ten, and a reviewer skimming text without listening would never notice.

A review workflow for accessible captions

  1. Generate a draft SRT or VTT file from the final edit of the video, so the timing matches what students will see.
  2. Play the audio and read the draft at normal speed, with the slides or script open.
  3. Fix technical terms, names, numbers and negations first; they carry the most meaning per character.
  4. Add speaker identification wherever the speaker changes and is not visible on screen.
  5. Add meaningful non-speech information in brackets, such as [demonstration plays without narration] or [audience laughs].
  6. Check readability: split any cue too long to read in the time it is on screen.
  7. Watch the finished captions once in the player students actually use, with the sound off.
  8. Publish the caption file; for audio-only content, publish the corrected transcript with a clear link next to the player.

Plan time realistically. A clean single-speaker lecture reviews quickly; a dense lab demonstration with jargon and crosstalk can take several times its running length. Instructors are often the fastest reviewers of their own terminology, while an accessibility team is better at speaker labels, sound cues and consistency.

Where AI captioning helps

  • It turns captioning from a typing job into an editing job, which is usually far faster.
  • Timing comes with the draft: cues are already synchronized, so reviewers adjust rather than create.
  • It makes captioning a back catalogue affordable. Suppose a department has 300 hours of course video: that is 18,000 minutes, so drafts cost 18,000 credits, or $18. The review is where the real budget goes.
  • The same files serve other purposes, such as searchable archives and study notes, and translated subtitles for international students are a separate service covered in captions vs subtitles.

Limitations of automatic captions for accessibility

  • Speech recognition does not know every field's vocabulary, and specialist terms are the main source of errors.
  • mydubly's transcripts do not label speakers, and automatic transcription focuses on speech, so expect to add speaker identification and sound cues yourself.
  • Captions do not cover visual information. An instructor pointing at a diagram and saying "this goes here" needs audio description or a descriptive transcript under criteria 1.2.3 and 1.2.5.
  • Text that appears on screen is not captioned unless someone reads it aloud.
  • Mathematics and code spoken aloud become words, so provide the notation in accompanying materials.
  • An unreviewed draft published as an accommodation risks failing the very students it is meant to serve.

Making caption files with mydubly

mydubly's transcript mode produces SRT and VTT caption files plus a timestamped transcript from a video or audio file on your device, at 1 credit per minute with a 5-credit minimum per file. Speech recognition runs on OpenAI's open-source Whisper model, the spoken language is detected automatically, and files of up to 2 hours in any of 21 languages are supported. The captions are delivered as closed caption files to upload to your player or learning platform; nothing is burned into the picture. If you are unsure which format your platform wants, see SRT vs VTT.

Both caption formats are plain text, so you can review them in any subtitle editor or text editor. The video file stays on your device, and only the audio is uploaded, then deleted within 30 minutes of the job finishing. Remember that speaker identification and sound cues are review tasks. The subtitle generator is the starting point.

A practical next step

Pick the course video students watch most, generate a draft with the subtitle generator and time how long a full review takes. That number, multiplied across your library, is your real captioning budget, and it tells you whether to spread review across instructors or fund it centrally. For lecture recordings you also want to make searchable, see searchable lecture transcripts.

Frequently asked questions

Are automatic captions from a video platform enough to meet WCAG?

WCAG asks for captions that convey the content of the audio, and unreviewed automatic captions often contain errors on exactly the terms that matter. Most accessibility guidance treats them as a starting point to correct, not finished captions. Check your institution's policy for its specific expectation.

Does a transcript alone make a lecture video accessible under WCAG?

Not for the captions criterion. Prerecorded video with sound needs synchronized captions under 1.2.2. A transcript, ideally a descriptive one, can serve as the media alternative under 1.2.3 at Level A, but it does not replace captions.

Do audio-only lectures and podcasts need captions?

Criterion 1.2.1 asks for a text alternative for prerecorded audio-only content, which in practice is a transcript that identifies speakers and notes meaningful sounds. Publish it next to the audio player with a clear link.

Should accessibility captions include filler words like um and uh?

Caption style guides generally let you omit fillers that carry no meaning while keeping everything substantive, and automatic transcription tends to drop many of them anyway. Keep hesitations that matter, for example in a recorded role-play. See verbatim vs clean verbatim transcription for the trade-offs.

Can translated subtitles count as accessibility captions?

Captions for deaf and hard-of-hearing viewers represent the audio in the language spoken. Translated subtitles serve viewers who do not understand that language, which is a different need. Provide accurate same-language captions first, then add translations if your audience needs them.