How to use this checklist
Accessibility problems are cheapest to fix before a video is published, while the script, the edit and the caption file are all still open. Work through the layers in the order below, because later layers depend on earlier ones: the captions and the transcript come from the same text, and description of visuals depends on knowing where the pauses in speech fall.
Mark each item pass, fail or not applicable, and write down who checked it.
The items map loosely to the Web Content Accessibility Guidelines (WCAG), the W3C standard many organizations use as their benchmark. Which level applies to you, and whether a law requires it, depends on your country, your sector and which version of the standard is referenced, and these rules change. Check the current official source and treat this list as practical guidance, not legal advice. Schools and universities have their own expectations, covered in caption requirements for educational video.
Layer 1: captions
Captions carry everything a viewer who cannot hear needs to follow the video: the words, who is speaking when that is not obvious, and sounds that matter.
- Every meaningful spoken word is captioned, including off-screen narration and background speech that carries information.
- Speaker changes are clear when the speaker is off screen or several people are visible.
- Meaningful non-speech sounds are tagged, such as [door slams], [phone buzzing] or [tense music].
- Each caption appears with its speech and stays on screen long enough to read.
- Names, technical terms and numbers are checked against a reliable source such as the script or a glossary.
- Captions do not cover important on-screen information, such as a name in a lower third or code on a slide.
- You know whether the captions are closed (a separate track viewers can switch on or off) or open (burned into the picture), and why you chose that.
Automatic captions are a draft, not a finished accessibility feature: recognition errors land hardest on names, figures and negations, and a deaf viewer has no audio to fall back on. How AI captioning affects access goes into the benefits and risks in more depth.
Done means someone who did not write the captions watched the whole video with the sound off and could follow it.
Layer 2: the transcript
A transcript serves people who cannot read captions at playback speed, deafblind users reading on a braille display, and people who want to search or skim. For audio-only content, such as a podcast episode embedded on a page, the transcript is the main accessible alternative.
- The transcript includes all speech, speaker names and meaningful sounds.
- For videos where important information is only visual, the transcript also describes it.
- It sits next to the player or is linked directly beneath it with a clear label.
- It is real text on the page or in an accessible document, not an image or a scanned PDF.
- It matches the final edit. Any recut after the transcript was made means checking it again.
Page layout, collapsible transcript panels and formatting are covered in publishing transcripts alongside your videos.
Layer 3: description of visual information
Blind and low-vision viewers miss whatever is shown but not said: a chart, a demonstration step, text on a slide, who has just walked into frame, a facial expression that changes the meaning of a line.
- List every visual that carries information not present in the audio.
- Fix as much as possible in the script: presenters can say what they are showing, for example the revenue line dropping sharply in March, instead of saying look at this.
- Where narration cannot cover a visual, plan an audio description track or a described version of the video.
- If your player supports a separate description track, test that it actually plays; otherwise publish the described version alongside the standard one.
- Where pauses are too short for a needed description, consider extended description, which pauses the picture while the description plays.
Audio description is a different deliverable from captions and subtitles; the comparison is laid out in audio description versus subtitles.
Layer 4: the player
An accessible caption file is useless if the viewer cannot reach the caption button. Test the player you will actually publish with, on the page where it will live.
- Every control can be reached and operated with a keyboard alone: play and pause, seek, volume, captions, settings and fullscreen.
- Focus is visible and moves in a logical order, and the keyboard never gets trapped inside the player or an embedded frame.
- Controls have accessible names that a screen reader announces, such as Play or Captions on.
- The caption toggle is easy to find, and viewers can adjust caption size and colour where the player allows it.
- The video does not autoplay with sound. WCAG is commonly cited as asking for a way to pause or stop audio that plays automatically for more than three seconds, and for moving content that starts automatically and lasts more than five seconds.
- The controls remain usable when the page is zoomed in and on a phone screen.
- The transcript and any described version are reachable from the player area.
What to look for when choosing or evaluating a player is covered in detail in what makes a video player accessible.
Layer 5: flashing, motion and burned-in text
- Nothing flashes more than three times in any one-second period, the threshold WCAG uses for general and red flashes. Check strobe effects, camera flashes in event footage, glitch transitions and very fast cuts. Dedicated photosensitivity analysis tools exist if you publish a lot of high-energy footage.
- Text burned into the picture has strong contrast against its background. Many teams borrow WCAG's text contrast ratios, commonly cited as 4.5:1 for normal text and 3:1 for large text, and use a solid or semi-transparent backing box behind text over busy footage.
- On-screen text stays up long enough to be read comfortably, ideally twice.
- Important on-screen text is also spoken or described, because a screen reader cannot read text inside the video picture.
- Information is never conveyed by colour alone, such as a red or green status light with no label.
- Text and graphics avoid the bottom area where captions usually appear, or captions are positioned to avoid them.
Who checks what before release
- Scriptwriter or producer
- Visual information spoken aloud, plain wording, a glossary of names and terms
- Editor
- Flashing, on-screen text contrast and duration, space left for captions
- Caption reviewer
- Caption accuracy, timing, sound tags, speaker identification, a full watch with the sound off
- Web or platform owner
- Keyboard and screen-reader test of the player, caption and description tracks, language tags, transcript placement
- Final approver
- Spot check of every layer, a dated sign-off record, a route for viewer feedback
A three-person marketing team plans a 12-minute product update. The producer rewrites two lines so the presenter names the menu items she clicks instead of saying this one here. The editor swaps a strobing transition for a simple cut and adds a dark box behind the feature names. A caption draft is generated from the final export, and a colleague corrects three product names, adds [notification chime] where the sound matters, and watches the whole video muted. The web editor tabs through the embedded player, confirms the captions toggle is announced, and links the transcript under the video.
Common mistakes and the limits of a checklist
- Publishing automatic captions without anyone reading them.
- Captions that contain the words but drop the sounds and speaker changes that give them meaning.
- Treating the caption file as the transcript. Short cue fragments with timestamps are hard to read as a document.
- Testing the player only with a mouse.
- Assuming narration means no description is needed, when the narration says see here over a chart.
A checklist catches omissions; it cannot tell you whether the video is genuinely usable. Testing with disabled viewers, or at least with a screen reader and the sound off, finds problems a list never will. Some needs are outside what captions and transcripts can meet at all: Deaf viewers whose first language is a sign language may need an interpreter, and no checklist item turns a caption file into sign language.
Where mydubly fits in the checklist
mydubly covers the first draft of two layers, captions and the transcript, and nothing else on this list. You choose a video (MP4, MOV, WebM, MKV or M4V) or an audio file from your device, up to 2 hours long. The spoken language is detected automatically, and you get a plain-text transcript, a timestamped transcript with [m:ss] labels, and SRT and VTT subtitle files from the subtitle generator or the video to text tool.
Those files contain recognized speech only. Each cue is a single line covering one recognized segment, with no sound tags, no speaker labels, no styling and no positioning. The caption reviewer's job is therefore to correct errors, add the sound and speaker information, and adjust cue breaks where needed before the file counts as accessible captions. mydubly does not create audio description, check flashing or contrast, provide a player, burn captions into the picture or read on-screen text. If a file has several audio tracks, only the default track is used.
Transcription costs 1 credit per minute with a minimum of 5 credits, so a 12-minute video costs 12 credits (1.2¢). The video picture stays on your device; only compressed audio chunks are sent over HTTPS, and they are deleted within 30 minutes of the job finishing.
Next step: run one video through the list
Pick the next video due for release and walk it through the six layers, writing down pass, fail or not applicable for each item and who checked it. If captions are the first gap, the guide on adding subtitles to a video covers attaching a caption file on common platforms once it has been reviewed.
Frequently asked questions
What is the minimum a video needs to be accessible?
For most prerecorded video with speech, the baseline is accurate captions, a way for blind and low-vision viewers to get important visual information, and a player that works with a keyboard and screen reader. A transcript is strongly recommended and is the main alternative for audio-only content. Which items are legally required depends on where you are and who you publish for, so check the current official rules.
Are automatic captions good enough for accessibility?
Not on their own. Automatic captions are a useful starting point, but errors in names, numbers and short words can change the meaning, and most automatic systems do not add sound descriptions or speaker identification. Treat them as a draft that a person reviews against the audio before publishing.
Does my video need audio description if it already has narration?
It depends on whether the narration covers everything shown. If the presenter describes the chart, names the people on screen and reads out on-screen text, separate description may not be needed. If important information appears only in the picture, it needs to be described, either by rewriting the narration or with an audio description track or described version.
Can a transcript replace captions?
No. A transcript and captions serve different situations: captions let a viewer follow speech in sync with the picture, while a transcript lets people read, search or use a braille display at their own pace. Most accessible videos publish both, and both can be made from the same reviewed text.
How do I check a video for dangerous flashing?
Watch for strobe effects, camera flashes, rapid cuts between bright and dark shots and glitch transitions, and compare them with WCAG's guidance of no more than three flashes in any one-second period. When in doubt, replace the effect.
Who should sign off on video accessibility?
Ideally one named person who did not produce the video checks each layer and records the result, supported by the editor, a caption reviewer and the web owner. The record matters because it shows what was tested when a viewer reports a problem later.