What a descriptive transcript is
An ordinary transcript records what was said. That works for a podcast, but in a video much of the meaning can sit in the picture: a presenter points at a diagram and says this part here, a slide lists the three deadlines, a chef shows how thick to slice an onion without saying it. Someone reading only the speech misses all of that.
A descriptive transcript closes the gap. It interleaves the dialogue with short descriptions of the visual content that matters, plus the non-speech sounds a caption would include. Read from start to finish, it should give the same information as watching and listening to the video. It is sometimes called a full text alternative or a media alternative.
Who relies on descriptive transcripts
- People who are deafblind, who often read text through a refreshable braille display and can access neither the soundtrack nor the picture directly.
- Blind and low-vision people who prefer reading with a screen reader to listening to a video, or who need to search or revisit a passage.
- Deaf people who want a single document containing the speech and the context, for example to study from.
- People with slow connections, data limits or devices that cannot play video.
- Anyone who wants to skim, quote or translate the content of a video.
WCAG 2 recognizes a full text alternative as one way to meet a level A requirement for prerecorded video with sound, while its level AA criteria call for audio description. Check the current success criteria and any rules that apply to your organization for exact obligations; this is a general description, not legal advice.
Plain, caption-style and descriptive transcripts compared
- Plain transcript
- The words spoken, possibly with timestamps. Useful for search, notes and quotes
- Caption-style transcript
- Speech plus speaker names and meaningful sounds such as [phone rings] or [laughter]. Close to what accessible captions contain
- Descriptive transcript
- Everything in a caption-style transcript plus descriptions of important visual information, so the text stands in for the whole video
The difference matters most for videos where the picture carries information: demonstrations, product walkthroughs, lectures with slides, explainer animations, drama and anything with on-screen text. For a talking-head video in which the speaker says everything, a caption-style transcript plus a brief description of the setting may be enough. How transcripts and timed captions differ in general is covered in transcript vs subtitles.
What visual information to include
Describe what a viewer needs in order to understand the content, not everything visible. A useful test is to ask: if someone missed this, would they misunderstand or miss something the video is trying to say?
- On-screen text: titles, names and roles in lower thirds, slide content, labels, captions in a chart, phone numbers and web addresses.
- Actions that carry meaning: a demonstration step, a gesture that replaces words, someone leaving the room, a product being assembled.
- Data and diagrams: the key point of a chart or the structure of a diagram, not every number unless the numbers matter.
- People and settings when they matter: who appears, where the scene takes place, a change of location.
- Expressions and reactions that change the meaning of speech, such as an eye roll that turns a statement into sarcasm.
- Visual jokes and reveals, described so they still land.
Usually leave out decorative backgrounds, routine camera movements, clothing and appearance unless relevant to the content, and anything already stated in the speech.
Writing descriptions that read well
- Keep descriptions short and factual. Say what is there or what happens, in the present tense.
- Describe, do not interpret. A man frowns and shakes his head is more useful than a man disagrees, unless the disagreement is unmistakable from context.
- Mark descriptions clearly so readers can tell them from speech, for example by wrapping them in square brackets or starting them with a label such as Description.
- Put each description where it happens in the sequence, usually just before or after the line of speech it relates to.
- Identify speakers by name once they have been introduced, and describe how they are introduced if their name only appears on screen.
- Write out on-screen text exactly when the wording matters, and summarize it when it does not.
- Use plain words. A descriptive transcript is read in sequence, often by braille or synthesized speech, so long, nested sentences are tiring.
Building one from a timestamped speech transcript
Writing descriptions from scratch while also typing out the speech is slow. Splitting the work is faster.
- Get an accurate timestamped transcript of the speech. A machine transcript is a reasonable start, but proofread it against the audio first, especially names, numbers and terms.
- Add speaker names. Mark who says each passage, since many machine transcripts do not label speakers.
- Watch the video with the transcript open. Each time the picture shows something a listener would miss, pause and note the timestamp and a rough description.
- Add meaningful sounds such as alarms, music that sets a mood, or laughter, in the same bracket style.
- Turn the notes into finished descriptions and place them in the transcript at the right points. Remove the timestamps afterwards if they clutter the text, or keep them so readers can jump to the moment in the video.
- Read the result with the video hidden. If any passage is confusing, something visual is still missing.
- Ask someone who uses a screen reader or braille display to try it, if you can.
If the video already has an audio description script, start your visual notes from it, though descriptions written to fit pauses may need expanding in a text version where time is not limited. Proofreading machine output is covered in how to proofread an AI transcript, and general layout choices in how to format a transcript.
A short sample
The plain transcript reads: Take your onion, cut it like this, then like this, and you're done. That is meaningless without the picture. The descriptive version below adds who is speaking, what is shown and how the action is done.
- [Description: A kitchen counter. A cook named Priya, introduced by on-screen text as Priya Shah, Head Chef, holds a halved onion, flat side down.]
- Priya: Take your onion and cut it like this.
- [Description: She makes vertical cuts toward the root, about a finger's width apart, without cutting through the root end.]
- Priya: Then like this.
- [Description: She turns the knife and slices across the first cuts, producing small, even dice.]
- Priya: And you're done.
- [Upbeat music. On-screen text: Recipe at the link below.]
Limits and mistakes to avoid
- Describing too much. Listing every object in the frame buries the information that matters.
- Describing too little. Leaving out on-screen text is the most common gap, especially names, slide bullets and links.
- Interpreting emotions or intentions that are not clear from the picture.
- Letting the transcript drift from the video after edits. If the video is re-cut, the descriptive transcript needs updating too.
- Hiding the transcript. A descriptive transcript that is hard to find, or offered only as an image or an inaccessible PDF, does not help. Where and how to publish it on a page is covered in publishing transcripts on your website.
- Treating it as a replacement for everything. A descriptive transcript does not replace captions for deaf viewers who want to watch, and some blind viewers prefer audio description. Many organizations provide all three.
Where mydubly fits
mydubly handles the first step, the speech. Upload the video and you get a plain transcript and a timestamped transcript with [m:ss] labels, which gives you an anchor for every visual note you add. You also get SRT and VTT subtitle files. A 12-minute video costs 12 credits (1.2¢), and a translated transcript can be included at the same rate if you need the text in another of the 21 supported languages.
The rest is human work. mydubly's output contains recognized speech only: no speaker labels, no sound descriptions and no visual descriptions. It does not analyze the picture at all; the video picture stays on your device and only the audio is processed. Names, roles and on-screen text all have to be added by someone watching the video, and the speech text should be checked against the audio before you build on it. The video to text page and the timestamped transcript format page describe the transcript outputs.
Next step: write one descriptive transcript
Choose a short video where the picture carries real information, such as a demonstration or a slide-based explainer under five minutes. Get a timestamped transcript, proofread it, add speaker names, then watch with the transcript open and add a note at every moment where a listener would miss something. Read it back with the screen off. Once one is done, the method becomes routine, and you can apply it to the videos people use most.
Frequently asked questions
What is the difference between a transcript and a descriptive transcript?
A transcript records the speech, sometimes with speaker names and sounds. A descriptive transcript also describes the important visual information, such as actions, on-screen text and charts, so that someone reading it gets the same information as someone watching and listening.
Who needs a descriptive transcript?
It is especially important for people who are deafblind and read with a braille display, since they cannot use captions or audio description. Blind and low-vision people who prefer reading, people who cannot play video, and anyone who wants to search or study the content also benefit.
Is a descriptive transcript the same as audio description?
No. Audio description is spoken narration added to the soundtrack, fitted into pauses in the dialogue. A descriptive transcript is a text document read at the reader's own pace. They cover similar visual information, and an audio description script is a good starting point for a descriptive transcript.
Should a descriptive transcript include timestamps?
It can, but it does not have to. Timestamps help readers jump to a moment in the video and help you keep the text in step with edits. Some publishers remove them from the final version because they interrupt reading, especially with braille or synthesized speech.
How detailed should visual descriptions be?
Detailed enough that the reader would not misunderstand or miss anything the video is trying to communicate, and no more. Prioritize on-screen text, meaningful actions, data and changes of scene. Decorative details, routine camera moves and anything already said aloud can usually be left out.
Can mydubly write a descriptive transcript automatically?
No. mydubly produces the speech part: a plain transcript and a timestamped transcript of what was said. It does not describe visual content, identify speakers or add sound descriptions, so a person watching the video needs to add those using the timestamps as anchors.