Accessibility & inclusive video

Writing an Audio Description Script, Step by Step

To write audio description, watch the video for visual information a blind viewer would miss, rank it by importance, and write short present-tense descriptions that fit into the natural pauses between dialogue without talking over speech. Where the pauses are too short for essential information, use extended description, which pauses the video, or build the description into the narration itself. A timestamped transcript makes the pauses easy to find before you start writing.

8 min read · Updated

What an audio description script does

Audio description is narration that tells blind and low-vision viewers what is happening on screen: actions, people, settings, on-screen text and anything else that carries meaning but is not spoken. It is usually delivered in the gaps between dialogue, so the original soundtrack and the description together give the full picture. How description compares with captions and subtitles, and who each one serves, is covered in audio description vs subtitles.

The script is the written plan for that narration. Each entry has a start time, sometimes an end time, and the words to be spoken. Writing it is a skill of selection: there is always more to describe than time allows, so most of the work is deciding what matters and saying it in as few words as possible.

What to describe, in priority order

Ask of every visual moment: would a viewer who cannot see this misunderstand, miss information or lose the thread? Then describe in roughly this order.

  1. Information the viewer needs to follow the content: on-screen text such as names, titles, slide points, phone numbers and instructions; demonstrations and steps; data shown in a chart.
  2. Who is present and who is speaking, especially when a new person appears or the speaker is not obvious from the voice.
  3. Actions that move things forward and are not obvious from the sound: someone hands over a document, a character leaves, a screen changes.
  4. Where and when: a change of location or time that the dialogue does not mention.
  5. Expressions, gestures and reactions that change the meaning of what is said.
  6. Atmosphere and visual detail, if time remains: lighting, costume, scenery that sets tone.

Do not describe what can be heard. If a door slams loudly, the viewer knows. If a person introduces themselves by name, there is no need to repeat it, though their role from a lower third may still need reading.

Finding the gaps: timing description to natural pauses

Standard audio description sits in pauses in the soundtrack, so the first practical step is to map where those pauses are and how long each one lasts.

  • Never talk over dialogue. Covering speech removes information blind viewers rely on most.
  • Avoid covering important sounds and music cues where you can, though music under a scene is often an acceptable place for description.
  • Place a description close to the moment it describes. Describing slightly before an action can help a viewer understand the sound that follows, but describing a reveal too early spoils it.
  • Write to the gap. Read each line aloud at a natural pace with a timer, then cut words until it fits with a little air at each end.
  • Leave some silences undescribed. Constant narration is tiring, and pauses can carry mood.

Using a timestamped transcript to map the gaps

A transcript with timings turns gap-hunting from repeated scrubbing into a list.

  1. Get a transcript of the dialogue with timestamps, ideally as a subtitle file such as SRT, which records a start and end time for each cue down to milliseconds.
  2. List each cue's end time next to the next cue's start time. The difference is the gap available for description.
  3. Mark every gap longer than a couple of seconds as a candidate slot. Short gaps can hold a name or a two-word action; long ones can hold a sentence or two.
  4. Watch the video with the list open and note what happens visually in each slot and in the stretches of speech around it.
  5. Check each candidate gap by listening. Recognition timestamps are approximate, and a cue's timing can extend into silence or clip the start of a word.
  6. Write descriptions for the most important visual information first, assigning each to the nearest suitable gap.

How recognized speech becomes cue timings, and why those timings drift, is explained in subtitle timestamp alignment.

Writing style: tense, tone and word choice

  • Use the present tense and the third person: She opens the laptop, not She opened the laptop.
  • Be objective. Describe what is visible rather than what you assume: He clenches his fist rather than He is furious, unless the emotion is unmistakable.
  • Choose precise verbs. Strides, shuffles or limps each say more than walks.
  • Name people once the video has introduced them; before that, use a short, neutral description.
  • Read on-screen text clearly, introduced with a cue such as A caption reads or On screen.
  • Match the tone of the content without performing it. Description for a children's cartoon can be livelier than description for a corporate policy video, but the describer should not compete with the cast.
  • Keep sentences short. Description is heard once, at speed, while other audio continues.

When the gaps are too short: extended and integrated description

Some videos have almost no pauses. Talking-head explainers, fast lectures and dense tutorials can leave no room for essential visual information. Two approaches help.

Extended description pauses the video while the description plays, then resumes. It allows full descriptions at the cost of a longer viewing time, and it needs a player or a separate described version that supports it. WCAG treats extended audio description as a level AAA criterion for prerecorded video, while standard audio description appears at level AA; check the current text for the exact requirements.

Integrated description, sometimes called inclusive or built-in description, writes the visual information into the main narration from the start. A presenter who says As you can see on this chart, sales doubled between March and June needs no separate description for that chart. For new productions this is usually the simplest route, and it also helps sighted viewers who look away. Turning finished description into text is the subject of descriptive transcripts, which deafblind viewers in particular may need.

A sample script excerpt

Hypothetical example: an onboarding video

A three-minute welcome video for new staff opens with a manager speaking to camera, cuts to the building entrance, shows a door code on screen and ends with a team photo. The dialogue never mentions the entrance or the code. Using the gaps found in the transcript, the writer adds four short descriptions.

0:00 to 0:03
An office reception area. Lena Park, Operations Manager, stands beside a front desk.
0:41 to 0:45
Outside, a glass door marked Staff Entrance at the side of the building.
0:46 to 0:50
On screen: Door code 4 7 1 9. Change it on your first day.
2:52 to 2:58
A group photo of about twenty people on outdoor steps, smiling and waving.

Each line names what matters, reads on-screen text exactly, and fits inside a measured pause.

Recording or synthesizing the description

Once the script is final, it needs a voice.

  • A human describer brings natural pacing and tone, and professional describers are trained in exactly this kind of delivery.
  • Synthetic speech is quicker to revise and can be acceptable for informational content, though listeners may find it less engaging over long programs. Pronunciation of names and terms needs checking line by line. How synthetic voices are produced is covered in what is text to speech.
  • Choose a voice clearly distinguishable from the main speakers, so listeners can tell description from dialogue.
  • Mix carefully. Lower the program audio slightly under each description so the words are clear, then bring it back up.
  • Decide how viewers will get it: a separate described version of the video, a selectable described audio track where the platform supports one, or text descriptions that some players read aloud. Player and platform support varies, so check yours before you produce.

Common mistakes and limits

  • Talking over dialogue to fit everything in.
  • Describing what is already audible, or what the speaker has just said.
  • Interpreting characters' thoughts or motives instead of describing actions.
  • Forgetting on-screen text, which is often the most important visual information in instructional video.
  • Writing the script from the edit decision list or memory rather than watching the final cut.
  • Skipping review. Ask a blind or low-vision viewer to watch the described version if at all possible; writers who can see routinely overestimate what is clear from sound.
  • Expecting standard description to work for every video. Dense content may need extended description or a rewrite of the narration.

What mydubly can and cannot do for audio description

mydubly does not create audio description. It does not analyze the picture, which never leaves your device, and it has no way to take a written script and voice it, because its inputs are audio and video files only. It also cannot add a described track alongside the original audio.

What it can do is produce the timing map. Upload the final cut and you get a timestamped transcript with [m:ss] labels plus SRT and VTT files whose cue times show where speech starts and stops, which is the raw material for the gap list above. A six-minute video costs 6 credits (0.6¢). The transcript contains recognized speech only, with no speaker labels or sound tags, and its timings should be checked by ear before you rely on a gap. One caution: if you upload a version with description already mixed in and dub it, the description would be treated like any other speech and voiced in the same single voice as the dialogue, so listeners could not tell them apart. Translating the description script separately keeps it distinct. The video to text page and the SRT generator page describe these outputs.

Next step: describe one short video

Pick a video under five minutes that relies on its picture. Get a timestamped transcript or SRT, list the gaps, and watch once noting only on-screen text and essential actions. Write those descriptions first, time each against its gap, and add lower-priority detail only where room remains. Then listen to the described version with your eyes closed and fix anything you cannot follow.

Frequently asked questions

What should audio description include?

Visual information a blind viewer would otherwise miss, in order of importance: on-screen text and instructions, who is present and speaking, actions that matter, changes of place or time, and meaningful expressions. Skip anything already clear from the dialogue or sound effects, and leave atmosphere for when time allows.

What tense should audio description be written in?

Present tense and third person, describing events as they happen: She picks up the phone. Keep sentences short, use precise verbs, and describe what is visible rather than interpreting thoughts or motives.

What is extended audio description?

Extended description pauses the video so a longer description can play, then resumes. It is used when the natural gaps in the dialogue are too short for essential information. It needs a player or a separate described version that supports pausing, so check your platform before relying on it.

Can I use text to speech for audio description?

Many organizations do, particularly for informational videos, because synthetic voices are fast to produce and easy to revise. Check pronunciation of every name and term, choose a voice that sounds different from the main speakers, and ask blind viewers for feedback. For long or dramatic content, many viewers prefer a skilled human describer.

How do I find the pauses for audio description?

Use a timestamped transcript or subtitle file and calculate the gap between the end of each cue and the start of the next. Mark gaps of a couple of seconds or more as possible slots, then confirm each one by listening, because recognition timestamps are approximate.

Does mydubly create audio description?

No. mydubly does not describe visual content or voice a written script. It can give you a timestamped transcript and SRT or VTT files of the dialogue, which help you find the pauses where your own descriptions can go.