Accessibility & inclusive video

Captioning Sound Effects, Music and Speakers: Conventions That Work

To caption sound effects, describe the sound briefly in square brackets at the moment it happens, such as [door slams] or [phone vibrating], and include it only when it matters to understanding the scene. The same approach covers music, speaker identification, tone of voice and speech that cannot be made out. Conventions differ between style guides, so pick one set, write it down and apply it consistently across every video.

9 min read · Updated

What non-speech information captions need to carry

Captions for deaf and hard-of-hearing viewers have to carry more than the words. A hearing viewer also gets who is speaking, how they say it, what sounds happen off screen and what the music is doing to the mood. Without that, a viewer reading captions can miss why a character turns around, why the audience laughs, or that a line was meant as a joke. Where captions end and subtitles begin is covered in captions versus subtitles; this article is about the extra information that makes captions complete.

The non-speech layer has five parts:

  • Sound effects: doors, alarms, phones, footsteps, crashes, notification chimes.
  • Music: whether it is playing, what it sounds like, lyrics when they matter.
  • Speaker identification: who is talking when the picture does not make it obvious.
  • Manner of speaking: whispering, shouting, sarcasm, a strong accent, stammering.
  • Speech that is off screen, overlapping, unclear or in another language.

Writing sound tags

Most caption style guides put sound descriptions in square brackets so they are visibly different from dialogue; some older broadcast styles use parentheses or all capitals. Within that frame, a few habits make tags quick to read:

  • Describe the sound, not the fact that there is a sound: [glass shatters] rather than [sound of glass breaking].
  • Use a short, present-tense phrase with an active verb: [dog barks], [engine starts], [crowd cheers].
  • Name the source when it matters and is clear: [car horn honks] is more useful than [honking].
  • Add meaning when the sound carries it: [distant siren approaching], [knock at door].
  • Avoid onomatopoeia such as bang or whoosh unless the exact sound is part of the joke or the content.
  • Keep tags short. A tag that takes longer to read than the sound lasts steals reading time from the dialogue around it.

Place the tag so it appears when the sound happens. A tag can sit in its own cue during a pause, or at the start of a cue before the line it explains, such as [phone buzzes] followed by the speaker answering it.

Captioning music

Music needs captioning when it sets the mood, carries the story or is the content itself. A few common patterns:

Background music starts
A short descriptive tag: [soft piano music], [tense orchestral music], [upbeat electronic music]
Mood changes
A new tag when the music shifts in a way that matters: [music turns ominous]
Music stops
[music stops] when the silence itself is meaningful
Known song
The title, and the artist if relevant, in a tag: [Clair de Lune playing]
Lyrics
Sung words between music note symbols, for example ♪ down by the harbour ♪, when the lyrics matter
No lyrics needed
A tag such as [song continues] when the words do not add anything

Describe music by its feel rather than by technical terms: tense, playful, melancholy and upbeat are more useful to a viewer than minor key or 120 beats per minute. For lyrics, consider whether they carry meaning in the scene; a song a character sings to another character usually does, while a backing track under a product demo usually does not. Lyrics are often copyrighted text, which matters if you are reproducing them in full, and speech recognition handles singing poorly, so lyrics are almost always typed by hand.

Identifying speakers

Speaker labels are needed when the viewer cannot tell from the picture who is talking: off-screen voices, narration, phone calls, a group shot where lips are hard to see, or a sudden cut to a new person.

  • Use a name or role in brackets or followed by a colon, for example [Maria] or NARRATOR:, and follow one style throughout.
  • Label at the first line and again whenever the speaker changes in a way the picture does not show.
  • Do not label every line when the speaker is clearly on screen; it adds reading load for no gain.
  • When two people speak within one cue, many styles start each person's line with a dash.
  • Before a name is introduced, use a description: [man on phone], [interviewer].

WebVTT has a voice span that marks who is speaking, but most players do not display the name visually, so a visible label in the caption text is the safer choice for accessibility. Automatic speaker detection is a separate technology with its own failure modes, explained in what speaker diarization is.

Manner of speaking, off-screen and unclear speech

Tone tags tell the viewer something they would otherwise miss: [whispering], [shouting], [sarcastically], [voice breaking], [slurring]. Use them where the delivery changes the meaning; do not tag ordinary speech. If a sarcastic line reads as sincere without the tag, it needs the tag.

  • Off-screen speech: label the speaker, such as [Dad, off screen], or follow your style guide if it uses italics for off-screen voices in players that support them.
  • Overlapping speech: caption the main thread and add a tag such as [overlapping chatter] or [people talking at once].
  • Unclear speech: [inaudible] or [indistinct] for words that cannot be made out. Do not guess and present the guess as fact.
  • Another language: either caption the words if your audience reads that language, translate them if the script calls for it, or use a tag such as [speaking Japanese].
  • Meaningful silence: [silence] or [no audio] when the absence of sound is the point, such as a deliberate audio dropout.

How much to include

The test for every tag is whether a viewer who cannot hear would miss something important without it. Include sounds that:

  • Explain a reaction on screen, such as laughter, a gasp or someone turning toward a noise.
  • Happen off screen and affect what follows, such as a doorbell or an explosion.
  • Carry the mood in a way the picture does not, such as ominous music under a calm shot.
  • Are part of the content, such as an alert sound in a software tutorial.

Leave out constant background noise that does not change, routine footsteps, and sound effects that simply match what is clearly visible. A hearing viewer filters these out, and a caption track crowded with tags makes the dialogue harder to follow.

Hypothetical example: a two-minute documentary scene

The raw speech captions for a scene in a fishing village show three lines of dialogue. A reviewer adds [gulls crying] and [waves crashing] at the start to set the place, [boat engine sputters] where the fisherman frowns at his motor, [Ana, off screen] before a voice calls him from the shore, and [gentle guitar music] when the score begins. She leaves out the continuous wind because it never changes and never matters to the story. The finished track has eight cues instead of three, and a muted viewer now understands why the fisherman stops talking.

Pitfalls and limits of sound captioning

  • Editorialising: [creepy music] tells the viewer how to feel, while [low, slow strings] describes; many guides prefer description unless the intent is obvious.
  • Inconsistency: [laughter] in one video, (laughs) in the next and [LAUGHING] in a third. A one-page style sheet fixes this.
  • Tags that are too long for the time available, pushing reading speed beyond what viewers can manage.
  • Captioning sounds that match the picture exactly, which adds reading load without adding information.
  • Guessing at unclear words instead of marking them inaudible.

Sound captioning is also interpretive. Two careful captioners can describe the same music differently, and no tag fully conveys a piece of music or a vocal performance. That is a reason to aim for consistent, useful descriptions, not exhaustive ones.

Adding sound information to mydubly captions

mydubly's subtitle files contain recognized speech only. Each SRT or VTT cue is a single line of what was said, with no sound tags, no music descriptions, no speaker labels and no tone tags, so everything in this article is added by hand. Because the speech recognizer uses a voice activity filter, stretches of music or noise without speech usually produce no cues at all, and those timing gaps are a good map of where sound tags may be needed.

  1. Generate the subtitle file from your video with the subtitle generator. The spoken language is detected automatically, and the timestamped transcript downloaded with it is easier to scan than the cue list.
  2. Watch the video with the timestamped transcript open and note every gap where something audible happens without speech.
  3. Open the SRT in a subtitle editor, as described in how to edit an SRT file, and add sound and music tags as new cues in those gaps.
  4. Add speaker labels at the start of cues where the speaker is off screen or changes without a visible cut.
  5. Check any cues that fall during music or singing. Recognition models can mishear or occasionally invent words in those passages, so delete or correct them and add lyrics by hand if they matter.
  6. Save the file as UTF-8 so music note symbols and accented letters survive.

If you also make a translated subtitle track, its cues are speech only as well, and the sound tags need to be written in the target language. The SRT generator page describes the file format mydubly produces.

Next step: write your house style for non-speech captions

Before editing your next caption file, write down five decisions: brackets or parentheses, capitalisation of tags, how speakers are labelled, how music and lyrics are shown, and the wording for unclear speech. Keep it to one page, share it with everyone who edits captions, and add to it when a new case comes up. For where sound captioning sits among the other checks before release, see the video accessibility checklist.

Frequently asked questions

Should sound effects in captions go in brackets or parentheses?

Square brackets are the most widely used convention in current caption style guides, because they separate sounds clearly from dialogue. Some broadcast and legacy styles use parentheses or capitals instead. Either is acceptable as long as you use one style consistently and it does not clash with your platform's own conventions.

Do I need to caption background music?

Caption it when it affects meaning or mood, starts or stops in a way that matters, or is the subject of the content. A brief tag at the start, such as [calm acoustic music], is usually enough, with a new tag only when the music changes noticeably. Music that is purely incidental and unchanging can often be left out.

How do I caption song lyrics?

Put the sung words between music note symbols, for example ♪ lyrics here ♪, timed to when they are sung. Only caption lyrics when they carry meaning; otherwise a tag such as [song playing] is enough. Remember that speech recognition handles singing poorly, so lyrics usually need to be typed by hand.

How do I show who is speaking in captions?

Add a name or role in brackets or followed by a colon at the start of the line, but only where the speaker is not obvious from the picture, such as off-screen voices, narration or quick changes between people. Keep the label style consistent across the whole video and your other videos.

What should I write when I cannot make out what someone said?

Use a tag such as [inaudible] or [indistinct] rather than guessing. If several people talk at once, caption the main speaker and add a tag like [overlapping conversation]. Presenting a guess as dialogue can mislead a viewer who has no way to check it against the audio.

Should laughter be captioned?

Usually yes when it explains a reaction or tells the viewer a line was a joke, such as [audience laughs] after a punchline. Constant canned laughter or a speaker's light chuckle mid-sentence may not need a tag each time. Use the same test as for any sound: would a viewer who cannot hear miss something without it.

Does mydubly add sound tags or speaker labels?

No. mydubly's SRT and VTT files contain recognized speech only, as single-line cues, with no sound descriptions, speaker labels or styling. Use the gaps between cues and the timestamped transcript to find where tags belong, then add them in a subtitle editor before using the file as accessibility captions.