Two audiences, two jobs
The difference starts with who the text is for. A deaf or hard-of-hearing viewer is missing the entire soundtrack, so the text has to carry everything audible that matters: the words, who said them, and sounds like a ringing phone or ominous music that change the meaning of a scene. A viewer watching a film in a language they do not speak can hear all of that perfectly well; they only need the words translated.
Those two jobs lead to different text even for the same moment.
Spoken in English: "Where did you leave the keys?" followed by a door slamming. English captions: "MAYA: Where did you leave the keys? [door slams]". Spanish subtitles: "¿Dónde dejaste las llaves?". A machine-generated English subtitle file: "Where did you leave the keys?" The caption carries the speaker and the sound; the subtitle carries only the translated words.
What captions contain
Captions aim to be a complete text equivalent of the audio. Beyond the dialogue, good captions include:
- Speaker identification when it is not obvious from the picture, such as an off-screen narrator or a voice on the phone.
- Meaningful sound effects: knocking, laughter, an alarm.
- Music cues, often with a short description of mood or lyrics where they matter.
- Manner of speaking where it changes meaning, such as whispering or sarcasm.
Captions are normally in the same language as the speech. For accessibility, the Web Content Accessibility Guidelines (WCAG 2) include captions for prerecorded audio in synchronized media at the most basic conformance level. If you have a legal or contractual accessibility obligation, check its exact requirements; this article is not legal advice.
What subtitles contain
Subtitles translate the dialogue for viewers who do not understand the spoken language. Because the viewer can hear the soundtrack, subtitles usually omit sound descriptions and speaker labels, and they are often condensed to keep reading speed manageable. A subtitle's job is to convey meaning quickly enough that the viewer can still watch the picture.
Where SDH fits in
SDH stands for subtitles for the deaf and hard of hearing. It is a hybrid: subtitle-style text, often in the same language as the audio, with caption-style additions such as sound cues and speaker identification. SDH tracks are common on streaming services and discs, partly because the formats used there handle subtitle tracks more easily than broadcast-style closed captions. For a viewer, SDH and captions do essentially the same job.
Regional differences in what the words mean
Usage is not uniform, which is a common source of confusion when briefing vendors or reading platform documentation:
- In the United States, "captions" usually means same-language text for accessibility, and "subtitles" means translation.
- In the UK, Ireland and much of Europe, "subtitles" is the everyday word for both, and "subtitles for the deaf and hard of hearing" is used when the accessibility version is meant.
- Many video platforms blur the line with a single menu that lists every text track together.
When commissioning work, describe the content you need, such as "same-language text with sound cues and speaker names", rather than relying on the label.
Same files, different content
Captions and subtitles usually live in the same file formats, such as SRT and WebVTT. The format does not decide which one you have; the content does. An SRT file can hold translated dialogue, same-language captions with sound cues, or anything in between. The format differences themselves are covered in SRT vs VTT.
When each one helps most
Captions pay off well beyond deaf and hard-of-hearing viewers. People watching on a muted phone, in a noisy open-plan office or in a language they read better than they hear all lean on same-language text. Learners in particular benefit from seeing and hearing the same words.
Subtitles pay off when a video's value depends on the original performance, such as an interview, a documentary or a vlog, and you want new audiences to hear the real voices. They are also far cheaper and faster to produce than a dubbed voice track, which makes them a sensible first step into a new language.
The limits of machine-generated caption files
Speech recognition produces the words and their timing, which is the hardest and most time-consuming part of making either track. It does not reliably produce the rest of what captions need:
- Speaker names: recognition models transcribe speech, not identity. mydubly's transcripts do not label speakers.
- Sound cues: non-speech sounds are generally not described in any consistent way.
- Editorial choices: line breaks, condensing long lines and reading speed may need adjustment.
So a machine-generated same-language file is a strong starting point for captions, not a finished caption track. For translated subtitles, the main review task is meaning; for captions, it is completeness.
Captions and subtitles with mydubly
mydubly's subtitle generator produces SRT and VTT files from your video or audio. In transcript mode you get subtitles in the spoken language, which you can turn into captions; choose a target language and you get translated subtitles instead. A full translation run includes SRT and VTT subtitles alongside the dubbed video. The spoken language is detected automatically, and 21 languages are supported.
The files are delivered separately, not burned into the picture, so viewers can switch them on or off and you can upload several languages to the same video; open vs closed captions explains why that matters. Transcripts and subtitles cost 1 credit per minute, so captioning a 20-minute video costs 20 credits, or 2 cents.
Turning a machine subtitle file into captions
If you need proper captions rather than plain subtitles, plan a short editing pass:
- Generate a same-language SRT or VTT file and open it in a subtitle editor or a plain text editor.
- Play the video and correct misheard words, especially names, numbers and technical terms.
- Add speaker labels where the speaker is off screen or not obvious.
- Insert sound cues in square brackets for sounds that matter to the story or instructions.
- Split overlong cues so viewers have time to read them, then upload the finished file to your player or platform.
Where to go next
To create a starting file for either track, upload a video to the subtitle generator. If your audience includes students, captions for educational video accessibility covers expectations in that setting.
Frequently asked questions
Are SDH subtitles good enough for accessibility?
In practice SDH and captions serve the same viewers, since both include speaker identification and sound cues. Whether a particular SDH track meets a specific legal or platform requirement depends on that requirement, so check its wording.
Can one file serve as both captions and subtitles?
Only for viewers of the same language. Same-language captions with sound cues work for deaf and hearing viewers alike, but viewers who do not speak the language still need a translated track.
Why do some subtitles leave out words that were spoken?
Subtitles are often condensed so that viewers can read them in the time available and still watch the picture. Captions for accessibility generally aim to be closer to verbatim.
Do automatic captions include sound effects?
Generally not in a consistent, reliable way. Speech recognition is built to transcribe speech, so sound cues and speaker names usually need to be added by a person.
Which should a short social video have, captions or subtitles?
Start with same-language captions, since many people watch social video with the sound off and they also serve deaf viewers. Add translated subtitles for each other language your audience speaks.