Subtitles & captions in depth

Live Captioning Explained: Stenography, Respeaking and Automatic Captions

Live captions are produced in one of three ways: a stenographer writes speech in phonetic shorthand on a steno machine that software turns into text (CART when it serves people in the room), a respeaker repeats what is said into speech recognition trained on their voice, or automatic speech recognition transcribes the audio directly. Each trades latency, accuracy and cost differently, and all of them make errors that cannot be taken back once displayed. That is why live captions are usually corrected or replaced before a recording is published.

9 min read · Updated

Three ways to caption live speech

All live captioning has the same constraint: text must appear within seconds of the words being spoken, with no chance to rewind. The methods differ in who or what turns speech into text.

Stenography and CART
A trained stenographer types phonetic chords on a steno machine; software translates them into words using a personal dictionary
Respeaking
A trained respeaker repeats or condenses the speech, adding punctuation by voice, into speech recognition tuned to their own voice
Automatic live captions
Streaming speech recognition transcribes the audio directly, with no person in the loop
Hybrid
Automatic captions with a human editor correcting key errors, or a respeaker supported by a second person

The terminology varies by country and industry. CART, short for Communication Access Realtime Translation, usually means a human captioner providing access for deaf and hard-of-hearing people at a meeting, class or event. Broadcast stenocaptioning is the same skill applied to television.

Stenography and CART

A steno machine has a small number of keys pressed together in chords, each chord representing a syllable, word or phrase by sound. Captioning software looks each chord up in the captioner's dictionary and outputs text. Skilled stenographers write at the speed of natural conversation, and the delay between speech and text is short, typically a few seconds.

Quality depends heavily on preparation. A captioner who receives the agenda, speaker names, slides and technical terms beforehand can add dictionary entries so that names come out right first time. Without preparation, unfamiliar names may appear as phonetic fragments or as untranslated steno. Typical errors are homophones, words split oddly and occasional garbled strokes, but the captioner can often correct as they go.

CART providers commonly work on site or remotely through an audio feed, and many deliver a lightly edited transcript after the session. The skill takes a long time to learn and qualified captioners are in demand, so booking ahead matters.

Respeaking

In respeaking, sometimes called voice writing, a captioner listens through headphones and speaks the words again into speech recognition software trained on their voice, dictating punctuation and sometimes speaker changes. Their clear, steady speech is far easier for software to recognize than a noisy room, overlapping speakers or a strong accent.

Respeakers often condense slightly, dropping repetitions and false starts so captions keep up with fast speech. That makes captions easier to read but means they are not verbatim. The delay is usually longer than stenography because the respeaker must hear a phrase before repeating it, and recognition adds its own processing time. Respeaking is widely used in some broadcasting systems for live programmes. Errors tend to be misrecognitions of the respeaker's own words, which a well-trained respeaker catches and corrects on screen.

Automatic live captions

Automatic live captions come from streaming speech recognition: the audio is fed to a model in small pieces, and the model emits text as it goes. Many systems show a provisional guess first and revise it as more audio arrives, which is why words sometimes change on screen a moment after they appear. Video calling apps, streaming platforms and presentation software increasingly offer this at little or no extra cost.

The advantages are availability and scale: it works for every meeting, at any hour, in many languages. The weaknesses are predictable. Names, acronyms and specialist terms are often wrong; crosstalk, room echo and distant microphones degrade results; punctuation and sentence boundaries are guesses; and there is no one to notice that a key word came out wrong. Quality varies with audio conditions, accent and language, so testing in your actual setup tells you more than any general claim.

How speech recognition works on recorded files, where the whole file is available at once, is explained in how speech to text works. Live recognition has less context to work with, which is one reason recorded-file transcripts tend to be more accurate.

Latency versus accuracy

Every live method balances speed against correctness. The longer a system or person waits before committing to text, the more context it has, and the fewer errors it makes. Waiting too long, though, separates captions from the speaker's face and slides, which makes them hard to follow.

  • Stenography generally offers low delay with high accuracy when the captioner is prepared.
  • Respeaking adds delay because of the repeat-then-recognize step, in return for clean, readable text.
  • Automatic captions can be fast, but low-latency settings and revising guesses can produce flicker and more errors.

Measuring live caption quality is not just counting wrong words. Some broadcasters and regulators use models such as NER, which weight each error by how much it affects understanding, so a misspelled filler word counts for less than a wrong number or a reversed meaning. Requirements and targets vary by country and change over time; check the current guidance that applies to you rather than relying on a figure you have seen quoted.

Roll-up versus pop-on display

Live captions are usually shown as roll-up: two or three lines at the bottom of the screen, with new text appearing on the bottom line and older lines scrolling up. Roll-up suits live speech because words can appear as soon as they are produced, without waiting for a complete sentence. Some web players show a similar word-by-word stream.

Prerecorded programmes generally use pop-on captions, where each complete caption appears at once, is timed to the speech, and clears before the next. Pop-on reads more comfortably and can be positioned to avoid covering faces, but it needs the full text and timing in advance, which live captioning cannot provide. The broadcast display modes behind both styles are described in CEA-608 vs 708.

Choosing a live method by situation

  • A deaf participant attending a seminar or meeting: ask what they prefer. Many people who rely on captions prefer human CART for accuracy, and some prefer a sign language interpreter, which captions do not replace.
  • A large public webinar: automatic captions provide a baseline for everyone; adding a human captioner improves accuracy for an audience that depends on it.
  • Live broadcast news or sport: broadcasters typically use stenocaptioners or respeakers, following their regulator's rules.
  • Internal stand-ups and routine calls: automatic captions are often adequate, provided participants know they may contain errors.
  • Events with heavy technical vocabulary: human captioners with preparation materials usually handle names and terms better than unprepared automatic systems.

Translating speech live into another language is a separate problem with its own trade-offs, covered in real-time speech translation.

Limits and risks of live captions

  • Errors are public and permanent in the moment. A misrecognized word can change meaning before anyone can fix it.
  • Names and numbers are fragile across all methods without preparation.
  • Audio quality drives everything. A captioner or recognizer working from a room microphone hears far less than one on a direct feed from the sound desk.
  • Live captions are often incomplete in crosstalk, and may omit speaker identification and sound information that viewers need.
  • Recorded live captions carry the delay with them: if saved with the recording, they appear seconds after the speech.
  • Automatic captions may not meet accessibility expectations for events where deaf or hard-of-hearing attendees depend on them; check what your institution or the applicable rules expect, which vary and are not covered here as legal advice.

Correcting the record afterwards

Once the event is over, there is time to do what live captioning cannot. For any recording you publish, the usual practice is to replace or correct the live captions rather than keep them as they were.

  1. Get the cleanest recording available, ideally from the sound desk or the platform's original file rather than a screen capture.
  2. Decide between editing the live captions and regenerating captions from the recording. Regenerating from the full file usually gives better text and timing, since nothing has to be guessed in real time, and avoids the built-in delay.
  3. Correct names, numbers and terms against the agenda, slides and speaker list.
  4. Convert to pop-on style cues timed to the speech, and add speaker identification and sound information if the captions are for accessibility.
  5. Have a person review the result, then publish it alongside or in place of the old captions.
  6. Keep any edited transcript from the CART provider as a reference.

Recording conditions at live events also affect the transcript; transcribing live event recordings covers room sound, audience noise and taking a desk feed.

Example: a council meeting archive (hypothetical)

A town council streams its monthly meeting with automatic live captions. The archived video keeps them, so the clerk's office receives complaints that a resident's name and two budget figures were captioned wrongly. For the next meeting, the clerk downloads the 1 hour 50 minute recording after it ends, generates a new subtitle file from it, corrects names against the agenda, checks every number against the published papers, and replaces the archive captions. Live viewers still see automatic captions; the permanent record is corrected.

Where mydubly fits: recorded files only

mydubly captions recorded files only. It has no live mode and cannot caption a meeting, stream or broadcast as it happens. Where it fits is the correction step: after the event, choose the recording from your device (up to 2 hours per file, with the tab kept open while it processes), and you get SRT and VTT subtitles plus a plain and a timestamped transcript. A 90-minute recording costs 90 credits (9¢).

The subtitle files contain recognized speech only, one single-line cue per segment, with no speaker labels, sound tags or positioning. For accessibility use they should be reviewed by a person and supplemented with speaker identification and sound information where needed. Recordings longer than 2 hours need to be split first. The subtitle generator handles the captions, video to text covers the transcript, and the meeting transcription use case shows the workflow for recorded meetings.

Next step: plan live and after-event captions together

For your next event, decide two things in advance: which live method serves the people attending, and who will produce corrected captions for the recording afterwards. Booking a captioner, sending preparation materials and arranging a clean audio feed solve most live problems; a reviewed caption file made from the recording solves the rest.

Frequently asked questions

What is the difference between CART and live captions?

CART usually refers to human real-time captioning, often by a stenographer, provided as an access service for deaf and hard-of-hearing people at a meeting, class or event. Live captions is the broader term and includes automatic captions produced by speech recognition with no person involved.

Why do live captions lag behind the speaker?

Every method needs a moment to hear a word or phrase before writing it. A stenographer's delay is usually short; respeaking adds the time to repeat the phrase; automatic systems wait for enough audio to make a confident guess. Network and display processing add a little more.

Why do words change on screen in automatic live captions?

Streaming recognition shows a provisional guess as soon as possible, then revises it when later audio provides more context. The final text is often more accurate than the first guess, at the cost of some visible flicker.

Are automatic live captions accurate enough for accessibility?

Sometimes, but not reliably. They struggle with names, technical terms, accents, crosstalk and poor audio. For events where attendees depend on captions, ask them what they need and consider a human captioner; check your institution's own policy too.

Should I keep the live captions on a recorded event?

Usually not as they are. They contain uncorrected errors and appear a few seconds after the speech. Correcting them, or regenerating captions from the recording and reviewing them, gives viewers of the recording a much better result.

Can mydubly caption a live stream or meeting?

No. mydubly processes recorded video and audio files only. You can use it after the event to produce subtitle files and a transcript from the recording, then review and publish those.