AI captions

AI Caption Generator: How It Works and How to Check It

AI captions are a speech recognition model's guess at what was said and when. On clear audio the guess is usually close; with names, numbers, music or overlapping voices it slips. This page explains how the captions are produced, where the errors come from, and how to review them quickly in the editor before you publish.

Subtitle editor

Load a video, get subtitles from AI, an SRT/VTT file or your own typing, style them, and download an MP4 with the subtitles in the picture.

MP4, MOV, M4V, WebM or MKV · free to edit and render · no watermark · the video stays on your device

From sound to timed text

Three things happen to produce an AI caption, and knowing them explains most of the mistakes you'll meet.

Recognition
The audio is transcribed by Whisper (large-v3-turbo), an open speech recognition model trained on speech in many languages and recording conditions. It predicts the most likely words, which is why it fills a gap with something plausible rather than leaving it blank.
Word timing
Every recognised word gets a start and an end time. These are what let a caption appear exactly as a phrase begins.
Line building
The timed words are grouped into short single lines of about 32 characters, each on screen for no more than about 2.6 seconds.

Before any of that, your browser extracts and compresses the soundtrack, so only sound reaches the server and the picture stays on your computer. The language is identified from the speech itself, with no setting needed. For more on the model, read what Whisper is.

Why AI captions go wrong

Errors aren't random. They gather around the same causes in almost every video.

  • Proper nouns. Names of people, places, products and companies the model has rarely heard get swapped for common words that sound alike.
  • Numbers. Fifteen and fifty, years, prices and phone numbers are easy to mishear and come out formatted inconsistently.
  • Music and noise. A soundtrack under speech, wind on an outdoor microphone or a busy room masks consonants.
  • Overlap. When two people talk at once, one voice usually wins and the other's words vanish or blend in.
  • Accents and pace. Strong regional accents and very fast delivery raise the error count, especially around specialist vocabulary.
  • Silence and instrumentals. Long quiet stretches occasionally produce text nobody said, a known quirk of this kind of model.

Each of these has its own explainer: background noise, accents and hallucinated text.

Two things AI captions never contain are speaker names and descriptions of sounds. If your viewers rely on captions because they can't hear the audio, add those by hand; SDH subtitles shows the conventions.

How accurate will yours be?

No honest tool can promise a figure for your video, because accuracy depends more on the recording than on the software. A single presenter on a decent microphone in a quiet room often needs only a handful of fixes. A street interview with traffic, two voices and brand names can need a correction every few lines. Researchers measure this as word error rate, but in captions the errors that hurt are the visible ones: a misspelled guest name in the first ten seconds does more damage than a dropped filler word.

Translated captions add a second layer. The translation can only be as good as the transcript beneath it, and line timing becomes approximate because translated text is divided by length rather than matched to each word.

A fast review routine

A careful full read is still the goal, but the right order gets you there sooner. In the editor:

  1. Before pressing play, scroll the line list and look at the warning markers. Lines flagged as too fast, too long, overlapping or empty are the first to fix.
  2. Skim the list for every person, product and place name, and correct each occurrence; the same mistake usually repeats throughout.
  3. Play from the start at normal speed with a hand on the keyboard. Press Space to pause at an error, click the line, type the fix and press Space again.
  4. When you're unsure what was said, press Left to jump back a second, or Shift+Left for a tenth of a second.
  5. Where a line breaks a phrase in an awkward place, put the playhead on the right word and press Split, or Merge a stray fragment with the line after it.
  6. Watch the final minute a second time. Endings often have music, applause or goodbyes that the model handled poorly.

Ctrl or Cmd+Z undoes any change, and the work is saved as a draft in this browser as you go. Fixing auto-generated subtitles goes deeper on particular error types.

Recording so the AI gets it right

The cheapest caption fix happens before you hit record.

  • Put the microphone close to the speaker. A lapel or headset mic beats the camera's built-in one almost every time.
  • Keep background music out of the recording. If you edit the video yourself, generate captions from a cut without the music bed, export the SRT, then import it onto the final version with music.
  • Ask people not to talk over each other, and to repeat a question if it came from off-mic.
  • Say unusual names slowly and clearly the first time they come up.

Speaking clearly for transcription has more recording advice along these lines.

Three clips and what to expect

Solo explainer, USB microphone, no music

Expect mostly punctuation and the odd technical term to fix. Ten minutes costs 10 credits (1¢), and the review takes roughly as long as the video.

Podcast clip with two hosts laughing over each other

Expect missing words where the voices overlap and short fragments around the laughter. Merge the fragments and type overlapping lines by hand.

Event recording with a band before the keynote

Expect stray lines during the music. Delete them, then check the keynote speaker's name and any figures quoted from the stage.

Checklist before you publish

  • Every name and brand spelled the way its owner spells it.
  • Numbers, dates and prices checked against the source.
  • No warnings left in the line list, or only ones you've chosen to accept.
  • No invented lines during music or silence.
  • Speaker names and sound cues added if the captions serve deaf or hard-of-hearing viewers.
  • The look checked on the preview at full size, including the very first and last lines.

Then render the captioned MP4 or export the caption file. The settings and re-run options are covered on the auto subtitle generator page, and if a file is all you need, the subtitle generator hub gets you there with fewer steps.

About the subtitle editor on this page

Opens
MP4, MOV, M4V, WebM, MKV videos your browser can decode
Subtitles from
AI (from the speech, optionally translated into one of 21 languages), an SRT or VTT file, or typing them in
You download
An MP4 with the subtitles drawn into the picture at the original resolution and frame rate, plus SRT and VTT files
Cost
Editing, styling, rendering and SRT/VTT export are free, with no account and no watermark. AI subtitles cost 1 credit per minute (5-credit minimum per video); signing in with Google gives 100 free credits and accounts get 10 free credits a day
Privacy
The video never leaves your device and is rendered in your browser. Only for AI subtitles is the audio sent, over HTTPS, and it is deleted within 30 minutes of delivery
Rendering needs
Chrome or Edge 94+, Safari 16.4+, or Firefox 130+ on a computer. Editing and SRT/VTT export work in any current browser

More: subtitle and caption generators

Subtitle tools

Frequently asked questions

Which AI model writes the captions?

Whisper large-v3-turbo, a speech recognition model that also identifies the spoken language. Any translation is applied to its transcript afterwards.

Can AI captions tell who is speaking?

No, they contain the spoken words only. Type a speaker's name at the start of a line yourself, for example in capitals followed by a colon.

Why do my AI captions include words nobody said?

Speech recognition models sometimes produce text during music, noise or long silence. Delete those lines; they usually sit in gaps between real speech.

Does a better microphone really change the captions?

Usually more than anything else you can do. Close, clear speech with little background sound leaves the model the least to guess.

Do AI captions work for Hindi?

Yes. Hindi is one of the 21 supported languages and the captions are written in Devanagari script. Romanised Hinglish isn't produced automatically.