What automatic speech recognition actually means
ASR answers one narrow question: which words were spoken in this audio? It does not decide what the words mean, who said them or whether they are true. Those are separate problems handled by other systems, such as natural language understanding, speaker diarization or translation. Keeping that boundary in mind explains a lot about what a transcription tool can and cannot do.
The terms "speech recognition", "speech to text" and "ASR" are used almost interchangeably. Engineers tend to say ASR; product pages tend to say speech to text. "Voice recognition" is sometimes used loosely for the same thing, but strictly it means identifying a person by the sound of their voice, which is a different task.
The input is audio, either a live microphone stream or a recorded file. The output is text, usually with timing information. Everything in between is the model's job: separating speech from noise, mapping sounds to candidate words and choosing the most plausible sentence.
What ASR output looks like
A raw transcript from a modern system is a list of segments. Each segment has a start time, an end time and a piece of text, typically a phrase or a sentence. Some systems add word-level timings, a confidence score per word and the detected language.
00:00:04.20 to 00:00:07.86 "Welcome back. Today we're looking at how the filter works." then 00:00:08.10 to 00:00:11.02 "First, let's open the settings panel."
- Segment
- A stretch of recognized text, usually a phrase or sentence, with a start and end time
- Timestamp
- The position in the recording where a segment or word begins or ends
- Confidence
- A model score for how sure it is; useful for flagging words to check, not a guarantee
- Language tag
- The language the system detected or was told to expect
- Plain transcript
- The segment texts joined together without timing
Those timestamps are what turn a transcript into subtitles: an SRT or VTT file is essentially the same list of timed segments written in a standard format, which is how a subtitle generator builds captions from recognized speech.
Template matching: the first recognizers
The earliest systems compared incoming sound against stored examples. In 1952 Bell Labs demonstrated "Audrey", which recognized spoken digits from a single speaker it had been tuned to. IBM's Shoebox, shown in the early 1960s, handled a small vocabulary of digits and command words. These machines matched acoustic patterns directly, and they broke down as soon as a different person spoke or a word was said at a different speed.
Dynamic time warping, developed in the following decades, improved template matching by stretching and compressing time so that a slowly spoken word could still be aligned with a quickly spoken template. It made isolated-word recognition practical for limited vocabularies, but it could not scale to continuous, natural speech from any speaker.
End-to-end neural models
The latest shift removed the separate components. End-to-end models learn the whole mapping from audio features to characters or word pieces inside one network. Important ideas along the way included connectionist temporal classification (CTC), which lets a network learn the alignment between audio and text without frame-by-frame labels; attention-based encoder-decoder models; and transducer models suited to streaming. Self-supervised pretraining on unlabeled audio, as in wav2vec 2.0, showed that a model could learn useful speech representations before seeing any transcripts.
OpenAI's Whisper, released in 2022, took a different route: a transformer encoder-decoder trained on a very large collection of audio paired with transcripts gathered from the web. Because that data was messy and varied, the model learned to cope with accents, background noise and many languages, and to write punctuation and capitalization directly. A step-by-step look at how such a model turns sound into text is in how speech to text works.
Where ASR is used today
- Transcribing interviews, meetings, lectures, podcasts and voice memos
- Generating captions and subtitles for video
- Dictation on phones, in word processors and in medical or legal documentation tools
- Voice assistants and voice commands in cars and smart speakers
- Call-center analytics and searchable archives of recorded calls
- The first stage of speech translation and AI dubbing, where the transcript is translated and then spoken in another language
In that last case, recognition quality sets the ceiling for everything after it. A misheard product name becomes a mistranslated product name and then a mispronounced one in the new voice track.
What ASR does well
- Speed: an hour of clear speech is transcribed far faster than a person could type it.
- Cost: machine transcription costs a small fraction of a human service per minute.
- Searchability: once speech is text, you can search, quote, index and translate it.
- Accessibility: transcripts and captions make audio and video usable for deaf and hard-of-hearing people and for anyone watching without sound.
- Consistency at scale: the system treats the hundredth file the same way it treated the first.
Limits of automatic speech recognition
Accuracy still depends heavily on the recording. Noise, echo and overlapping speakers cause most errors, which is why background noise gets its own article. Rare names, brand terms and jargon are often spelled phonetically or replaced with a common word. Accents that are underrepresented in training data fare worse than well-represented ones.
ASR also produces text, not understanding. It will not tell you who is speaking unless a separate diarization step is added, and it cannot fix a speaker who misspoke. Neural models can occasionally invent words during silence or music, a failure called hallucination. And a single headline accuracy figure says little about your own files; word error rate explains how accuracy is measured and why testing on your own audio matters more.
How mydubly uses automatic speech recognition
mydubly's transcription runs on Whisper, using the large-v3-turbo model by default. When you open a video or audio file, your browser extracts the audio, splits it into windows of roughly 30 seconds cut at quiet moments, and uploads only those audio chunks; the video file itself stays on your device. Recognition returns timestamped segments, which become a timestamped transcript plus SRT and VTT subtitle files in the spoken language. You don't need to say which language is spoken, because it is detected automatically.
Transcripts cost 1 credit per minute, so a 45-minute interview is 45 credits, or 4.5 cents, with a minimum of 5 credits per file. The transcript can optionally be translated, and it is also the starting point when you produce a translated voice track with the video translator. Transcripts do not label speakers, so recordings with several people need names added by hand.
Try ASR on your own recordings
- Pick a short file with clear speech, such as a two-minute clip from a talk or a voice memo.
- Open it in video to text or, for audio-only files, audio to text.
- Read the transcript while listening and note which words are wrong, starting with names and numbers.
- Try a noisier file and compare. The difference shows you how much the recording itself matters.
Frequently asked questions
What is the difference between ASR and voice recognition?
ASR identifies what was said; voice recognition, in the strict sense, identifies who is speaking from the characteristics of their voice. The terms are often mixed up in everyday use. Most transcription tools, mydubly included, perform ASR only and do not identify or label speakers.
Is ASR part of natural language processing?
ASR is usually treated as a neighbor of natural language processing rather than part of it. It converts audio to text; NLP tasks such as summarization, sentiment analysis or translation then operate on that text. Modern pipelines chain them, so a recognition error can carry into every later step.
Why does ASR output include timestamps?
Timestamps link each piece of text to the moment it was spoken. They let you jump from a transcript line back to the audio, verify a quote and build subtitle files, which are essentially timed segments of text. Without them a transcript cannot be turned into captions.
Does automatic speech recognition work in real time?
Some systems are designed for streaming and produce text while someone is still talking, which is what voice assistants and live captioning need. Others, including Whisper as normally used, process audio in windows and work best on recorded files. mydubly is file-based: you transcribe a recording, not a live stream.
How accurate is ASR today?
On clear, single-speaker recordings, modern systems can produce transcripts that need only light correction. Accuracy drops with noise, echo, overlapping speech, heavy jargon and less-resourced languages. The honest answer for any specific project is to transcribe a sample of your own audio and check it.