What decides accuracy for each
Accuracy is not a fixed property of either option. It depends heavily on the recording.
For AI transcription, the main drivers are audio quality, background noise and music, how far speakers are from the microphone, overlapping speech, accents and dialects that were less common in the model's training data, and specialist vocabulary such as drug names, legal terms or product codes. On a clean, single-speaker recording, a modern model like OpenAI's Whisper can produce a transcript that needs only light correction. On a noisy, crosstalk-heavy group discussion, it can miss or garble whole passages.
Human transcribers are affected by the same audio problems, but they bring things a model lacks: they can look up a name, recognize when a word does not make sense in context, mark a passage as inaudible rather than guessing, and ask a question. Their accuracy depends on skill, familiarity with the subject and the time they are given.
Turnaround
Machine transcription generally runs much faster than real time, so the text for a long recording arrives far sooner than a person could type it. A human transcriber has to listen, type and check, which takes considerably longer than the recording itself, and services add queue time on top. Rush delivery is usually available from human services at a higher price.
For one recording this may not matter. For a backlog of dozens of hours, or a workflow that needs text the same day, it decides the question.
Cost structure
The two are priced on the same unit, the audio minute, but at very different levels. AI transcription costs little enough that it is reasonable to transcribe everything and decide later what is worth reading. Human services usually charge more for difficult audio, many speakers, strict verbatim, specialist subjects and fast turnaround, so get a quote based on a representative sample. With mydubly, transcription is 1 credit per minute: an hour of audio is 60 credits, or 6 cents.
Confidentiality
The privacy question differs in kind. With a human service, a person listens to the whole recording, so you rely on their confidentiality agreement and the service's handling of files. With AI transcription, nobody needs to listen, but the audio is processed on a provider's servers, so you rely on its retention and access policies. Neither option is automatically compliant with any particular regulation or research ethics approval, and some recordings should not leave your control at all. The questions worth asking any provider are listed in cloud speech processing privacy.
Speaker labels, timestamps and formatting
Human services routinely label speakers, follow style guides, apply clean or strict verbatim on request and add timestamps at set intervals. AI tools vary. Many provide segment timestamps by default; speaker labeling, called diarization, is offered by some and not others, and it makes mistakes when voices are similar or overlap. mydubly's transcripts carry timestamps on every segment but do not label speakers.
Where AI transcription falls short
- Hard audio: noise, distance, echo and overlapping voices degrade machine output faster than human output.
- Silence and music: Whisper-style models can occasionally produce text that was never said in long silences or music; Whisper hallucinations explains why.
- Names and jargon: unfamiliar proper nouns and technical terms are often spelled phonetically or replaced with common words.
- Verbatim detail: models tend toward clean output and drop many fillers and false starts, which matters for linguistics or some legal work.
- No judgment calls: a model will not flag a passage as uncertain the way a careful transcriber marks something inaudible.
Where human transcription falls short
- Cost and time make it impractical to transcribe large archives or every meeting.
- Quality varies between transcribers, and long jobs may be split among several people with different habits.
- A person hears everything, which may not be acceptable for some sensitive material.
- Fixes and re-runs are slow and cost money each time.
Hybrid workflows that combine both
Many teams now use AI for the draft and people for the parts that matter. A typical workflow:
- Transcribe everything with AI to get searchable, timestamped text quickly.
- Use the timestamps to find the passages you will actually quote, publish or analyze.
- Have a person, yourself, a colleague or a professional, review and correct only those passages against the audio.
- Send recordings with very poor audio, heavy crosstalk or strict verbatim needs to a human service from the start.
- Keep a list of names and terms that the AI got wrong, and check for them first in later transcripts.
Suppose a researcher has 12 interviews of 50 minutes each, 600 minutes in total. An AI draft of all of them with mydubly costs 600 credits, which is 60 cents. Rather than paying for a full human transcription, the researcher reads the drafts, marks the passages relevant to the analysis, and corrects those carefully against the audio, sending only the two noisiest interviews to a professional transcriber.
AI transcription with mydubly
mydubly's transcription runs on Whisper, by default the large-v3-turbo model. Use audio to text for MP3, WAV, M4A, AAC, OGG or FLAC files, or video to text for MP4, MOV, WebM, MKV or M4V; files up to 2 hours are supported. You get a timestamped transcript and SRT and VTT files, with the spoken language detected automatically from 21 supported languages, and an optional translation.
There are no speaker labels, so plan to add them yourself for interviews and group discussions. Uploaded audio and results are deleted within 30 minutes of a job finishing. For reviewing what comes back, how to proofread an AI transcript gives a quick method.
Where to go next
To see how AI handles your own audio, transcribe a representative 5-minute clip on audio to text; it costs the 5-credit minimum. Read it against the recording, and you will know quickly whether AI alone, a hybrid or a human service fits the project. For interview-heavy work, interview transcription covers the specifics.
Frequently asked questions
Is AI transcription accurate enough for publishing quotes?
It can be, but verify every quote against the audio before publishing. Names, numbers and technical terms are the most common errors, and a misquote is costly even when the rest of the transcript is fine.
Do human transcribers make mistakes too?
Yes. Fatigue, unfamiliar accents and specialist vocabulary affect people as well, which is why professional services often include a second review pass for higher-priced tiers.
How should I add speaker names to an AI transcript?
Play the recording with the timestamped transcript open and add a name at the start of each turn. Doing this during your proofreading pass is efficient, because you are already listening along.
Can AI handle strict verbatim transcription?
Not reliably. Models tend to smooth speech, dropping many fillers and false starts. If every 'um' and repetition matters, use a human transcriber or plan a careful manual pass.
Is it cheaper to have a person correct an AI draft than to order a human transcript?
Often, because correcting is faster than typing from scratch, especially on clear audio. On very poor audio the draft may need so much work that a full human transcript is the better use of money.