Three transcription styles at a glance
Transcription vendors and research teams use slightly different labels, but most styles fall into three groups.
- True verbatim
- Every word and sound as spoken: fillers, stutters, false starts, repetitions, interruptions, laughter, pauses and nonstandard grammar.
- Clean verbatim
- The speaker's own words and sentence structure, minus fillers, stutters, false starts and most repetitions. Sometimes called intelligent verbatim.
- Edited
- Clean verbatim further polished: grammar corrected, rambling condensed, sometimes paraphrased for clarity while keeping the meaning.
What true verbatim captures
A true verbatim transcript treats the recording as evidence. Fillers such as "um", "uh", "like" and "you know" stay in. So do false starts ("I went to the, we drove to the station"), repeated words, stammers, self-corrections and trailing sentences. Non-speech events are marked in brackets, for example laughter, coughs or crosstalk, and long pauses may be noted with their length.
Specialist fields go further. Conversation analysts use detailed notation systems, such as the one developed by Gail Jefferson, to mark overlaps, pause lengths in tenths of a second, pitch movement and emphasis. That level of detail is far beyond general transcription and is nearly always done by trained people.
What clean verbatim removes, and what it must keep
Clean verbatim removes material that does not carry meaning: fillers, stutters, abandoned starts and accidental repetitions. It should not change what the speaker said. Word choice, slang, dialect forms and sentence structure stay, even when they are not standard written grammar. Deliberate repetition for emphasis ("it was very, very expensive") is usually kept.
The line can be blurry. Is "like" a filler or a meaningful comparison? Is a restarted sentence a false start or a correction that changes meaning? Good clean verbatim style guides answer those questions explicitly, so different transcribers make the same calls.
True verbatim: "So, um, we, we basically, uh, I think we launched it in, like, March? No, April. Yeah, April." Clean verbatim: "So we basically, I think we launched it in March? No, April. Yeah, April." Edited: "We launched it in April."
Notice that clean verbatim keeps the self-correction from March to April, because it changes the meaning, while the edited version keeps only the final fact.
Why Whisper-style models lean toward clean output
AI transcription is not neutral on style. Whisper learned from large amounts of audio paired with transcripts and subtitles found on the web, and most of that text was written to be read: subtitle writers and transcribers routinely drop fillers and stutters. The model learned that pattern, so it tends to omit "um" and "uh", often smooths over stutters and short false starts, and writes punctuated, readable sentences.
The behavior is a tendency rather than a rule. Fillers sometimes survive, especially when they are long or stressed, and repetitions may be kept or dropped inconsistently within the same file. Whisper's reference implementation accepts an initial text prompt, and a prompt written with fillers can nudge the output toward keeping them, but this is not reliable enough for work that needs consistent true verbatim. Non-speech events like laughter are generally not marked at all.
When true verbatim matters
- Legal and evidentiary work, such as depositions, police interviews and recorded statements, where hesitation, interruption and exact wording can matter to how a statement is interpreted. Courts and agencies often have their own transcript rules, so follow those.
- Linguistics, conversation analysis and discourse research, where disfluencies, overlaps and pauses are the object of study.
- Some qualitative research, where an interviewee's hesitation around a sensitive question is part of the finding.
- Speech-language therapy and assessment, where stuttering and word-finding pauses are what is being measured.
- Compliance reviews of sales or support calls, when a regulator or policy requires exact wording.
When clean verbatim is the better choice
For most other uses, clean verbatim is easier to read and loses nothing important. Subtitles benefit most: viewers have limited reading time, and fillers waste it. Published interviews, articles, show notes and study notes read better without them. Clean text is also better input for translation and dubbing, because fillers and false starts translate awkwardly and would be voiced in the dub. General research interviews that use thematic analysis usually work fine with clean verbatim; see research interview transcription.
Limitations of AI transcripts for verbatim work
If your project genuinely needs true verbatim, an AI transcript is a draft at best.
- Fillers, stutters and false starts are dropped inconsistently, so you cannot assume either style throughout.
- Non-speech events such as laughter, sighs and crosstalk are not marked.
- Pause lengths are not reported directly, although segment timestamps give a rough clue.
- Overlapping speech is usually rendered as one speaker's words, or partly lost.
- Speaker turns are not labeled by the model itself, which verbatim conventions usually require.
Producing true verbatim from an AI draft means listening to the whole recording, not just flagged passages, which is where the cost of verbatim work really lies.
Producing transcripts in either style with mydubly
mydubly's transcription runs on Whisper, and it does not offer a transcription-style setting, so expect output close to clean verbatim: readable sentences with most fillers removed. Its transcripts also do not label speakers. For subtitles, show notes, translation and dubbing, that is usually exactly what you want.
For verbatim projects, use the transcript as a time-saving draft:
- Transcribe the recording with audio to text, or video to text for video files, and download the timestamped transcript.
- Agree on a style guide first: which fillers to include, how to mark pauses, laughter and crosstalk, and how to label speakers.
- Listen to the full recording at normal speed, inserting fillers, false starts and bracketed events into the draft.
- Add speaker labels as you go, using the timestamps to keep your place.
- Do a final check of passages where exact wording matters most.
Suppose a team has twelve 60-minute interviews to transcribe in true verbatim. Running them through transcript mode costs 720 credits, or 72 cents, and gives each transcriber a timestamped draft with the words in place. The transcribers then spend their time adding disfluencies, pauses and speaker labels rather than typing every word from scratch.
The broader comparison of when a person should do the job is in AI vs human transcription, and proofreading an AI transcript covers the clean-verbatim review workflow.
Choosing your style and next step
Pick the style before transcription starts: true verbatim if the way something was said is part of the evidence or the data, clean verbatim for nearly everything else, and edited when the transcript is really the raw material for a written piece. Then upload a recording to audio to text and judge the draft against the style you chose.
Frequently asked questions
Is intelligent verbatim the same as clean verbatim?
Usually, yes. Many transcription services use intelligent verbatim and clean verbatim interchangeably for transcripts that drop fillers and false starts but keep the speaker's words. Some use intelligent verbatim for a slightly more edited style, so check the provider's definition.
Does AI transcription keep filler words like um and uh?
Whisper-style models usually drop most of them, because the transcripts they learned from mostly omitted fillers. Some survive, so the output is close to clean verbatim but not perfectly consistent.
Which transcription style do academic researchers use?
It depends on the method. Conversation analysis and linguistics need detailed true verbatim, while many interview studies using thematic analysis work from clean verbatim. Check your methodology and ethics approval, which sometimes specify the style.
Should subtitles be verbatim?
Generally not. Subtitles are usually clean verbatim or lightly edited so viewers can read them in time, although subtitles for deaf and hard-of-hearing viewers add sound descriptions such as music or laughter.
Is true verbatim more expensive?
With human transcription services it usually costs more, because it takes longer to produce and check. With AI tools the machine step costs the same, but turning the draft into true verbatim requires a full human listening pass.