Why conventions matter
Every transcript is a set of decisions. Speech has pauses, overlaps, changes of pace and volume, laughter and false starts; written text has none of these unless someone chooses symbols for them. Without agreed conventions, two transcribers will mark the same recording differently, and a reader cannot tell whether a missing pause means there was none or that nobody marked it.
Conventions do three jobs. They make transcripts consistent across a project and across transcribers. They tell readers of your publication what the symbols in your extracts mean. And they record, as part of your method, what you chose to capture and what you left out. That last point matters: a transcript with no pause markings is not wrong, but it reflects a decision that should be visible.
The level of detail is a separate question from whether you keep fillers and false starts. That choice, between verbatim and cleaned-up text, is covered in the comparison of verbatim and clean verbatim styles. Notation sits on top of whichever style you choose.
Detailed notation: the Jefferson system
Conversation analysis (CA) studies talk-in-interaction in fine detail, and it relies on a notation system developed by Gail Jefferson. It is widely used, though individual researchers and journals adapt it, so always check the key that accompanies a given transcript. Commonly used symbols include:
- (0.5)
- A timed pause, here half a second
- (.)
- A micropause, audible but too short to time meaningfully
- [ and ]
- Start and end of overlapping talk, aligned across the two speakers' lines
- =
- Latching: one turn follows another with no gap, or one speaker continues across lines
- word
- Underlining marks stress or emphasis
- WORD
- Capitals mark talk noticeably louder than the surrounding speech
- °word°
- Degree signs mark talk noticeably quieter
- wo::rd
- Colons mark a stretched sound; more colons, longer stretch
- wor-
- A dash marks a cut-off
- ↑ and ↓
- Marked rise or fall in pitch
- >word< and <word>
- Talk faster or slower than the surrounding speech
- .hh and hh
- In-breath and out-breath
- ( ) and (word)
- Unclear talk, or the transcriber's best guess
- (( ))
- Transcriber's description, such as ((coughs)) or ((door closes))
Punctuation also changes meaning in CA transcripts. A full stop marks falling intonation, a comma continuing intonation and a question mark rising intonation, regardless of grammar. That is one reason a CA transcript cannot simply be produced by tidying a standard transcript.
Detailed notation is slow to produce. Timing pauses to a tenth of a second and aligning overlaps requires listening to short stretches many times, often with audio software showing the waveform. Researchers commonly transcribe only selected extracts in full CA detail and keep the rest of the recording in a simpler form.
Simpler conventions for interview and thematic research
Most interview-based research does not need CA detail. A simple set of conventions, applied consistently, covers what thematic, narrative or content analysis usually needs. A workable set looks like this:
- Speaker labels at the start of each turn: INT for interviewer, a pseudonym or code for each participant.
- [pause] for a noticeable pause, or [long pause] if length matters to you. Some projects use (3s) style timings for pauses above a threshold.
- [laughs], [sighs], [crying] and similar tags for non-verbal sounds that affect meaning.
- [inaudible m:ss] for words that cannot be heard, with the timestamp so someone can try again later.
- [word?] for a best guess.
- A dash for an interrupted or abandoned sentence.
- [overlapping] or a short note where two people speak at once, without aligning exact overlap points.
- Italics or capitals for strong emphasis, if you choose to mark it at all.
- [name removed] or [workplace] for anonymized details, decided with your de-identification rules.
The value lies in deciding in advance. If one transcriber writes (laughs) and another writes [laughter], search and coding both get harder.
Choosing the right level of detail
Match notation to what you will analyze. A few questions help:
- Will you interpret how something was said, or only what was said? Tone, hesitation and laughter matter more in studies of identity, emotion or interaction than in studies of practices or opinions.
- Will you analyze turn-taking, repair or overlap? If so, you need CA-style notation for at least those passages.
- Who will read the extracts? Dense notation is hard for non-specialists to follow; a policy audience may need simpler extracts with a note on what was omitted.
- How much time do you have? Detailed notation multiplies transcription time. Plan for it rather than trimming it halfway through.
Mixed approaches are common and defensible: a readable transcript for the whole dataset, with selected key passages re-transcribed in detail.
Writing a conventions key
A conventions key is a short document listing every symbol and rule used. It belongs in your project files, your thesis appendix and, in shortened form, in publications that show extracts.
A useful key includes:
- The overall style: clean verbatim, verbatim, or CA detail for selected extracts.
- Each symbol with its meaning and a one-line example.
- How speakers are labeled, and how pseudonyms were assigned.
- How unclear speech, guesses and anonymized details are marked.
- How timestamps are written and how often they appear.
- Spelling rules for dialect, nonstandard forms and code-switching, for example whether you write gonna or going to.
- What is deliberately not marked, such as intonation or breathing.
- The date and version, so later changes are traceable.
Write the key before the second transcript, test it on a short stretch with a second transcriber if you have one, and update it with dated changes.
A research team records twenty lessons to study how teachers respond to wrong answers. They agree on a simple key for whole-lesson transcripts: speaker codes for teacher and pupils, [pause] for gaps over about two seconds, [laughter], [inaudible m:ss] and [overlapping]. For the forty or so exchanges they select for close analysis, one team member re-transcribes the passage in Jefferson notation, timing pauses with waveform software and aligning overlaps. Their methods section explains both levels, and the appendix contains both keys.
Adding notation to a machine draft
Automatic transcripts are a reasonable starting text for simple conventions, provided you know what they leave out. Recognition models output words and standard punctuation. They do not produce pause timings, overlap brackets, emphasis or non-verbal tags, and they tend to drop fillers, false starts and many repetitions. Their punctuation reflects grammar, not intonation.
A practical sequence:
- Correct the words first, listening to the whole recording. The proofreading guide for AI transcripts covers an efficient method.
- Add speaker labels. General-purpose transcribers often do not provide them; the explainer on speaker diarization explains why automatic labeling is hard.
- Restore fillers and false starts if your style needs them.
- Add bracketed tags for pauses, laughter and unclear words in a second pass, using the timestamps to jump around.
- For passages needing CA detail, re-transcribe them from the audio rather than editing the machine text, since its punctuation and segmentation will mislead you.
Keep the uncorrected draft, the corrected text and the notated version as separate files, so your audit trail shows each stage.
Common mistakes with conventions
- Changing symbols partway through a project without updating earlier transcripts.
- Using CA symbols loosely, such as a full stop for sentence end in a transcript that claims Jefferson notation.
- Marking pauses in some transcripts and not others, then comparing pause frequency across participants.
- Leaving guesses unmarked, so readers cannot tell confident text from uncertain text.
- Publishing extracts without a key, or with a key that does not match the symbols shown.
- Treating a machine draft's punctuation as evidence of intonation.
Where mydubly fits and where it stops
mydubly produces the base text, not the notation. You choose a recording from your device, and it returns a plain transcript, a timestamped transcript with [m:ss] labels, and SRT and VTT files, in any of 21 languages with the spoken language detected automatically. The audio to text page lists the accepted formats and limits.
The timestamped transcript is useful for adding conventions, because you can find [inaudible m:ss] passages and selected extracts quickly. But mydubly adds no speaker labels, pause markings, overlap notation or sound tags, and its output leans toward clean text. Every symbol in your key is added by a person. If your research audio is sensitive, check with your ethics board or institution whether an external transcription service is permitted before uploading.
A one-hour recording costs 60 credits (6¢) to transcribe.
Next step
Draft your conventions key now, even if it is only ten lines, and apply it to one recording before transcribing the rest. If you work with interviews generally rather than detailed interaction, the interview transcription use case shows the standard workflow; for detailed CA work, budget time for selective re-transcription from the start.
Frequently asked questions
What are Jefferson transcription conventions?
They are a notation system for conversation analysis developed by Gail Jefferson. Symbols mark timed pauses, overlapping talk, latching, emphasis, volume, stretched sounds, pitch movement and breathing. Researchers and journals adapt the system, so every transcript using it should come with its own key.
Do I need conversation analysis notation for interview research?
Usually not. Thematic, narrative and content analysis generally need accurate words, speaker labels and a few tags for pauses, laughter and unclear speech. Detailed notation is worth the time when you analyze how talk is organized, for example turn-taking, repair or hesitation in specific passages.
How should I mark words I cannot hear?
Use a consistent tag such as [inaudible] with the timestamp, so you or a colleague can listen again later. For a best guess, put the word in brackets with a question mark, or in single parentheses if you follow Jefferson notation. Never leave a guess unmarked.
Can automatic transcription produce Jefferson notation?
General-purpose speech recognition does not. It outputs words with grammatical punctuation and tends to drop fillers and false starts. For detailed passages, re-transcribe from the audio with waveform software rather than editing the machine text, because its punctuation does not represent intonation.
Where does the conventions key go in a thesis or article?
In a thesis, the full key normally goes in an appendix, with a short reference in the methods chapter. In an article, include a brief key in a note, table or supplementary file covering the symbols that appear in your extracts. Check the journal's guidance, as some have preferred formats.
Should I time every pause?
Only if pause length is part of your analysis. Timing pauses precisely is slow and needs waveform software. Many interview projects mark only noticeable pauses with a tag, or time pauses above a set threshold, and say so in the key.