Research & academic workflows

Doing Content Analysis on Interview, Media and Meeting Transcripts

Content analysis applies a defined coding frame to transcripts so that categories can be counted and compared systematically. To do it well, decide your unit of analysis, write category definitions clear enough for a second coder to apply, test intercoder reliability, and only then code the full set. If transcripts come from speech recognition, check them first, because misrecognized words, dropped fillers and inserted phrases change counts in ways that look like findings.

8 min read · Updated

What content analysis is, and how it differs from thematic analysis

Content analysis is a family of methods for systematically describing the content of communication. In its classic quantitative form, researchers define categories in advance, apply them to units of text according to explicit rules, and count how often each category occurs. Qualitative content analysis keeps the systematic coding but places more weight on interpreting categories in context, and may develop categories from the data.

What distinguishes content analysis from thematic analysis is the emphasis on a stable coding frame, reliability between coders, and, often, frequencies and comparisons across groups or time. Thematic analysis is more interpretive and builds patterns of meaning; the article on preparing transcripts for thematic analysis covers that approach. Many projects use both, for instance counting how often a topic arises and then interpreting what participants say about it.

Transcripts suit content analysis well: interview sets, parliamentary debates, broadcast news, podcasts, earnings calls, public meetings and recorded consultations can all be coded the same way.

Building a coding frame

A coding frame (or codebook) lists your categories and the rules for applying them. Each category should have:

  • A short name.
  • A definition that says what the category covers.
  • Inclusion and exclusion rules, especially for borderline cases.
  • One or two typical examples, and an example of something that looks similar but does not belong.

Categories can be deductive, drawn from theory or prior work, or inductive, developed from a first reading of part of the data. Many frames mix both: a deductive starting set refined after coding a pilot sample.

Good frames are mutually exclusive where the design requires it (each unit gets one category on a given dimension), exhaustive (every unit can be coded, even if only as "other"), and specific enough that two coders reach the same decision. If your coders keep debating a category, its definition needs work, not more debate.

Choosing the unit of analysis

The unit of analysis is the piece of text that receives a code. The choice shapes every number you report:

Word or phrase
Useful for frequency of specific terms; loses context
Sentence
Easy to define in edited text; spoken transcripts have unclear sentence boundaries
Speaker turn
Natural for interviews, debates and meetings; turns vary greatly in length
Thematic segment
A stretch about one topic; meaningful but harder to delimit reliably
Whole transcript
Records presence or absence of a category per interview or broadcast

Spoken transcripts complicate sentence-level units, because sentence breaks are often the transcriber's or the software's decision rather than the speaker's. If you code by sentence, standardize how sentences are split before coding. Turns are often more defensible for spoken data, but they require speaker labels, which some automatic transcripts lack.

Separate the coding unit from the context unit: you may code each turn but read the surrounding exchange to decide what it means.

Intercoder reliability

Reliability testing checks whether different coders apply the frame consistently. It is central to quantitative content analysis and increasingly expected in qualitative versions too.

A typical process:

  1. Train coders on the frame using material outside the main sample.
  2. Have two or more coders independently code the same subsample.
  3. Calculate agreement. Simple percentage agreement is easy to understand but does not account for agreement by chance. Chance-corrected measures such as Cohen's kappa, Scott's pi or Krippendorff's alpha are commonly used; Krippendorff's alpha handles more than two coders and missing data.
  4. Discuss disagreements, refine definitions, and repeat on a fresh subsample if agreement was low.
  5. Report the measure, the subsample size and how disagreements were resolved.

Acceptable values are debated, and expectations vary by field and journal; check what is usual in your discipline rather than relying on a single threshold. Low reliability on one category is common and informative: it usually means the category is ambiguous or overlaps with another.

Word-frequency and concordance tools

Computational tools complement manual coding. They are fast and transparent, but they count strings of characters, not meanings.

  • Word-frequency lists show which words occur most often. Stop-word lists remove very common function words, and lemmatization groups forms such as speak, speaks and spoke.
  • Concordance (keyword-in-context) views show every occurrence of a word with surrounding text, so you can check how it is used before counting it.
  • Collocation analysis finds words that occur together more often than chance.
  • Keyword comparison contrasts frequencies between two sets of transcripts, such as two parties or two time periods.
  • Dictionary-based coding counts words from predefined lists for categories such as emotion or certainty.

Free concordancers such as AntConc and the text features of qualitative analysis software cover most needs. Always read concordance lines before trusting a count; the word "cool" in a climate debate and in a youth interview are not the same thing. The article on searching AI transcripts has techniques for finding terms despite spelling variation.

How recognition errors distort counts

Transcripts produced by speech recognition introduce specific biases into content analysis. They are easy to miss because the text looks fluent.

  • Misrecognized key terms. Technical vocabulary, names, local terms and acronyms are often replaced with common words. A term you are counting may be undercounted, and the replacement word overcounted.
  • Inconsistent spellings. The same name or organization may appear in several spellings across files.
  • Dropped words and passages. Quiet speech, crosstalk and some chunk boundaries can lose words; the article on transcripts missing words explains the causes.
  • Inserted text. In long silences or music, recognition models can produce short phrases that nobody said, sometimes repeated. The article on hallucinations in Whisper-style models describes typical patterns.
  • Dropped fillers and hesitations. If you count hedging or disfluency, automatic transcripts will understate it.
  • Number formatting. Numbers may appear as digits in one place and words in another, splitting counts.
  • Uneven error rates. Recognition is typically less accurate for some accents, speakers and recording conditions. That can create apparent differences between groups that are really differences in transcript quality.

The last point deserves emphasis. If one group's recordings were made in noisier settings, or its speakers have accents the model handles less well, its transcripts will contain more errors, and counts will differ for reasons unrelated to what people said.

Hypothetical: council meeting transcripts

A policy researcher analyses two years of recorded city council meetings to see how often housing affordability is raised. Automatic transcripts make the corpus searchable in days rather than months. A concordance check shows that the name of a local housing scheme is recognized in four different spellings and that one meeting with poor audio contains a repeated phrase in a silent stretch. She builds a variant list for key terms, corrects the transcripts for the meetings in her coded sample against the audio, and excludes the inserted text. She reports that counts come from corrected transcripts and lists the variant spellings she merged.

A practical workflow

  1. Define research questions, sample and unit of analysis.
  2. Draft the coding frame, pilot it on a few transcripts, and revise.
  3. Check transcripts: correct at least the sample you will code manually, and check key terms across the whole set if you will count them computationally.
  4. Run a reliability test with a second coder and refine the frame.
  5. Code the full sample, logging decisions on difficult cases.
  6. Run frequency and concordance checks alongside manual coding.
  7. Report the frame, unit, reliability, transcript checking and tools used.

Limits and pitfalls

  • Treating counts as meaning. Frequency shows salience, not importance or attitude.
  • Coding uncorrected machine transcripts and reporting precise frequencies.
  • Changing category definitions mid-coding without recoding earlier material.
  • Comparing groups whose transcripts differ in quality.
  • Forgetting that interviewers shape interview content; a topic raised often may reflect the interview guide rather than participants' priorities.
  • Using sentence counts when sentence boundaries were set by software.

How mydubly fits a content analysis project

mydubly produces the transcripts you analyse, not the analysis. You choose an audio or video file from your device, such as an interview, a recorded meeting or a broadcast you have the right to use, and get a plain transcript, a timestamped transcript and SRT and VTT files, in any of 21 languages with the spoken language detected automatically. The audio to text and video to text pages describe inputs and outputs.

Its output has no speaker labels, so turn-based units need labels added by hand. In the default deployment, a voice activity filter skips silence before recognition, which reduces but does not eliminate inserted text. It does not code, count or search; export the text into your concordancer or analysis software. Check with your ethics board before uploading participant recordings. Transcription costs one credit per minute, so 30 hours of recordings cost 1,800 credits ($1.80).

Next step

Write your coding frame and unit of analysis on one page, pick five transcripts, correct them against the audio, and run a small reliability test with a colleague. The disagreements will show you where definitions need work, and the corrections will show you how much transcript checking the full project needs. For recorded meetings specifically, the meeting transcription use case covers capture and transcription.

Frequently asked questions

What is the difference between content analysis and thematic analysis?

Content analysis applies a defined coding frame systematically and often counts category frequencies, with an emphasis on reliability between coders. Thematic analysis is more interpretive and builds patterns of meaning across data. Many projects combine them, counting topics and then interpreting what is said about them.

How do I calculate intercoder reliability?

Have two or more coders independently code the same subsample, then compute agreement. Percentage agreement is simple but ignores chance; Cohen's kappa suits two coders, and Krippendorff's alpha handles several coders and missing data. Many analysis programs and statistics packages calculate these.

Can I run word-frequency analysis on automatic transcripts?

You can, but check key terms first. Misrecognized words, inconsistent spellings of names, dropped fillers and occasional inserted phrases all distort counts. Build a variant list for terms you count, read concordance lines, and correct transcripts for any sample whose numbers you will report.

What unit of analysis should I use for interview transcripts?

Speaker turns or thematic segments are common for interviews, because sentence boundaries in speech are often the transcriber's choice. Whole-interview units work if you only record whether a category appears. Choose based on your research question and keep it consistent.

Do recognition errors affect groups equally?

Not necessarily. Accents, recording conditions and speaking styles affect accuracy, so one group's transcripts may contain more errors than another's. That can create apparent differences that are really differences in transcript quality. Check samples from each group before comparing counts.

Which tools are used for transcript content analysis?

Qualitative analysis software such as NVivo, ATLAS.ti or MAXQDA supports coding and simple counts. Concordancers such as AntConc handle frequency lists, keyword-in-context views and collocations. Spreadsheets are often enough for tallying codes in smaller projects.