Education & research

Getting Usable Transcripts from Focus Group Recordings

Focus group transcription is harder than transcribing a one-to-one interview because typically six to ten people talk, interrupt and overlap, and the analysis usually depends on knowing who said what. The most effective fixes happen before the session: microphone placement, a moderator who manages turn-taking, and a seating plan or separate audio tracks that make speakers identifiable later. Automatic transcription then gives you a timestamped draft, and a reviewer adds speaker labels, marks overlaps and anonymizes the text before coding.

7 min read · Updated

Why focus groups are harder than interviews

  • More voices, sitting at different distances from the microphone.
  • Overlapping speech: agreement noises, interruptions and sometimes two conversations at once.
  • Similar voices that are hard to tell apart from audio alone.
  • Speaker identity matters to the analysis, for example when comparing participant groups or checking whether one dominant voice drove an apparent consensus.
  • Speech recognition transcribes a single stream. When two people talk at once, the output may capture only the louder one, garble both or drop a phrase.

The general research interview transcription page covers one-to-one interviews; this article is about the problems that only appear when a group is in the room.

A recording setup that saves hours later

  1. Put an omnidirectional boundary or conference microphone in the center of the table, not a phone at one end.
  2. Run a backup recorder at the other end of the table, because batteries and storage cards fail at the worst moments.
  3. For in-person groups, record video as well if your consent forms cover it. Faces resolve most speaker questions during review.
  4. For online groups, check whether your meeting tool can save a separate audio track for each participant. Zoom, for example, offers this for local recordings; confirm the current setting in your tool's documentation.
  5. Seat participants in a known order and write down the seating plan with participant codes.
  6. Ask participants to say their code or pseudonym before speaking in the opening round.
  7. Remove noise sources near the microphone, such as cups, paper shuffling and fans.
  8. Record a short test and listen back before the session starts.

Moderating for a cleaner transcript

  • Open with a round of introductions where each person speaks alone. It gives reviewers a clean voice sample for every participant.
  • Ask people to speak one at a time, and redirect crosstalk politely: "Let's hear from P3 first, then P5."
  • Name the next speaker when handing over, as in "P4, you were nodding. What's your view?" The transcript then records who spoke next.
  • Repeat or summarize quiet contributions so they reach the microphone.
  • Describe relevant reactions aloud, such as "several people are shaking their heads", because a transcript cannot see them.

These habits feel slightly unnatural at first, but they are the cheapest speaker labeling available.

Adding speaker labels when the transcript has none

Many automatic transcripts, mydubly's included, are a single stream of timestamped text with no speaker labels. Automatic speaker diarization exists in other tools, but it struggles with exactly what focus groups produce, namely short turns, similar voices and overlap, so its labels need checking regardless. The speaker diarization explainer covers why. Your practical options are these:

Manual labeling during review
Listen through with the seating plan and introductions in mind, adding P1, P2 or MOD at each turn. Slowest, and the most reliable.
Video-assisted labeling
Watch the recording alongside the transcript; faces settle most ambiguous turns.
Per-participant tracks
Transcribe each participant's track separately, so every line in a transcript belongs to one person, then merge the transcripts by timestamp.
Moderator cues
Names spoken by the moderator at each handover anchor the labels.
Six tracks, one conversation

Suppose an online focus group with five participants and a moderator runs for 90 minutes and is recorded as six separate audio files. Transcribing each file costs 90 credits, so all six cost 540 credits, or 54¢, compared with 90 credits for a single mixed recording. Each transcript gets its participant code in a column, all six go into one spreadsheet, and sorting by start time produces the conversation with every line already attributed. This only works if the tracks share a common start time, so check that first, and expect to delete a few stray lines where one person's microphone picked up another's voice.

Handling overlapping speech

  • Expect overlaps to be the most error-prone passages. Mark them consistently, for example [crosstalk 00:34:10–00:34:18], and recover what you can by listening at reduced speed.
  • Decide in your transcription conventions how many backchannels, such as "mm-hm" and "yeah", to keep. Usually they matter only when they signal agreement relevant to the analysis.
  • When two threads run at once, transcribe the one on topic and note that a side conversation occurred.
  • Per-participant tracks help most here, because each person's speech is clear in their own track even when the mixed recording is a tangle.

Anonymizing before analysis

  1. Replace names with participant codes, including the names participants use for each other.
  2. Remove or generalize other identifiers, such as employers, schools, small towns, unusual job titles and family details.
  3. Store the key that links codes to identities separately and securely, as your ethics approval or data management plan requires.
  4. Use consistent placeholders, such as [employer] or [small town in the region], so the meaning of a sentence survives.
  5. Check file names as well as content; recordings are often saved under a participant's name.

As a final pass, search the transcript for every name on the participant list. People mention each other far more often than moderators expect.

Coding the transcript

Once labeled and anonymized, the transcript is ready for whatever your method calls for, whether that is thematic, framework or content analysis. Qualitative analysis software such as NVivo, ATLAS.ti and MAXQDA imports plain text, and for smaller projects a spreadsheet with columns for start time, speaker, text and codes works well.

  • Keep the timestamps, because they let you return to the audio when tone matters, such as deciding whether a remark was sarcastic.
  • Code individual turns, but also note group dynamics: how agreement built up and who changed their mind.
  • Record who stayed silent on a topic as well as who spoke; silence in a group is data too.

Where automatic transcription helps and where it falls short

Automatic transcription removes the slowest part of the job. A 90-minute session becomes an editable draft in one job instead of a long day of typing, the timestamps make it easy to return to the audio, and the per-minute price makes transcribing separate participant tracks affordable.

  • There are no speaker labels, so attribution is a manual step or comes from per-participant tracks.
  • Overlapping speech is often garbled or partly missing.
  • Quiet or distant participants transcribe less accurately than the moderator.
  • Output leans clean: fillers, false starts and repetitions are often dropped, which is fine for thematic work but means extra effort for conversation analysis that needs strict verbatim.

Running focus group recordings through mydubly

mydubly accepts audio files (WAV, MP3, M4A, AAC, OGG, FLAC) and video files (MP4, MOV, WebM, MKV, M4V) from your device, up to 2 hours each, and in transcript mode returns a timestamped transcript plus SRT and VTT files at 1 credit per minute, with a 5-credit minimum per file. It detects the spoken language automatically and can translate the transcript into another of its 21 languages, which helps when a group is run in a language the analysis team does not share. Transcripts do not label speakers, which is the most important limit for this kind of recording.

For ethics applications, describe the processing accurately: the audio track is uploaded in chunks, uploaded audio and results are deleted within 30 minutes of a job finishing, unfinished jobs expire after 24 hours, and video files never leave the device. Name files with codes rather than names. The audio to text page lists formats, and M4A to text covers recordings from phones and many meeting tools.

Next step for your study

Before your first real session, run a 10-minute mock group with colleagues, record it with your planned setup and transcribe it with audio to text. Label the speakers and time how long it takes; that pilot tells you whether you need per-participant tracks, video or a stricter moderation style. For archival interviews rather than group discussions, see oral history transcription.

Frequently asked questions

How many participants can a focus group transcript handle?

Speech recognition has no participant limit, but every extra voice makes labeling and overlaps harder. Groups at the smaller end of the usual range, clear moderation and per-participant tracks all keep the review manageable.

Can I record a focus group on a phone?

It works in a pinch: put the phone flat in the center of the table, switch on airplane mode to avoid interruptions and make a test recording first. A dedicated boundary microphone picks up distant participants noticeably better.

Should backchannels like mm-hm go in the transcript?

It depends on your method. Conversation analysis keeps them, while thematic analysis usually includes them only when they show agreement that matters. Automatic transcription tends to drop many of them, so add back the ones your conventions require.

How do I merge transcripts from separate participant tracks?

Put each transcript in a spreadsheet with columns for start time, speaker and text, combine them into one sheet and sort by start time. Confirm first that all tracks start at the same moment, otherwise the merged conversation will be out of order.

Can I use an online transcription service under my ethics approval?

That depends on what your approval and consent forms say about third-party processors. Name the service and describe how audio is handled and deleted, and check your institution's data rules before uploading any participant audio.