Business video

Working With User Research Sessions Recorded in Other Languages

To translate user research interviews, get consent that covers transcription and translation, transcribe the recording in its original language, produce a translated transcript for analysis, and remove personal details before sharing. Do your coding and synthesis on the translation, but go back to the original audio with a fluent colleague to verify every quote that will appear in a report or influence a decision.

7 min read · Updated

Where translation fits in a multilingual study

International studies are usually run in one of three ways: a native-speaking moderator runs sessions in the participant's language, a researcher moderates with a live interpreter, or participants who speak a shared language take part in that language. Only the first two produce recordings the core team cannot follow directly, and they produce different material.

With a native-speaking moderator, the recording is entirely in one language and needs a full translation for the team. With an interpreter, the recording contains both languages, often overlapping, and the interpreter's version is already a summary filtered through one person. Translating the participant's original words afterward lets you check what the interpreter conveyed, which matters when a finding rests on nuance.

Either way, translation turns hours of recordings into text the whole team can read, search and code. The task is to do that without losing the participant's meaning or exposing their identity.

Anonymization before anything is shared

Session recordings contain personal data: names, employers, locations, account details read aloud, faces and screens. A good default is to keep raw recordings in a restricted location and share only cleaned transcripts.

  • Assign participant codes such as P1 and P2, and replace names in transcripts, including names of family members and colleagues mentioned in passing.
  • Remove or generalize identifying details: specific employers, small towns, unusual job titles.
  • Check translated transcripts separately. A name the speech recognizer spelled oddly can slip past a search for the correct spelling.
  • If the session was a screen recording, remember the picture may show email addresses or account data even when the audio does not.

Where possible, export only the audio for transcription and translation, which keeps faces and screens out of the process entirely.

Analyzing translated transcripts

Translated transcripts make synthesis possible for the whole team, but they need some preparation and a few habits.

Machine transcripts do not say who is speaking. Before coding, mark the moderator's questions and the participant's answers, which is quick in a one-to-one session with a clear question and answer pattern. The article on speaker diarization explains why this step is still manual in many tools. For group sessions, the separate article on focus group transcription covers multi-speaker material.

Keep the original-language transcript alongside the translation, with timestamps, so any line can be traced back to the recording. Code themes on the translation, but treat hedging, politeness and intensity with care: translation tends to flatten them. A participant who says, in effect, "it is perhaps a little difficult" may be expressing strong frustration in a culture where direct criticism of a product is unusual. A native-speaking researcher or moderator should review how such statements are coded.

Be cautious about counting. Words and phrases do not map one to one between languages, so frequency counts across translated transcripts in different languages are rarely comparable.

Verifying key quotes before they reach a report

Quotes carry weight in research reports, and a mistranslated quote can push a product decision in the wrong direction. Before a translated quote goes into a report or presentation:

  1. Find the timestamp and listen to the original audio around it, including the question that prompted it.
  2. Ask a fluent colleague, the moderator or a professional translator to confirm the translation and the tone.
  3. Check that removing surrounding context does not change the meaning.
  4. Mark it clearly as translated, for example "P4, translated from Japanese," so readers know it is not verbatim.
  5. Remove any identifying details from the quote itself.

Transcripts for analysis are often cleaned of filler words, which is usually fine; the trade-offs are covered in verbatim vs clean verbatim transcription.

Example: onboarding usability tests in Japan and Germany

A product team at a payments app runs ten moderated usability sessions on its onboarding flow, five in Japanese and five in German, each led by a native-speaking moderator and about an hour long. The research lead exports the audio of each session, produces transcripts and English translated transcripts, and replaces participant names with codes. The team tags moments of confusion in the English translations. Before the readout, the Japanese moderator reviews every Japanese quote selected for the report and corrects two where the translation made a polite suggestion sound like neutral approval, which changes how severe one usability issue looks.

A workflow from recording to report

  1. Confirm that consent covers transcription, translation and the tools you plan to use.
  2. Store raw recordings in a restricted location and export the audio track for processing.
  3. Produce a transcript in the original language and a translated transcript for each session.
  4. Mark moderator and participant turns, and replace personal details with codes.
  5. Share only the cleaned transcripts with the wider team for coding and synthesis.
  6. Have a native-speaking moderator or researcher review coding of nuanced or culturally loaded statements.
  7. Verify every quote that will appear in a report against the original audio.
  8. Delete recordings and transcripts according to your retention policy and consent terms.

Pitfalls and limits of machine-translated research data

  • Think-aloud sessions often contain quiet, fragmented speech that is harder to recognize; see transcribing quiet audio for recording tips.
  • Overlapping speech between moderator and participant, or with an interpreter, reduces transcript accuracy in the overlapping parts.
  • Product names and interface labels may be misheard or translated when they should stay as they appear in your product.
  • Sessions where participants mix languages, such as English product terms inside Japanese sentences, can confuse recognition and translation.
  • Machine translation cannot tell you what a participant meant beyond their words. Cultural interpretation remains the job of researchers who know the context.

How mydubly fits a research workflow

mydubly takes a recorded audio or video file uploaded in the browser and returns a transcript in the spoken language, with a timestamped version, and optionally a translated transcript, in 21 languages. It also produces subtitles as SRT or VTT, which some teams find handy for watching a session clip with translated subtitles. The spoken language is detected automatically. See the audio translator for audio files and audio to text for transcription alone.

It does not label speakers, so marking moderator and participant turns stays manual. It does not anonymize transcripts, and it processes recorded files only, not live sessions. Files can be up to two hours long. For privacy, only compressed audio chunks are sent over HTTPS, a video file itself stays on your device, results are deleted within 30 minutes of completion, and the privacy policy says user data is not used for training. Whether that fits your consent terms and data policy is for your team to decide. As a cost example, a 60-minute session costs 60 credits (6¢) for the transcript and translated transcript, so ten such sessions cost 600 credits (60¢).

Your next study

Before your next international round, update the consent form to cover transcription and translation, agree who will verify quotes in each language, and decide where raw recordings will live. Then process one pilot session end to end, from audio export to verified quote, to check the workflow before the full study. The research interview transcription page and the audio translator describe the transcription and translation outputs.

Frequently asked questions

Should the moderator speak the participant's language?

Usually yes. A native-speaking moderator builds rapport, follows up on subtle cues and keeps the conversation natural, which tends to produce richer data than sessions run through an interpreter. Translation then happens afterward, on the recording. Interpreted sessions can work for observing tasks, but the participant's words pass through one person live, so check important moments against the original audio later.

Can I rely on machine-translated transcripts for synthesis?

For finding themes, patterns and moments worth revisiting, a machine-translated transcript is usually workable, provided audio quality is reasonable. It becomes risky for nuance, intensity and exact wording. Have a native speaker review coding of culturally loaded statements, and verify every quote used in a report against the original audio. Treat findings that rest on a single translated phrase with extra caution.

How should translated quotes be presented in reports?

Label them clearly, for example "P3, translated from German," and keep the original-language text in an appendix or linked notes for anyone who reads the language. Avoid polishing the translation so much that a hesitant participant sounds articulate. Keep identifying details out of the quote, and include enough context, such as the task being performed, for readers to interpret it fairly.

What about sessions where the participant switches languages?

Code-switching is common, especially with product or technical terms. Recognition and translation handle long stretches of one language better than frequent switching, so expect more errors at switch points. Note in the transcript where switching happens, and have a bilingual reviewer check those passages. The switches themselves can be a finding, for example if participants consistently use English terms for certain features.

How long should we keep research recordings and translations?

Keep them only as long as your consent terms and retention policy allow, and no longer than the research needs. Many teams keep cleaned, anonymized transcripts longer than raw recordings, since they carry less personal data. Write the retention period into the consent form, set a reminder to delete, and make sure copies in shared drives and analysis tools are deleted too.