Why recordings need special care
Audio and video recordings are among the most identifying kinds of research data. A voice can identify a person even when no name is spoken, and interviews often contain details about health, work, family, politics or migration status that participants shared on the understanding they would be protected. Transcripts carry much of the same risk in text form, plus the ease of copying and searching.
Research data management (RDM) is the set of practices that keeps that data safe, usable and, where appropriate, shareable. Most funders and institutions ask for a data management plan (DMP), and many have specific rules for personal and sensitive data. Those rules differ between institutions, countries and funders, and data protection law varies by jurisdiction, so this article describes common practice rather than any particular requirement. Always check your institution's policy and talk to your research data or information governance team.
What a data management plan should say about recordings
A DMP for an interview study usually addresses these items for recordings and transcripts:
- What data you will create: audio or video files, formats, expected volume, transcripts, translations, field notes, consent records.
- How it will be captured and transferred from devices to storage.
- Where each kind of data will be stored during the project, and how it is backed up.
- Who will have access, and how access is controlled.
- Whether, when and how data will be de-identified.
- Which external services will process the data, such as transcription or translation providers.
- How long each version will be kept, and how it will be deleted.
- What will be shared or archived at the end, under what conditions, and in which repository.
- Who is responsible for each of these tasks.
Write the DMP alongside your ethics application, since the two should say the same thing. If your consent form says recordings will be deleted after transcription, your DMP and your practice must match.
Storage and backups
During the project, store recordings and transcripts on storage your institution approves for identifiable research data. That is often a managed network drive or institutional cloud service with access controls and backups, rather than a personal cloud account or a laptop's desktop folder.
Good practice for storage:
- Keep one master copy of each original recording, never edited.
- Work on copies, and name versions clearly: original, corrected transcript, de-identified transcript.
- Follow a backup arrangement that keeps at least two copies in different places, using approved storage for both.
- Encrypt laptops and portable drives that hold identifiable data, as most institutions require.
- Move files off recorders, phones and portable drives onto approved storage promptly, and clear the devices once copies are verified.
A consistent folder structure helps everyone. One common layout separates raw recordings, working transcripts, de-identified data, consent records and the linking key into separate folders, each with its own access rights.
Access control
Limit who can open each kind of data to people who need it:
- Raw recordings and the key linking codes to identities: the smallest possible group, often only the principal investigator and the person who conducted the interviews.
- Identifiable transcripts: researchers and transcribers covered by the ethics approval and confidentiality agreements.
- De-identified transcripts: the wider team, co-coders, and later, if consent allows, other researchers.
Review access when people join or leave the project. Remove shared links that are no longer needed. Avoid sending recordings by email or chat apps; share through approved storage with named access instead.
De-identification
De-identification means removing or altering details that could identify participants. For transcripts, common practice includes:
- Replacing names with pseudonyms or codes, applied consistently across all transcripts.
- Generalizing specific details: "a hospital in the north" instead of the hospital's name, an age range instead of a date of birth.
- Removing or bracketing rare details that could identify someone in combination, such as an unusual job title in a small town.
- Keeping a separate, access-restricted key that links codes to identities, if you need to re-contact participants.
- Marking every change in brackets, so readers know text has been altered: [colleague's name], [town].
Audio is much harder to de-identify than text. Voices are identifying, and bleeping names leaves the voice intact. Many projects therefore keep raw audio under the strictest controls, delete it at an agreed point, and share only de-identified transcripts. Anonymization is a matter of degree: a transcript with names removed may still identify someone to people who know them. Your ethics board and data team can advise on what level is appropriate for your data.
A research team interviews twenty-four teachers in one city about workplace stress. Recordings are copied from the recorder to the university's managed research drive on the day, and the recorder is cleared. Transcripts are drafted, corrected by the interviewer, and then de-identified: names become codes, school names become [primary school], and two participants' unusual subject specialisms are generalized. The linking key is stored in a separate folder that only the principal investigator can open. As the consent form promised, the audio is deleted once transcripts are verified, and the de-identified transcripts are deposited in a repository with access limited to registered researchers.
Retention and deletion
Decide how long each version is kept, and write it down:
- Raw recordings: some projects keep them until transcripts are verified; others keep them for the full retention period because tone and pauses matter for analysis. The consent form should say which.
- Identifiable transcripts and the linking key: often kept until analysis is complete, then deleted or destroyed.
- De-identified transcripts: often kept for the institution's or funder's retention period, which may be several years after publication.
- Consent records: usually kept as long as the data they cover.
Deletion should be real. Remove files from all locations: working folders, backups when they rotate, laptops, portable drives, recorders, email attachments, chat messages and any external service. Record what was deleted and when. If you are unsure about deleting from managed backups, ask your IT or data team how their retention works.
Risks and common mistakes
- Recordings left on personal phones, recorders or laptops long after the interview.
- Sending audio to a transcription service, freelancer or translation tool that the DMP and consent form did not mention.
- Using a personal cloud account for convenience.
- Pseudonyms applied in some transcripts but not others, or the linking key stored next to the data.
- Promising deletion in the consent form and keeping files anyway.
- Planning to share data without consent wording that allows it.
Where mydubly fits in a data management plan
If your DMP and ethics approval allow an external transcription service, mydubly's data handling can be described factually as follows. When you choose a file, the browser decodes the audio locally; for video files, the picture never leaves your device. The audio is cut into chunks of roughly 30 seconds, compressed, and sent over HTTPS for speech recognition and, if requested, translation. Uploaded audio chunks and results are deleted within 30 minutes of a job finishing, and unfinished jobs expire after 24 hours. The privacy policy states that user data is not used for training. Results are downloaded to your device, where your own storage rules take over. The private video translation page describes this processing model, and the article on cloud speech processing privacy lists questions to ask any provider.
This is not a statement that mydubly meets any regulation, ethics standard or institutional policy. Whether it is acceptable for your data is for your institution, ethics board and data protection team to decide; give them these facts and the privacy policy. mydubly does not store, organize or archive recordings for you, and it does not de-identify transcripts.
Next step
Draft the recordings section of your DMP using the checklist above, then compare it line by line with your consent form and ethics application. Ask your research data team to review it before fieldwork. When transcripts start arriving, the interview transcription use case shows a simple workflow that can sit inside the plan, and the explainer on what stays on your device helps when describing browser-based processing to a reviewer.
Frequently asked questions
Do I need a data management plan for an interview study?
Many funders and institutions require one, and even when they do not, it is good practice for recorded interviews. It records where data lives, who can access it, how it is de-identified, how long it is kept and what may be shared. Ask your institution's research data team for their template.
Should I keep the original recordings after transcription?
It depends on your analysis and consent. Keeping audio lets you check tone and verify quotes, but it is the most identifying form of data. Some projects delete audio once transcripts are verified; others keep it securely for the full retention period. Whatever you choose must match what participants were told.
Is replacing names enough to anonymize a transcript?
Often not. Combinations of details such as job, location, age and specific events can identify someone, especially to people who know them. Generalize or remove rare details as well, mark changes in brackets, and keep the linking key separate. Your ethics board can advise on the level needed.
Can I share interview transcripts with other researchers?
Only if participants consented to sharing and the data is suitably de-identified. Sharing usually happens through a repository with controlled access rather than openly. If your consent form did not mention sharing, consult your ethics board before doing anything.
How long does mydubly keep uploaded audio?
Uploaded audio chunks and results are deleted within 30 minutes of a job finishing, and unfinished jobs expire after 24 hours. For video files, the picture is not uploaded. Results you download are stored on your own device and fall under your own data management arrangements.
Does using a transcription service need ethics approval?
Ethics boards often want to know about any third party that processes participant data, and participants may need to be told. Requirements differ between institutions, so check with your ethics board before sending recordings to any service, including mydubly.