Why event recordings are harder than studio recordings
In a studio, the microphone hears the speaker. At an event, a camera or recorder in the room typically hears the speaker through a public address system: the voice comes out of loudspeakers, bounces off walls, floor and ceiling, and arrives a fraction of a second later from several directions at once. Add a hundred people shifting in chairs, coughing, laughing and applauding, and the speech the recognizer needs is only part of what was captured.
Events also bring their own structure. Several speakers take turns at a lectern or share a panel table. Audience members ask questions from the floor, often without a microphone. Music plays at the start, during breaks and for walk-ons. And there are long stretches with no speech at all. Each of these needs a decision before transcription, not after. Recordings of streamed events that already exist as stream replays are a different case, covered in translating livestream recordings.
The sound desk feed versus the camera microphone
Where you take the audio from decides most of the transcript's quality.
- Sound desk feed
- A cable from the mixing desk's record or auxiliary output into a recorder or camera. Carries the microphones directly, with no room echo. It is usually the cleanest source for amplified speech.
- Camera's built-in microphone
- Hears the PA, the room's reverberation and the audience. Fine as a backup, usually poor as the only source in a large or echoey room.
- Recorder near a loudspeaker
- Louder than the camera microphone but prone to distortion and still full of room sound.
- Lavalier recorded on the presenter
- A small recorder worn by the speaker captures them very cleanly, but only them, and depends on someone remembering to start and stop it.
- Room or audience microphone
- Captures questions, laughter and applause that the desk feed misses. Most useful as a second track alongside the desk feed.
The desk feed has its own pitfalls. It carries only what goes through the desk, so an unamplified speaker in a small room, an audience member speaking without a microphone or a laugh from the crowd won't be on it. The desk is mixed for the room, not the recording: in a small venue, the engineer may keep the lectern microphone low because the room doesn't need much reinforcement, which leaves the feed quiet. Ask for a dedicated record send, set to a sensible level and independent of the room mix if the desk allows it, and test it before the doors open.
Connecting a recorder to a desk can introduce mains hum through a ground loop, especially when the camera and desk are plugged into different circuits. An isolation transformer or a direct box usually fixes it; audio hum and buzz explains the cause and other cures.
Recording two sources and combining them
The most reliable setup records the desk feed and a room microphone at the same time, on separate tracks or separate devices. The desk feed provides the clean speech; the room track provides audience questions and reactions and acts as a backup if the feed fails. In editing, line the tracks up on a clap or a sharp sound, keep the desk feed as the main source, and bring up the room track only where needed, such as during floor questions.
Blending both tracks at full level throughout brings the PA echo back and can cause comb filtering, because the same voice arrives at slightly different times. How to combine several microphones without those problems is covered in recording with multiple microphones. Export the finished mix as the only audio track of the file you transcribe.
Applause, music and long gaps
Applause is loud, broadband noise with no speech in it, and music is structured sound that isn't speech. Recognizers handle both imperfectly. A burst of applause usually produces nothing, but sometimes a recognizer writes a short phrase that nobody said; Whisper-style models are known for producing stock phrases such as a thank-you during non-speech passages. Long instrumental walk-on music or a 15-minute break can produce stray lines. The behavior is explained in Whisper hallucinations.
Before transcribing, trim the pre-event music, the breaks and the post-event chatter. Keep short applause in place, since cutting it would shift timestamps against the video, but expect to delete the odd stray line during review. If you need timestamps to match the full recording, trim only from the beginning and end and note the offset.
A worked example: a town council meeting
A hypothetical town council records its 150-minute meeting with a camera at the back of the chamber and a cable from the sound desk carrying the councilors' microphones and the public comment microphone. The camera audio is echoey, so the clerk uses the desk feed, laid against the camera picture in an editor, with the camera audio brought up only for two residents who spoke from their seats. The file is split at the recess into two parts of 75 minutes, which costs 150 credits (15¢) to transcribe. The clerk then checks each speaker's name against the sign-in sheet and the agenda, and checks every motion and vote count against her notes.
That last step matters most. For public meetings, an AI transcript is a working draft that helps prepare minutes; it isn't an official record. The rules on what counts as the official record differ by jurisdiction, so check yours, and treat this as general information rather than legal advice.
Event recording and transcription, step by step
- Before the event, ask the sound engineer for a record feed and confirm the connector type, level and whether the room microphones are all routed to it.
- Set up a backup room recorder or camera microphone, with fresh batteries and enough storage for the full event plus overruns.
- Collect the program, speaker list, slides and any list of names, such as award recipients or graduates.
- Ask the moderator to repeat audience questions into a microphone, or use a roving microphone that goes into the desk.
- Record continuously from before the first speaker to after the last, and make a clap or announce the session at the start of each recording for syncing.
- After the event, line up the feed and backup, choose the desk feed as the main source, and bring in the room track only where needed.
- Trim music, breaks and chatter from the start and end; split any recording longer than 2 hours at a natural break such as an intermission.
- Transcribe each session, then correct names, titles and organization names against your collected materials, and add speaker names by hand.
Trade-offs and risks with event audio
- The desk feed misses everything that wasn't miked. Without a room track, floor questions and audience reactions can disappear entirely from the recording.
- A single camera microphone in a large hall can produce a transcript that is only a rough guide, especially for quieter speakers and panelists who sit far from their microphones.
- Names are the most frequent error. Ceremonies that read long lists of names will produce misspellings even with clean audio, because many names can be spelled several ways; the program, not the transcript, is the source of truth.
- Panels bring crosstalk. When panelists interrupt each other, expect gaps or merged sentences in those passages.
- Recording people at an event carries obligations. Tell attendees that sessions are recorded and transcribed, and follow venue rules and any speaker agreements.
Where mydubly fits for event recordings
mydubly works on recordings after the event; it has no live captioning or real-time translation. Upload the finished file as MP4, MOV, WebM, MKV or M4V to video to text, or an audio-only mix as WAV, MP3, M4A, AAC, OGG or FLAC to audio to text. Each file can be up to 2 hours long. If a file has several audio tracks only one is used, and stereo channels are combined when the browser converts the audio to 16 kHz mono, so a camera that recorded the desk feed on one channel and its own microphone on the other gives you a blend of both. Export the mix you want as the only track.
The browser sends only compressed audio chunks over HTTPS for transcription with Whisper; the video stays on your device. You get a plain transcript, a timestamped transcript and SRT and VTT subtitles, and optionally a translated transcript. There are no speaker labels, so panel and meeting transcripts need names added during review. Translating a recorded talk for another audience, with subtitles or a dubbed voice, is covered on the conference talk translation page; services and worship recordings have their own sermon translation page.
Next step: plan the audio before the next event
Contact the venue's sound engineer a week before your next event, agree on a record feed and test it during setup, with a backup recorder running in the room. For an event you've already recorded, compare the desk feed and the camera audio on the same five minutes, transcribe whichever sounds cleaner, and check names against the program.
Frequently asked questions
What cable do I need to record from a sound desk?
It depends on the desk's outputs and your recorder's inputs. Desks commonly offer XLR or quarter-inch outputs for auxiliary sends, and sometimes RCA record outputs or a USB connection. Recorders and cameras may accept XLR, quarter-inch or a 3.5 mm input. Ask the venue for the output type in advance, bring adapters, and confirm the level is line or microphone level so you set your recorder's input to match.
Can I record a conference session on my phone from the audience?
You can, and it is better than nothing, but expect a reverberant recording with audience noise close to you. Sit near the front, centered between the loudspeakers, and avoid placing the phone on a table that people bump. If the session matters, ask the organizers whether they record a desk feed and whether they share it with speakers or attendees.
Should I transcribe each speaker's session as a separate file?
Usually, yes. Splitting by session or speaker keeps files under the 2-hour limit, makes review easier to divide among people, and lets you publish or translate each talk separately. Split at the applause or changeover between speakers, where there is no speech to cut through, and name files by event, date and speaker.
How do I handle a panel where everyone shares one microphone?
Expect uneven levels, because whoever holds the microphone closest is loudest and others fade in and out. During review, add speaker names by recognizing voices and checking the video. For future panels, ask for one microphone per panelist routed to the desk, or at least a microphone that the moderator passes deliberately, so only one person speaks into it at a time.
Can attendees get a translated transcript of a session?
Yes, if you have the recording and the right to share it. Choose a target language when transcribing and you receive the original transcript and a translated one in the same job, at 1 credit per minute. Have a fluent reader check names and key terms before distributing it, since recognition errors in the original carry into the translation.