About WAV files
WAV is uncompressed audio — the default for professional recorders, audio interfaces and studio software. Files are large but lossless.
Common sources: field recorders (Zoom, Tascam), studio and DAW exports, research and interview recorders.
Every browser reads standard PCM WAV.
WAV to text
WAV files are large, but only compressed chunks of the audio are uploaded, so you don't need to convert first.
A qualitative researcher transcribes a 70-minute WAV interview captured on a field recorder.
Cost: 70 credits (7.0¢).
Translating a WAV
The translated voice comes back as M4A; import it into your editor alongside the original WAV.
An audio producer translates a 12-minute WAV narration into German for a museum installation.
Cost: 600 credits (60¢) with a translated voice, or 12 credits (1.2¢) for translated text only. Switch Full translation output on in the tool above for a voice track.
Why WAV files are huge, and why it doesn't matter
WAV stores sound as uncompressed PCM samples, so its size is simple arithmetic: sample rate × bit depth × channels. A stereo file at 48 kHz and 24-bit takes roughly 1 GB per hour; a mono 44.1 kHz 16-bit file is about a third of that. None of that weight is uploaded as-is, so a multi-gigabyte interview doesn't need converting to MP3 first.
Bigger numbers on the recorder don't make a better transcript. Recognizers work at a resolution suited to speech, so 96 kHz or 32-bit float buys you editing headroom, not accuracy. Audio bit depth and sample rate explain what those settings are really for.
Field recorder quirks
Dedicated recorders add complications you won't meet with phone audio:
- Split files. A standard WAV can't exceed 4 GB, so many recorders start a new file partway through a long take and number them in sequence. Join the parts in an audio editor before uploading, or upload each one and remember that every part's times begin from zero.
- Polyphonic WAV. Multitrack recorders can write every input into one file with four, six or more channels, some silent and some holding a safety copy recorded at lower gain. Mix down to the channels that carry speech (often one lavalier) in your DAW, so unused inputs and duplicate safety channels don't shape what gets transcribed.
- RF64 and 32-bit float. Very long recordings may use the RF64 extension, and some recorders write 32-bit float. If your browser refuses a file, export a standard 16- or 24-bit PCM WAV.
- Timecode. Broadcast WAV (BWF) files store the recorder's start timecode in a metadata chunk. Transcript and subtitle times count from the beginning of the file instead, so add the start time yourself when you line text up with camera footage.
An interview workflow that keeps the original safe
WAV is the default for researchers and journalists recording on dedicated recorders, and the file often becomes evidence or research data. Work from a copy and keep the original untouched:
- Copy the card to two places and rename the file with the date and an interviewee code before doing anything else.
- Listen to the first and last minute on headphones to catch a wrong input or clipping.
- If each person had their own lavalier channel, decide between one mixed file (one transcript) and one file per person (separate transcripts you interleave by time to attribute lines by hand). Per-person channels pick up some bleed from the other voice, so expect a few of their words too.
- Upload the WAV, download the timestamped transcript, and check names, numbers and technical terms against the audio.
- Store the transcript beside the WAV under the same base name.
Recording fieldwork interviews covers the capture side, and thematic analysis of transcripts covers what happens after.
Level problems a lossless file won't fix
Lossless means the file keeps exactly what the microphone and preamp delivered, faults included. Three problems show up repeatedly in recorder WAVs:
- Clipping. If the input gain was too high, peaks are flattened and speech distorts on loud words. A 32-bit float file can't clip internally, but the preamp in front of it still can, so check the waveform for flat-topped peaks. Declipping tools help a little; a cleaner transcript usually needs a re-recorded take or patient review.
- Hiss from low gain. A recording made far too quiet and raised afterwards brings the noise floor up with the voice. Speech still transcribes, but quiet words near the noise get lost first.
- Handling and cable noise. Thumps from a recorder held in the hand, or rustle from a lavalier under clothing, land on top of words. Recognition can't separate them, so note those stretches for listening during review.
Costing a set of recordings
Twelve 55-minute WAV interviews come to 660 credits (66¢) for transcripts, timestamps and subtitles. A separate 7-minute studio narration dubbed into French costs 350 credits (35¢) and arrives as an M4A that follows the original WAV's timeline, with any room sound or music from the recording kept under the French voice.
For the full interview process, including consent and review, see interview transcription. If your recordings are compressed rather than uncompressed, the MP3 page covers how bitrate affects results.
Using the tool on this page
- The spoken language is detected automatically; you select the output language — the spoken language itself for an untranslated transcript, or another language for a translation. Files can be up to 2 hours long.
- Audio is sent over HTTPS in small chunks and deleted with its transcripts within 30 minutes of delivery (how files are protected).
- You can download a timestamped transcript, subtitles (SRT and VTT) and plain text. Every output, step and limit is explained on the audio to text page.
Frequently asked questions
Does a WAV give better transcripts than an MP3?
A clean recording matters more than the format, but WAV avoids compression artifacts, so it is an excellent starting point.
Can I get the translated audio back as WAV?
The translated voice is delivered as M4A. Convert it to WAV in any audio editor if your workflow requires it.
The interviewer is much louder than the guest in my WAV. Will that hurt the transcript?
Quiet speech is more likely to be missed or misheard. Raise the quieter channel, or apply gentle normalisation, before uploading. Avoid heavy noise reduction, which can smear consonants and make recognition worse rather than better.
Should I archive my WAVs as FLAC instead?
For storage, FLAC is lossless and noticeably smaller, and it opens in most audio tools. FLAC is also accepted for transcription, so you can keep the archive in FLAC and upload the same files.
Does a 96 kHz or 32-bit WAV cost more to transcribe?
No. Billing is by length only: 1 credit per minute for a transcript, whatever the sample rate, bit depth or file size.