What makes a collection of recordings a corpus
A folder of interview recordings is not yet a corpus. What distinguishes a corpus is design: a clear statement of what language variety or situation it represents, a sampling plan that follows from that statement, consistent transcription and annotation, and documentation that lets other people use it correctly.
Well-known spoken corpora, such as the spoken components of large national corpora or collections like the Santa Barbara Corpus of Spoken American English and the CHILDES child language archive, differ enormously in size and purpose. What they share is that a user can find out who was recorded, in what situation, how the speech was transcribed and what the symbols mean. That is the standard to aim for, whatever the scale of your project.
Design and sampling
Start with the research questions the corpus should support, then work out what it needs to contain:
- Population: whose speech? Speakers of a regional dialect, learners of a language, children at certain ages, a profession, a bilingual community.
- Situations and genres: conversation, interviews, narratives, classroom talk, service encounters, broadcast speech. Different situations produce different language.
- Balance: how many speakers per group, how much speech per speaker, how many hours per genre. Balance matters if you will compare groups.
- Size: enough to find the features you care about. Frequent features such as common function words appear in small corpora; rare constructions need far more data.
- Time frame: a snapshot, or repeated recordings to track change.
Spontaneous conversation is the hardest data to collect and the most valuable for many questions. Interviews are easier but produce a particular, interviewer-shaped kind of talk. Be explicit about the trade-off in your documentation.
Plan transcription effort alongside recording. A corpus with many hours of audio and too few hours of transcription time ends up with a small transcribed core and a large untouched archive. Pilot a few files to estimate how long transcription to your standard really takes.
Consent for a corpus
Corpus consent differs from consent for a single study, because the data is meant to be reused, often by people you will never meet. Consent wording usually needs to cover:
- Recording and transcription, including any use of external transcription services.
- Use by the research team and by other researchers in the future.
- Sharing through an archive or repository, and under what access conditions.
- Whether audio, transcripts or both will be shared.
- How speakers will be identified or anonymized, and whether voices themselves will be available.
- Whether speakers can withdraw, and until when.
Conversations often include people who did not sign up, such as family members or bystanders. Decide in advance how you will handle their speech. For community-based work, especially with Indigenous or minority language communities, consent may need to be negotiated with the community as well as individuals, and community expectations about access may differ from academic norms. Check with your ethics board and, where relevant, community representatives before recording.
Speaker and recording metadata
Metadata turns recordings into analyzable data. Without it, you cannot compare groups or explain variation. Record it at the time of collection using a fixed form, and store it in a structured file linked to recording IDs.
- Speaker ID
- A code used in all transcripts and files, never the real name
- Demographics
- Age or age band, gender as self-described, region of origin and residence
- Language background
- First language or languages, other languages and proficiency, age of acquisition
- Education and occupation
- If relevant to your questions, in categories you define in advance
- Relationships
- How speakers in a recording know each other
- Recording situation
- Date, place, setting, genre, who was present, interviewer if any
- Technical
- Recorder, microphone, format, sample rate, duration, file name
Collect only what your questions need and your consent covers. Fine-grained demographics can make speakers identifiable in small communities, which matters if the corpus will be shared.
Transcription standards
Every corpus needs a written transcription guide that all transcribers follow. Key decisions include:
- Orthographic or phonetic. Most corpora use standard orthography with conventions for nonstandard forms, and add phonetic transcription only for selected material.
- Spelling of dialect and colloquial forms. Will you write gonna or going to? A regional pronunciation as standard spelling or as heard? Consistency matters more than the choice, because these decisions directly affect frequency counts.
- Fillers, false starts and repetitions. Corpora studying spontaneous speech usually keep them all. The comparison of verbatim and clean verbatim transcription explains what each style keeps.
- Overlaps, pauses and non-verbal events, using a defined notation. The article on transcription conventions covers notation systems in more depth.
- Code-switching: how to mark stretches in another language.
- Unclear speech and anonymized items, marked consistently.
- Segmentation: by utterance, intonation unit, turn or time-based segments.
Test the guide by having two transcribers do the same short file and comparing results. Disagreements show where the guide needs to be clearer.
Annotation layers
Annotation adds information on top of the transcript. Common layers in spoken corpora include:
- Time alignment: linking each utterance, or each word, to its position in the audio. Forced alignment tools can align words automatically given an accurate transcript.
- Speaker turns and overlaps.
- Part-of-speech tags and lemmas, often produced automatically and checked.
- Prosodic annotation: stress, intonation, pauses.
- Phonetic or phonological annotation for selected features.
- Pragmatic or discourse annotation, such as speech acts or discourse markers.
- Translation or glosses, for corpora in languages your users may not read.
Keep layers separate rather than mixing them into one line of text, so users can work with the layer they need and you can update one without disturbing others. Multi-tier annotation tools such as ELAN, and phonetic tools such as Praat, are built for this.
A sociolinguistics group sets out to build a corpus of conversational speech from a coastal town, aiming for forty speakers balanced across three age groups and two genders, with an hour of conversation each. They record pairs of friends talking with a lavalier microphone on each speaker, store WAV masters, and collect metadata on a paper form scanned into a spreadsheet. A machine draft gives them a starting text for each file, but their transcription guide requires local dialect words to be written as pronounced, so correctors change many standardized forms back and log which files started from machine drafts. Time-aligned transcripts are stored in ELAN with tiers for each speaker, and part-of-speech tags are added automatically and spot-checked.
File formats for the long term
Choose formats that will still open in decades and that other tools can read:
- Audio masters: uncompressed WAV is the common archival choice. Keep compressed copies, such as MP3, for convenience, but never as the only copy.
- Transcripts: plain text in UTF-8 for maximum portability, plus structured formats where annotation needs them.
- Structured annotation: XML formats such as TEI, ELAN's EAF files, Praat TextGrids, or CHAT files for the CLAN tools used with CHILDES. Choose the one your field and your intended archive support.
- Metadata: a spreadsheet exported to CSV, or a standard metadata schema if your archive requests one.
- Documentation: the design rationale, transcription guide, annotation guidelines, metadata codebook and licence, as plain text or PDF.
Use consistent file names linking audio, transcript, annotation and metadata through the recording ID.
Using machine drafts responsibly
Speech recognition can produce a first draft quickly, but corpus linguists need to understand what it does to the data:
- It standardizes. Recognition models trained on large amounts of web audio tend to output standard spellings for dialect forms and may smooth grammar toward the standard variety, which is exactly what many corpora study. The article on accents in speech recognition explains why.
- It drops disfluencies. Fillers, repetitions and false starts are often omitted.
- It punctuates by grammar, not prosody.
- It varies by language. Quality differs between languages and is weaker for languages and varieties with less training data; the explainer on multilingual speech recognition covers why.
- It can insert words that were not said, particularly in silence or noise.
If you use machine drafts, correct every file fully to your transcription guide, listening to the whole recording. Record in the metadata which files started from a machine draft, which tool and version, and who corrected them. Never publish uncorrected machine output as part of a corpus, and avoid running frequency analyses on uncorrected text.
Limits of the approach
Corpora are expensive in time. Transcription and annotation typically take many times longer than recording, and checking automatic annotation adds more. A corpus also only represents what was sampled: a corpus of interviews says little about casual conversation, and a corpus of one town says little about a whole region. State these limits in the documentation so users do not over-generalize.
Where mydubly can help, and where it cannot
mydubly can produce draft transcripts from corpus recordings. You choose a file from your device, such as a WAV master, MP3 or a video recording, up to two hours each, and receive a plain transcript, a timestamped transcript with [m:ss] labels and SRT and VTT files with one cue per recognized segment. It supports 21 languages and detects the spoken language automatically; many minority languages and dialects are not among them. The audio to text page lists the details.
For corpus work, its limits are significant. Output follows standard spelling and tends to drop disfluencies. There are no speaker labels, no word-level timestamps, no overlap notation and no annotation. Mixed-language recordings may be handled inconsistently. Treat it as a typing aid for a transcriber who corrects to your guide, and check with your ethics board that external processing is covered by your consent. One hour of audio costs 60 credits (6¢).
Next step
Write a two-page corpus design document covering population, situations, sampling targets, consent, metadata fields, transcription standard and formats, then pilot it on three recordings, including full transcription and one annotation layer. Time each step. If your aim is vocabulary study for language learners rather than research, the article on building vocabulary lists from transcripts is a better fit.
Frequently asked questions
How big does a spoken corpus need to be?
It depends on the features you study. Frequent words and constructions can be studied in a modest corpus, while rare structures need much more data. Pilot your research questions on a small sample to see how often the features appear, then plan size from there rather than aiming for an arbitrary number of hours.
Should I transcribe dialect words in standard spelling?
Decide based on your research questions and apply the rule consistently. Many sociolinguistic corpora write certain dialect forms as pronounced, often with a standard-spelling layer for searching. Whatever you choose, document it in the transcription guide, because it directly affects word counts and searches.
Can I build a corpus from automatic transcripts?
Only if every file is corrected to your transcription standard by someone listening to the audio. Machine transcripts tend to standardize spelling, drop disfluencies and occasionally insert words, all of which distort corpus counts. Record which files began as machine drafts and how they were corrected.
What format should corpus audio be stored in?
Uncompressed WAV is the usual archival master because it avoids compression artifacts and opens in nearly every tool. Compressed copies such as MP3 are convenient for listening but should not be the only copy. Ask your intended archive whether it has preferred formats.
What metadata should I collect about speakers?
Collect what your research questions need and your consent covers, typically a speaker code, age band, gender, regional background, language background, and the recording situation. Avoid collecting detail you will not use, since fine-grained information can make speakers identifiable in small communities.
Which tools are used for annotating spoken corpora?
Common choices include ELAN for multi-tier, time-aligned annotation, Praat for phonetic work and the CLAN tools for CHAT-format transcripts. Part-of-speech taggers and forced aligners add automatic layers. Choose tools that export open formats your archive accepts.