Troubleshooting

Twenty-Five or 25? Making Numbers Consistent in AI Transcripts

Speech recognition has to decide how to write every number it hears, and spoken numbers are ambiguous: "one twenty" might be 120, 1:20 or a price, and "twenty twenty-five" might be a year or an amount. Models learned from transcripts written in many styles, so the same file can contain "25", "twenty-five" and "25.00". Fix it by choosing a short house style, searching the transcript for every digit and number word, and checking anything consequential, such as prices, dates and quantities, against the audio.

8 min read · Updated

What inconsistent numbers look like

Number problems are usually formatting problems rather than recognition errors, but both turn up together, and they need different handling.

  • Words and digits mixed. "Twenty-five people" in one sentence, "25 people" in the next.
  • Dates in several shapes. "March 3rd", "3 March", "the third of March" and "3/3" in the same document.
  • Times that drift between styles. "3 pm", "three o'clock", "15:00" and "fifteen hundred".
  • Money written loosely. "$5 million", "five million dollars", "5M" and "five mil".
  • Units spelled out or abbreviated. "Ten kilometres", "10 km" and "10k", which could also mean ten thousand.
  • Phone numbers and codes split up. A phone number or order reference broken into separate groups, partly in words, sometimes with stray full stops between the groups.
  • Version numbers and decimals. "Version two point one" instead of "v2.1", or "two and a half" next to "2.5".

Separately, sound-alike numbers can be misheard outright: fifteen and fifty, thirteen and thirty, or "a hundred and five" heard as "105" when the speaker said "a hundred and fifty". Those are recognition errors and can only be settled by listening.

Why recognition output varies

Traditional speech recognition systems first produced spoken-form words ("twenty five dollars") and then ran a separate step, often called inverse text normalisation, that converted them into written form ("$25") using rules. End-to-end models such as Whisper skip that step and write numbers directly, choosing whichever form looks most likely given the words around it and the transcripts they were trained on. Those training transcripts followed many different conventions, so the model has no single style to imitate.

Context steers the choice. In a sentence about money, digits and a currency symbol are likely; in casual conversation, words are common. When context is thin, as at the start of a chunk of audio or in a short reply, the model has less to go on and the style can flip. The explainer on how speech to text works covers how tokens and decoding produce this behaviour.

Spoken number systems add their own ambiguity:

  • English speakers say years in pairs ("nineteen eighty-four") and can say amounts the same way ("fifteen hundred").
  • French counts some tens by twenties (quatre-vingt-dix is ninety), and German says the units before the tens (fünfundzwanzig, literally five-and-twenty).
  • Indian English and Hindi often group large numbers in lakhs and crores, where a lakh is 100,000 and a crore is 10 million.
  • Japanese and Chinese group large numbers in units of ten thousand rather than thousands.
  • Many European languages use a comma as the decimal separator and a full stop or space to group thousands, the reverse of English.

Choose a house style first

Fixing numbers without a target style just moves the inconsistency around. Write down a short set of rules once and apply them to every transcript. Many English style guides spell out one to nine and use digits from 10 upward, but use digits for money, measurements, times, dates and percentages regardless of size. Pick whatever suits your readers, as long as it is written down.

Small counts
Words for one to nine, digits from 10, unless a sentence mixes both
Money
Currency symbol and digits, with words for millions and billions: $5 million
Dates
One order for the whole document, with the month written out to avoid 03/04 confusion
Times
Either 3 pm or 15:00, never both in one transcript
Units
Digits with a standard abbreviation: 10 km, 2 kg
Phone numbers and codes
Digits only, grouped the way your audience writes them
Verbatim exceptions
Keep words where the wording itself matters, as in a quoted statement or a legal record

If you translate or subtitle the material, add the number rules to a wider translation style guide so every language follows one policy.

Efficient fixes

  1. Search for every digit. Most editors with regular expressions accept [0-9] as a pattern; step through every match and correct format as you go.
  2. Search for number words. Run a regular expression or a series of searches for one, two, three and so on, plus hundred, thousand, million and billion, and convert the ones your style says should be digits.
  3. Search for symbols and abbreviations such as $, €, £, km and k, and confirm each one is consistent.
  4. Listen to every number that matters. Prices, dates, quantities, doses, scores and anything a decision rests on should be checked against the timestamped audio, not inferred from context.
  5. Rebuild phone numbers and reference codes in one piece after listening, rather than editing each fragment.
  6. Do a final read for numbers split across two lines or two subtitle cues.

The order matters: format passes are fast and safe, listening is slow and essential, so do the fast passes first and spend your listening time only on numbers that carry weight. A wider review routine is in proofreading an AI transcript.

A worked example: a quarterly meeting recording

Finance update, hypothetical

Suppose a team transcribes a 40-minute quarterly review. The transcript contains "revenue of 4.2 million", "four point two million" and "$4.2M" for the same figure, dates in three formats, and one line that reads "we expect fifteen new hires" where the speaker actually said fifty. The editor applies the house style with a digit search and a number-word search in about ten minutes, then listens to the eleven sentences that contain figures used in the minutes. That second pass catches the fifteen-or-fifty error, which no formatting rule would have found.

What numbers mean for subtitles

In subtitles, digits are usually better than words: "2,500" reads faster and takes less space than "two thousand five hundred", which matters when a single-line cue has limited room and a viewer has limited time. Times, scores and statistics are easier to take in as digits too. The main subtitle-specific checks are numbers split across two cues, which are hard to read, and units that change meaning when abbreviated. When subtitles will be published from a subtitle generator output, apply your number style to the SRT or VTT file directly, keeping cue numbers and timestamps untouched.

What numbers mean for translation

Translation is where number mistakes become expensive. Each language has its own conventions for decimals, thousands separators, dates and currency position, so a correct translation may legitimately rewrite "1,500.50" as "1.500,50". What should never change is the value. Check every figure in a translated transcript against the source, especially dates, which can silently swap day and month. The overview of common machine translation errors includes numbers and units alongside other categories worth checking.

Decide in advance whether units and currencies should be converted for the audience or left as spoken. Converting miles to kilometres in a voice track changes what the speaker said; leaving them may confuse viewers. There is no universally right answer, but there should be a consistent one.

A synthetic voice also has to read numbers aloud, and ambiguous digits such as "1/2", "2025" or "10k" can be read in an unexpected way, as covered in multilingual text-to-speech.

Pitfalls and limits of number cleanup

  • Global replacements without review. Replacing every "one" with "1" wrecks phrases such as "no one" and "one of them".
  • Fixing format but not value. A neatly formatted wrong figure is more dangerous than a messy right one.
  • Converting quotes. In verbatim records and published quotations, keep the speaker's words when the wording matters.
  • Trusting context guesses. If a number is unclear in the audio, mark it as unclear rather than picking the plausible value.
  • Applying English conventions to other languages. A comma decimal in a German transcript is correct, not an error.

How mydubly handles numbers

mydubly transcribes with Whisper large-v3-turbo in its default deployment, so numbers appear in whatever form the model predicts, and there is no setting to choose a number style. You get the transcript as plain text, as a timestamped transcript with [m:ss] labels for finding each figure in the audio, and as SRT and VTT subtitle files. Translated transcript jobs give the transcript in both languages, which makes it practical to compare every figure side by side.

mydubly cannot import a corrected transcript or subtitle file, so number fixes are made in the downloaded files. For a dubbed video, the voice reads the translated lines as generated; if a figure is read wrongly, the usual remedy is to replace that line in an editor. A transcript costs 1 credit per minute with a 5-credit minimum, so a 40-minute meeting costs 40 credits (4¢).

Next step: write your number rules and test them

Transcribe a short recording that is dense with figures using audio to text, then apply a draft house style with the search passes above and note which rules you hesitated over. Settle those rules in writing before the next file, and the cleanup becomes a routine ten-minute pass followed by targeted listening. If the files are recurring team calls, the page on meeting transcription covers the rest of that workflow.

Frequently asked questions

Why does my transcript write some numbers as words and others as digits?

Speech recognition models learned from transcripts written in many different styles, and they choose a form based on the surrounding words. Money and measurements tend to come out as digits, casual counts as words, and short or context-poor passages can go either way. A written house style and a quick search pass make the result consistent.

How do I find every number in a long transcript?

Search for digits with a regular expression such as [0-9], then search separately for number words like hundred, thousand and million, plus the words one to twelve. Step through each match rather than replacing all at once, because words such as 'one' appear in many non-numeric phrases.

Can formatting rules catch numbers that were misheard?

No. Formatting rules make numbers consistent, but a misheard figure such as fifteen for fifty looks perfectly normal in text. Listen to every number that matters, using timestamps to jump to the right place in the audio.

Should subtitles use digits or words for numbers?

Digits are usually easier and quicker to read in subtitles and take less space, especially for times, prices and statistics. Some subtitle style guides still spell out very small numbers or numbers at the start of a sentence. Whatever you choose, apply it consistently across the file.

Will a translation convert currencies and units?

It should keep the values, but separators, date order and currency position may change to suit the target language. Whether miles should become kilometres or prices should be converted is an editorial decision, not something to leave to chance. Compare every figure in the translated transcript with the source.

Why are phone numbers broken into pieces?

People say phone numbers in groups with pauses between them, and models may treat each group as a separate phrase, sometimes writing parts as words or adding punctuation between groups. Listen once to the full number and retype it in your chosen format rather than editing each fragment.