Use case

Voice Memo to Text

Record your thoughts, then get them back as text. Drop in a Voice Memo, Android recording or MP3 and get an editable transcript — or a translation into another language.

01 — Upload

Your audio

Audio + text outputText only — timestamped transcript with no audio in results.

The audio file stays in this browser. Language is required: pick the spoken language for a plain transcript, or another language to translate it.

02 — Result

Your transcript

Your timestamped transcript will appear here.

Waiting for upload

25 languages available

Why it helps

Voice memos are the fastest way to capture ideas but the slowest to review. A transcript makes them skimmable and easy to turn into notes, drafts or tasks.

Step by step

  1. On iPhone, share the memo from Voice Memos to Files or AirDrop it to your Mac; on Android, share the recording file.
  2. Drop it into the audio workspace with text-only output.
  3. Copy the transcript into your notes app or document.

Tips

  • Several short memos combined into one file avoid paying the 5-credit minimum on each.
  • Speak your punctuation naturally — pauses become sentence breaks.

Example

7 minutes

A novelist turns a 7-minute walking memo into a scene outline for 7 credits (0.7¢).

Cost: 7 credits (0.7¢).

Limitations to know

  • Files under 5 minutes are billed at the 5-credit minimum (0.5¢).

Getting the memo file off your phone

mydubly works on a file stored on your device; there is no app to install and no link import. Where the file usually comes from:

iPhone Voice Memos
M4A. Share the memo, choose Save to Files, then pick it from Files in the browser upload
Android recorder apps
Commonly M4A, AAC or MP3. Some older recorders save AMR or 3GP, which aren't accepted, so change the format setting or convert first
Messaging-app voice notes
Often Opus audio saved as .ogg or .opus. OGG is accepted; convert an .opus file to M4A or MP3 if it is refused
Handheld voice recorders
Usually MP3 or WAV, copied over USB or a card reader

Talking so the transcript reads well

This transcribes a finished recording; it isn't dictation software, so spoken commands such as "comma" or "new paragraph" are likely to be written out as words. What helps instead:

  • Pause between thoughts. Pauses tend to become sentence breaks, and a clear gap separates one idea from the next.
  • Start each idea with a short label, such as "Newsletter idea" or "Call the landlord", so you can scan the transcript for them later.
  • Spell an unusual name once, letter by letter, the first time you say it.
  • Keep the phone a short, steady distance from your mouth, and out of the wind if you record while walking.
  • On site visits or inspections, say the location or photo number before each observation, so the transcript lines up with the pictures you took.

Recording audio on a phone covers placement and settings, and dictation vs transcription explains why the two work differently.

Many short memos or one long file

Transcripts are billed at 1 credit per minute with a 5-credit minimum per file, so a 40-second memo costs as much as a 5-minute one. If you collect a week of quick notes, join them into one file in an audio editor such as Audacity or GarageBand before uploading. Say the date or topic at the start of each memo, or write down where each one starts in the joined file; the timestamped transcript then shows where each note begins, and you can split the text afterwards. Leave a second or two of silence at each join rather than butting memos together, so the last word of one note doesn't run into the first word of the next.

Turning a rambling memo into usable text

Thinking out loud circles back on itself, so expect to edit. Read the timestamped transcript once, cut repetition, pull out the labels you spoke as headings, and move the result into your notes app or draft. mydubly does not write summaries or to-do lists from the memo; that step is yours, but the timestamps let you replay the exact moment a half-formed idea was said. If you keep notes in a linked system, transcripts in a personal knowledge base suggests how to file them. Long silences, music or café noise can occasionally produce stray text that you never said, so check any passage that doesn't sound like you; Whisper hallucinations explains why.

Memos in more than one language

The spoken language is detected automatically, so a memo in one language needs no source-language setting. Memos that switch languages mid-sentence can come out uneven; recording each memo in a single language avoids that. To pass a memo to someone who reads another language, pick a target language and download the transcript, which then contains the translated text. The result screen shows your original words next to the translation, so you can spot a misheard phrase before sending it on. If you also want your own words as a file, run the memo once more with the language you spoke selected, which returns your words untranslated; at 1 credit per minute that rarely costs more than the 5-credit minimum.

A week of voice notes

Ten memos between 1 and 4 minutes long add up to 26 minutes. Uploaded one by one, each is billed at the 5-credit minimum, 50 credits (5¢) in all. Joined into a single 26-minute file they cost 26 credits (2.6¢), and a separate 40-minute walking dictation adds 40 credits (4¢).

Using the tool on this page

  • The spoken language is detected automatically; you select the output language — the spoken language itself for an untranslated transcript, or another language for a translation. Files can be up to 2 hours long.
  • Audio is sent over HTTPS in small chunks and deleted with its transcripts within 30 minutes of delivery (how files are protected).
  • You can download a timestamped transcript, subtitles (SRT and VTT) and plain text. Every output, step and limit is explained on the audio to text page.

Frequently asked questions

Does it work on my phone?

Yes, in a mobile browser. Choose the file from your Files or Downloads.

Can I translate my memo?

Yes — pick a target language to get the transcript in another language too.

Will the transcript be punctuated?

Whisper-based recognition normally returns punctuated, capitalised text, and pauses usually become sentence breaks. Long unbroken stretches of speech can come out as run-on sentences; fixing transcript punctuation covers quick repairs.

My memo is really a recorded meeting or interview. Does that change anything?

The upload is the same, but several voices and no speaker names in the transcript call for a different review. The meeting transcription and interview transcription pages describe those workflows.

Is a memo recorded somewhere noisy worth uploading?

Usually, yes, and testing it costs 5 credits (0.5¢). If traffic, wind or a coffee machine covered parts of your voice, those stretches may come back garbled or missing; listen to them at their timestamps and re-record the key points somewhere quieter.