Speech recognition & transcription

How to Run Whisper Locally on Your Own Computer

Whisper is OpenAI's open-source speech recognition model, released under the MIT license, and you can run it on your own computer for free. You need Python, ffmpeg and a few minutes in the terminal. Speed depends on your hardware: a modern GPU transcribes far faster than real time, while a laptop CPU can be slow on the larger models. Here's how to set it up and when a hosted service is the easier choice.

4 min read · Updated

What you need

  • Python 3 (a recent version) and pip.
  • ffmpeg, which Whisper uses to read audio and video files.
  • Disk space for the models, from tens of megabytes for the smallest to a few gigabytes for the largest.
  • Optionally, an NVIDIA GPU with CUDA for much faster transcription.

Install ffmpeg first. On a Mac with Homebrew: brew install ffmpeg. On Ubuntu or Debian: sudo apt install ffmpeg. On Windows, package managers like Chocolatey or Scoop can install it, or download a build and add it to your PATH.

Installing Whisper

The official package installs with pip:

  1. Open a terminal.
  2. Run pip install -U openai-whisper.
  3. Check it works with whisper --help.

Using a virtual environment keeps Whisper's dependencies separate from other Python projects. The first time you use a model, Whisper downloads it automatically.

Choosing a model

Whisper comes in several sizes. Smaller models are faster and need less memory; larger ones are more accurate, especially on hard audio and less common languages. The project's README lists approximate memory needs:

tiny and base
About 1 GB of video memory; fast, rougher results.
small
About 2 GB; a reasonable middle ground on modest hardware.
medium
About 5 GB; noticeably better on difficult audio.
large
About 10 GB; the most accurate of the original sizes.
turbo
About 6 GB; an optimised version of large-v3 that's much faster with a small accuracy trade-off.

For tiny through medium there are also English-only versions (with .en in the name) that do slightly better on English. Whisper large-v3-turbo explains the turbo model in detail.

Transcribing from the command line

The basic command is whisper followed by your file:

  • whisper interview.mp3 --model turbo
  • whisper lecture.mp4 --model medium --language Spanish
  • whisper call.m4a --model small --output_format srt

Whisper writes outputs next to the file: plain text, SRT and VTT subtitles, a TSV and JSON with timestamps. Setting --language skips language detection, which avoids mistakes on short or mixed clips. The --task translate option translates speech into English; Whisper doesn't translate into other languages.

GPU versus CPU

On an NVIDIA GPU, Whisper uses CUDA through PyTorch and runs many times faster than on a CPU. On a Mac or a laptop without a suitable GPU, the official package runs on the CPU, where the larger models can take longer than the recording itself.

That's where whisper.cpp helps: a C/C++ reimplementation that runs efficiently on CPUs and Apple Silicon. You build it from source or install a packaged version, download a converted model and run it from the terminal. Several desktop apps for Mac and Windows wrap whisper.cpp in a graphical interface, if you'd rather not use the command line.

Example: a journalist with sensitive interviews

A reporter has eight hours of interviews with confidential sources and doesn't want them on any server. On her laptop with an NVIDIA GPU, she installs Whisper, runs the turbo model with --language English, and gets SRT and text files for each interview overnight. For less sensitive material, like press conferences, she uses a hosted service from her phone.

Limits and common problems

  • "ffmpeg not found": install ffmpeg and make sure it's on your PATH.
  • Out-of-memory errors: choose a smaller model.
  • Repeated or invented text in silent stretches: a known Whisper behaviour; Whisper hallucinations explains causes and fixes, such as trimming long silences.
  • Wrong language detected: pass --language explicitly.

Processing many files

For a folder of recordings, a short shell loop runs Whisper on each file in turn, or you can pass several files to one whisper command. Keep the model loaded by processing in one session where possible, and write outputs to a separate folder with --output_dir so source files stay tidy. Check a few outputs before leaving a long batch running overnight, especially the language detection on the first files. Transcribing long audio files covers splitting very long recordings, which also helps recover from crashes partway through. The how to transcribe audio guide covers preparing files, whichever tool you use.

When mydubly is easier than local Whisper

Running Whisper yourself is free per file and keeps audio on your machine. A hosted service makes more sense when:

  • Your computer is slow or has no suitable GPU, and you have hours of audio.
  • You need a translation into a language other than English.
  • You want a dubbed version, not just text.
  • You'd rather not maintain a Python environment.

mydubly's transcription is built on Whisper, run on GPUs in the cloud. Upload to video to text: the video stays in your browser and only the audio is sent over HTTPS, deleted within 30 minutes of delivery. It costs 1 credit per minute, translates into any of 21 languages from the same run and produces the same SRT and VTT outputs. What Whisper is covers the model's background, and on-device speech AI covers the trade-offs of keeping everything local.

Frequently asked questions

Is Whisper free to run locally?

Yes. Whisper's code and models are open source under the MIT license, so running it on your own computer costs nothing per file.

How do I install Whisper?

Install ffmpeg, then run pip install -U openai-whisper in a terminal. Models download the first time you use them.

Which Whisper model should I use?

turbo is a good default on a capable GPU. On modest hardware, small or base are faster; medium and large are more accurate on hard audio.

Can Whisper run on a Mac?

Yes. The official package runs on the CPU; whisper.cpp runs efficiently on Apple Silicon.

Can Whisper translate into Spanish or Hindi?

No. Whisper's built-in translation goes into English only. Other languages need a separate translation step.

Does mydubly use Whisper?

Yes. mydubly's transcription is built on Whisper, run on cloud GPUs, with translation into 21 languages added.