What Chatterbox is and where it comes from
Resemble AI, a company known for commercial voice generation and voice cloning, published Chatterbox as an open-source model with its code and weights available to download. That makes it different from voices you can only reach through a paid API: developers can run it on their own hardware, inspect how it behaves and build it into products, subject to the license terms in the repository.
The official home is the resemble-ai/chatterbox repository on GitHub, with model weights published on Hugging Face under Resemble AI's organization. Because open models are updated frequently, the README there is the authoritative source for supported languages, installation steps, hardware requirements and any newer variants. Treat the descriptions below as an orientation rather than a specification.
Chatterbox Multilingual: one model, many languages
The original Chatterbox release focused on English. Chatterbox Multilingual extends the same approach so a single model can generate speech in a range of languages, selected with a language setting at generation time. For translation workflows this is the key property: one voice identity can be carried into whichever target language you need, rather than switching to a different model and a different-sounding speaker per language.
Multilingual TTS has its own wrinkles. A voice conditioned on a reference recorded in one language can carry a trace of that language's accent into another, and quality is rarely identical across all languages because training data is uneven. Listening tests in the specific language you care about matter more than a general impression of the model. Our article on multilingual text to speech goes into these effects.
Emotion exaggeration: a dial for intensity
Most TTS models give you a voice and little say over how animated it sounds. Chatterbox exposes an exaggeration control that scales the intensity of delivery. Low settings give a steady, restrained reading suited to documentation or calm narration; higher settings push toward more dramatic, emphatic speech. The README also describes a second guidance setting that interacts with pace, and suggests combinations for more expressive output.
It helps to be clear about what this dial is. It adjusts overall energy and emphasis; it does not let you mark up "say this word sarcastically" or direct an individual line the way a voice director would. Turned up too far, delivery can become theatrical or rushed. In practice teams pick a setting per voice or per content type and leave it.
Voice conditioning from a reference clip
Like other recent models, Chatterbox can imitate a voice from a short audio prompt, a technique usually called zero-shot voice cloning. You supply a clip of the target speaker, and the model generates new text in a voice resembling it. The general mechanics are explained in how AI voice generation works.
The reference clip shapes the result heavily. A clean, single-speaker recording with steady delivery yields a clean, steady voice; a clip with room echo, music or background noise passes those traits along. The repository also describes a watermark embedded in generated audio to help identify synthetic speech; check the current README for how that works and how to detect it, since details may change between releases.
Suppose a transcript arrives as three short segments: "So first," then "we open the settings panel" then "and turn on backups." Synthesized separately, each piece tends to sound like its own clipped sentence with a falling ending. Merged into "So first, we open the settings panel and turn on backups." the model reads one continuous thought with natural intonation, which is why sentence assembly before synthesis makes a noticeable difference.
How Chatterbox is built, in broad strokes
Chatterbox belongs to the speech-token family of TTS models. In broad terms, a transformer language model predicts a sequence of discrete speech tokens from the input text and the voice prompt, and a separate decoder stage turns those tokens into a waveform. This general design is what makes reference-based voice imitation work well and what makes output vary slightly between runs. For exact model sizes, component names and training data descriptions, rely on the official repository and any accompanying technical documentation rather than secondhand summaries, including this one.
Running Chatterbox yourself
If you want to experiment directly, the outline below matches how most open TTS projects are used. Follow the README for exact commands, since packages and versions change.
- Set up a recent Python environment, ideally on a machine with a GPU; CPU generation works for many models but is much slower.
- Install the package as the repository describes and let it download the model weights on first use.
- Generate a test sentence with the default voice to confirm the setup.
- Add a reference clip of a few seconds of clean speech to condition the voice, and compare with the default.
- For the multilingual model, set the language explicitly for each request rather than relying on the text alone.
- Adjust exaggeration and guidance in small steps, listening each time, and note the settings you settle on.
- Feed it whole sentences, not fragments, and normalize numbers and abbreviations in your text first.
Strengths and limitations of Chatterbox
On the plus side, Chatterbox combines open weights, expressive control, multilingual coverage and reference-based voices in one package, which is unusual among freely available models. Being open, it can be self-hosted, audited and integrated without a per-character API bill.
The limitations are the ones shared by its model family, plus the practical cost of running it yourself.
- Output is sampled, so occasional mispronunciations, odd pauses or clipped endings appear and may need a regenerate.
- Quality varies by language and by voice; test in each target language rather than assuming parity.
- Self-hosting means managing GPUs, latency and updates, which is real engineering work.
- Voice imitation from a sample creates consent and misuse responsibilities; the license and your own policies apply.
- Expressiveness is global, not line-by-line direction, so it will not replace an actor for dramatic material.
How mydubly uses Chatterbox Multilingual
mydubly's default voice engine is Chatterbox Multilingual. When you translate a video or audio file with full output on, your speech is transcribed with Whisper, translated by a neural machine translation engine, and then each translated sentence is voiced by Chatterbox in the target language. Some of the 8 preset voices are produced by conditioning the model on a reference recording; you choose one voice for the whole video, and there is no option to upload your own voice or clone anyone.
Before synthesis, mydubly merges short fragments into whole sentences of up to about 220 characters, as in the example above. Afterwards each clip is normalized, fitted to the original timing with at most a modest speed change, and joined into one continuous AAC track that is muxed with your video in the browser. You do not need to install anything or touch model settings. The languages offered are the 21 that both the speech recognition and voice models support; the AI dubbing page lists them along with the voices, and the audio translator applies the same voice step to podcasts and recordings.
Where to go next
If you are a developer, the official repository is the place to start experimenting. If you simply want to hear Chatterbox Multilingual voicing your own content in another language, try a short clip on AI dubbing: a 4-minute clip costs 200 credits ($0.20) and returns the dubbed MP4, the translated audio file, subtitles and both transcripts. Our guide to choosing an AI voice helps you pick among the presets.
Frequently asked questions
Is Chatterbox TTS free to use?
The code and model weights are published openly, so you can download and run them without paying Resemble AI per use. You still pay for the hardware or cloud compute you run it on. Commercial use is governed by the license in the official repository, so read it before building a product.
Does Chatterbox support languages other than English?
Yes. The multilingual version generates speech in a range of languages from a single model, chosen with a language setting. The official repository lists the current set. mydubly offers the 21 languages that both its speech recognition and voice models cover.
What does the exaggeration setting in Chatterbox change?
It scales how intense and emphatic the delivery sounds, from restrained to dramatic. It is a global dial for a generation rather than a way to direct individual words. Very high values can make speech sound theatrical or hurried, so most users settle on a moderate value per voice.
Can I choose Chatterbox settings when dubbing in mydubly?
No. mydubly exposes 8 preset voices rather than raw model parameters, and you pick one voice for the whole video. The engine settings behind each preset are fixed so that results stay consistent across a series.
Can Chatterbox copy a specific person's voice?
The model can imitate a voice from a short reference clip, which is why consent matters when using it directly. mydubly does not offer this to users: you cannot upload a voice sample, and only the stock voices are available. Our voice cloning vs stock voices article covers the trade-offs.