Workflows

Translating an Audiobook or Long Narration, Chapter by Chapter

To translate an audiobook with AI, first confirm you hold the rights to translate both the text and the recording, then process the book as chapter files under two hours each, choose one narrator-style voice for every chapter, and review the translated transcripts before generating or publishing the voice. The result can work well as a draft, an accessible companion or a market test; for a commercial release, literary translation and human narration may still be needed.

8 min read · Updated

Rights come before files

An audiobook carries at least two layers of rights, and translating it touches both. This is general information, not legal advice.

The text
The right to translate a book is normally an exclusive right of the copyright holder, usually the author or publisher. Owning a copy, even a purchased audiobook, does not include it. Self-published authors often hold it; traditionally published authors should check their contract, which may have assigned translation or audio rights to the publisher.
The recording
The narrator's performance and the producer's recording can carry their own rights, governed by the narration contract. Re-voicing a recording you commissioned may need the narrator's agreement depending on that contract.
Public domain works
A text whose copyright has expired can be translated freely, but a specific recording of it, and any modern translation used as a source, may still be protected.

Distribution adds a third check. Some audiobook retailers and distributors restrict, label or separately review AI-narrated titles, and these policies have been changing. Check the current terms of every platform you plan to use before investing in a full book. If any of this is unclear for your situation, ask a publishing lawyer or your literary agent.

Audio-first or text-first translation

There are two ways to make a translated audiobook, and the choice affects quality more than any tool setting.

  • Text-first. The manuscript is translated into the target language, ideally by a literary translator, and the translation is then narrated by a person or synthesized from text. The source is the clean, edited text the author approved.
  • Audio-first. The existing recording is transcribed, the transcript is translated, and a synthetic voice reads the translation, timed to the original narration. This is what speech translation tools such as mydubly do.

Audio-first is faster when the recording is all you have, or when you want a quick draft in several languages. Its weakness is that recognition is an extra step where errors can enter, even though the book's text already exists. A practical compromise for authors who have the manuscript: use the audio-first route for speed, then check the translated transcript against a translation of the manuscript, or have a reviewer compare it with the original text.

Preparing chapter files

Long-form audio has to be organized before it is translated:

  1. Work from the highest-quality master you have, not a compressed retail copy. Retail downloads may also be protected by DRM, which tools cannot and should not process; use your own production files.
  2. Keep one file per chapter. Most audiobook chapters are well under two hours. Split any longer section at a scene or section break, never mid-sentence.
  3. Convert formats if needed. M4B, a common audiobook container, is not on mydubly's accepted list, so export chapters as M4A, MP3, WAV or FLAC.
  4. Remove or set aside non-narration audio: opening and closing credits, music stings and retailer-specific intros. They either need their own translation or should be re-recorded.
  5. Name files so they sort: 01_opening-credits, 02_chapter-01, and so on, with the target language added to every output file.

How long recordings are split and processed behind the scenes is covered in transcribing long audio files; the practical point for audiobooks is that clean chapter files make every later step easier to review and repeat.

Choosing a narrator voice

Listeners spend hours with one voice, so the choice matters more than for a short video. With mydubly's eight stock voices, the narrator styles (Female narrator and Male narrator) are designed for long, steady reading, and the calm voices suit reflective nonfiction. Test two or three on the same chapter excerpt and listen for at least five minutes each, because a voice that sounds pleasant for thirty seconds can become tiring over a chapter. The general approach is covered in choosing an AI voice.

Then keep that voice for every chapter of the book. One chosen voice reads each file, so dialogue between characters is read in the same voice as narration. Many single-narrator audiobooks work this way, but a human narrator shifts tone and pace for each character; a stock voice will not. Books driven by dialogue, or full-cast productions, lose the most.

Reviewing the translated transcript

The translated transcript is where to catch problems before they become hours of audio. Review it chapter by chapter:

  • Character and place names must be identical in every chapter. Each chapter is a separate job, and long recordings are processed in segments, so a name can be rendered differently in different places. Keep a name list and check each chapter against it; translating names and technical terms explains the choices.
  • Invented words, in fantasy and science fiction especially, are often treated as misheard real words.
  • Wordplay, rhyme, dialect and deliberately unusual prose are where machine translation flattens an author's voice the most.
  • Numbers, dates and units in nonfiction should be checked against the manuscript.
  • Dialogue attribution: without speaker labels, make sure the translation keeps clear who is speaking.

For a commercial release, this review is a job for a fluent reader of the target language, ideally one with literary experience. For an internal or personal edition, the author's own spot checks against the manuscript may be enough.

Example: a nonfiction author testing a Spanish edition

Suppose a self-published author has a 6-hour-20-minute English audiobook across 18 chapter files and wants to test interest in Latin American Spanish. She first runs transcript jobs with a Spanish translation, which cost 380 credits (38¢) for the whole book, and asks a bilingual friend to review two chapters. The friend finds the author's coined term for her method translated three different ways. After agreeing on one rendering, she dubs three sample chapters, about 70 minutes, for 3,500 credits ($3.50) using the Male narrator voice and the Spanish (Latin America) option, and offers them free to her Spanish-speaking newsletter readers. A full dub of the book would cost 19,000 credits ($19.00); she decides based on the response, and on whether a professional translator should handle the full text.

Steps for a chapter-by-chapter translation

  1. Confirm text, recording and distribution rights for the target language.
  2. Export clean chapter files from your master, under two hours each, in an accepted format.
  3. Run transcript jobs with the target language to get translated transcripts for review.
  4. Build a name and term list, and review the translated transcripts against it and the manuscript.
  5. Choose and test a narrator voice on one chapter.
  6. Dub each chapter with the same voice and language variant, naming outputs consistently.
  7. Listen to every chapter in full, then master the files to your distributor's technical requirements.

Limits of AI narration, and when a professional is still needed

AI narration in this workflow is a translated reading timed to the original. That is useful, but it is not a performance:

  • Literary translation is a craft. Voice, rhythm, humor and subtext are what readers buy fiction for, and they need a translator, not just a translation.
  • A human narrator interprets: characters, emotion, comic timing and emphasis. A stock voice reads every line in the same style.
  • Hours of listening expose synthetic monotony and occasional odd stresses that pass unnoticed in a five-minute video.
  • Retailers and listeners may treat AI narration differently, and some markets expect human narrators for commercial titles.
  • Pronunciation of names and invented words cannot be corrected line by line with a stock voice.

AI translation fits drafts, market tests, educational and organizational narration, accessibility editions and personal or internal use where you hold the rights. The broader comparison is in AI dubbing vs human dubbing.

What mydubly does with audiobook files

mydubly accepts audio files (MP3, WAV, M4A, AAC, OGG and FLAC) up to 2 hours each and detects the spoken language automatically. A transcript job gives a timestamped transcript and, if you pick a target language, a translated transcript, for 1 credit per minute with a 5-credit minimum per file. A dubbing job adds the translated voice as an M4A track, at 50 credits per minute with a 2-minute minimum. Translated lines are placed near where the original sentences were read, with gentle speed changes where the translation runs long, so chapter pacing roughly follows the original narration.

It does not accept text manuscripts as input, clone the author's or narrator's voice, give different characters different voices, accept pronunciation dictionaries, or batch a whole book into one job. Files leave your device only as compressed audio, and results are deleted within 30 minutes of completion, so download each chapter as it finishes. Start with the audio translator; for podcasts and serialized audio, see podcast translation.

Next step

Before processing anything, write down who holds the translation rights for the text and the recording, and check the AI narration terms of the platform you would publish on. Then translate one chapter, have a fluent reader review it, and decide from that whether the book needs an AI edition, a professional one, or both in sequence.

Frequently asked questions

Can I translate an audiobook I bought for my own listening?

Buying an audiobook gives you a license to listen, not the right to translate it, and retail files are often DRM-protected, which you should not try to circumvent. For personal understanding, an ebook or print edition in your language, if one exists, is the straightforward route. Rules differ between countries, so this is not legal advice.

Can the translated version keep my own voice?

Not with mydubly. It offers eight stock voices and no voice cloning, so the translated edition will sound like a different narrator. If keeping the author's voice matters, the options are recording the translation yourself, if you speak the language, or hiring a narrator with a similar style. Consent and disclosure issues around cloning are covered in voice cloning vs stock voices.

How should I handle the opening and closing credits?

Process them separately or re-record them. Credits contain names, titles and legal lines that should be translated deliberately, and retailers often have specific requirements for what they say. Leaving them inside chapter one means a machine translation of your copyright notice, which is rarely what you want.

Will the translated chapters be the same length as the originals?

Roughly. Each translated line is placed near where the original sentence was read, and lines that run long are sped up gently, so a translated chapter generally follows the original's pacing and length. Pauses between sentences may be shorter where the target language needs more words.

What file format should I deliver to an audiobook distributor?

Distributors publish their own technical specifications covering file format, bitrate, sample rate, loudness and silence at the start and end of each file. mydubly's voice output is an M4A file, so plan to master and convert each chapter in an audio editor to meet the specification of the platform you are using.