Transcription · Japanese

Japanese Transcription: Video & Audio to Text

Get a Japanese transcript of any video or recording spoken in Japanese — plain text, timestamped, or as subtitles. Select Japanese as the language to keep it untranslated, or another language to have it translated.

01 — Upload

Your video

Full translation outputTimestamped transcript only — no video or audio in results.

The video file stays in this browser. Language is required: pick the spoken language for a plain transcript, or another language to translate it.

02 — Result

Your transcript

Your timestamped transcript will appear here.

Waiting for upload

25 languages available

How Japanese speech transcribes

Japanese speech transcribes well into standard written Japanese. Because Japanese has no spaces and many homophones, names and specialist terms can come out with the wrong kanji.

What a Japanese transcript looks like

Japanese transcripts mix kanji, hiragana and katakana as in normal writing, with Japanese punctuation (、 and 。) and no spaces between words. Homophones are common in Japanese, so check names and technical terms where the wrong kanji could be chosen.

Examples

52-minute video

A Tokyo product team transcribes a 52-minute recorded Japanese user interview to tag feature requests.

9-minute recording

An engineer converts a 9-minute Japanese voice memo into text for a bug report.

Costs: 52 credits (5.2¢) and 9 credits (0.9¢), including a translated transcript if you select another language. Audio files go in the audio to text workspace.

Counting errors by character, not by word

Written Japanese has no spaces, so word error rate doesn't apply neatly and researchers measure Japanese transcription by character error rate instead. For your own check, correct a two-minute passage carefully and compare it with the original output. Are the errors spread thin, a wrong particle here and a wrong kanji there, or clustered around names and jargon? Clustered errors mean a terms list plus find-and-replace will fix most of the file; scattered errors usually point to the audio, such as echo or people talking over each other. Word error rate explained shows how both measures are calculated.

Katakana variants and full-width characters

Many loanwords have more than one accepted spelling: コンピューター and コンピュータ, サーバー and サーバ, ウイルス and ウィルス. Some technical style guides drop the final long-vowel mark while others keep it, so check which form the transcript uses and normalise to your own style. Alphanumerics need the same attention: product codes, years and English acronyms can appear in full-width (AI, 2024) or half-width (AI, 2024) characters. Mixed widths look untidy in subtitles and break searches, so pick one and convert the rest.

Kanji numerals, digits and era years

Spoken numbers can be written as 三百 or 300, and dates as 令和六年 or 2024年. Era names deserve special care: a speaker who says Reiwa 6 means 2024, and an international reader or a translated transcript needs the Western year. Counters (三人, 五本, 二枚) change with the thing being counted, so a counter that doesn't fit the noun is a sign that one of the two was misheard. The article on numbers in transcripts suggests a house style for figures, dates and units.

What happens to dropped subjects in translation

Japanese routinely omits subjects and objects when context makes them clear: 行きました can mean I, we, he or they went. The transcript reproduces this faithfully, but if you pick a target language, machine translation has to supply a pronoun and may guess wrong, especially in interviews where a person introduced minutes earlier is never named again. Keigo raises a related issue: the difference between いらっしゃる and 来る carries social information that an English translation flattens. When a quote matters, read the Japanese line next to the translation on the result screen before using it.

Fillers, aizuchi and sentence-final particles

Japanese conversation is full of hesitation sounds (えー, あの, えっと) and listener responses called aizuchi (はい, うん, そうですね, なるほど) that overlap the main speaker. In an interview, a transcript may drop many of these, keep some, or attach a listener's はい to the speaker's sentence. For subtitles that is usually fine. For conversation analysis or a user-research study where hesitation is data, listen through and restore them. Sentence-final particles such as ね, よ and よね are different: they change nuance, so check they are present in any line you quote.

Breaking Japanese subtitle lines

When editing the SRT, break lines after a particle (は, が, を, に) or punctuation rather than inside a word, and never split a kanji compound across lines. Avoid starting a line with small kana such as っ or ゃ, or with closing punctuation. Subtitle segmentation compares line-breaking conventions across scripts.

A 25-minute cooking tutorial from Osaka

Phone video, Kansai host

Recorded on an iPhone as MOV, the tutorial costs 25 credits (2.5¢) to transcribe, and the creator downloads the Japanese SRT for YouTube captions. A second run with English as the target costs the same again. The host speaks Kansai dialect, so the editor checks whether forms such as あかん and 〜へん were kept or turned into standard Japanese, and confirms that 大さじ and 小さじ amounts and gram figures use half-width digits.

The MOV to text page covers the phone-video side of this workflow.

Using the tool on this page

  • The spoken language is detected automatically; you select the output language — the spoken language itself for an untranslated transcript, or another language for a translation. Files can be up to 2 hours long.
  • Only the audio is sent for processing — the video stays on your device — and audio and text are deleted within 30 minutes of delivery (how files are protected).
  • You can download a timestamped transcript, subtitles (SRT and VTT) and plain text. Every output, step and limit is explained on the video to text page.

Other languages

Frequently asked questions

Will Japanese names be written correctly?

Common words, yes. Personal and company names can be written with many different kanji, so the model has to guess — check names in the transcript.

Will the transcript include furigana for difficult kanji?

No. The downloads are plain text, SRT and VTT, and those formats carry no ruby annotations in everyday use, so readings aren't included. Add furigana yourself in a document or design tool if learners need them.

Is Kansai or another regional dialect kept as spoken?

Check a sample. Dialect endings and vocabulary may be kept or nudged toward standard Japanese, and the choice can vary within one file. If you need a faithful record of dialect, listen through the passages where it is strongest.

How do I count words in a Japanese transcript?

Because there are no spaces, run the text through a morphological analyser (MeCab and similar tools split Japanese into words) before counting or building a vocabulary list. Character counts are the simpler measure for subtitle length.

Why can't I find a word I know was said?

It may be written differently from what you typed: in kanji instead of hiragana (分かる or わかる), with different okurigana, or with a homophone kanji. Search for the reading in kana and for a distinctive single kanji, then play the timestamp to confirm.