Troubleshooting

Garbled Subtitles: Why Accents Turn Into Symbols and Scripts Into Boxes

Subtitles show weird characters when the file's text encoding does not match what the player expects, or when the player's font has no letters for that script. Accents that turn into pairs of symbols usually mean a UTF-8 file read as an older Western encoding; empty boxes usually mean a missing font. Open the file in a text editor that shows its encoding, re-save it as UTF-8 if needed, and choose a font that covers the language.

7 min read · Updated

Match the symptom to the cause

The exact shape of the garbling is a strong clue, so compare what you see with this table.

é appears as é
A UTF-8 file is being read as a Western legacy encoding such as Windows-1252
Black diamonds with question marks
A file in a legacy encoding is being read as UTF-8
Plain question marks replace letters
The file was saved in an encoding that could not hold those letters, and they were lost
Empty boxes or rectangles
The font has no glyphs for that script
Stray  before the first line, or the first cue missing
A byte order mark is being read as text
Hindi or Arabic letters look disconnected
The player lacks the text shaping those scripts need
Punctuation at the wrong end of a right-to-left line
A bidirectional display issue, covered in right-to-left video localization

The first two rows are reversible, because the bytes are intact and only being misread. The third is not: the original letters are no longer in the file.

How text encoding causes garbling

A text file stores bytes, and an encoding is the rulebook that maps bytes to characters. Plain English letters map the same way in almost every encoding, which is why English subtitles rarely break. Accented letters and other scripts do not. Before UTF-8 became common, each region used its own code page: Windows-1252 for Western European languages, Windows-1251 for Cyrillic, Windows-1256 for Arabic, Windows-1255 for Hebrew, GBK for Simplified Chinese, Shift JIS for Japanese, and others.

UTF-8 can represent every script in one encoding, which is why it is the safe default today. The catch is that an SRT file has no header declaring its encoding, so the player has to guess. When it guesses wrong, each UTF-8 accented letter, stored as two bytes, is shown as two unrelated characters, producing the familiar é pattern.

A byte order mark is an optional invisible marker at the start of a file that signals UTF-8. Some Windows software relies on it to detect UTF-8, while a few older or strict parsers treat it as text and stumble on the first cue.

Fonts and text shaping

Correct encoding is not enough if the font cannot draw the letters. A font without Devanagari, Arabic or Chinese glyphs shows empty boxes, sometimes nicknamed tofu, even though the file is perfect. Some scripts also need shaping: Arabic letters change form depending on their neighbors, and Devanagari combines consonants into conjunct forms. Players and devices without a shaping engine show the right letters in the wrong forms.

Smart TVs, older set-top boxes and some embedded players ship with limited fonts and may not offer a way to add more. Desktop players such as VLC let you choose the subtitle font in their preferences.

Diagnose in this order

  1. Open the subtitle file in a text editor that shows encoding, such as VS Code, whose status bar displays it, or Notepad++, which has an Encoding menu.
  2. If the text looks correct in the editor, the file is fine. The problem is the player's encoding guess or its font.
  3. If the text looks garbled in the editor too, use the editor's reopen-with-encoding option and try the likely legacy encoding for the language. When the letters snap into place, you have identified the file's real encoding.
  4. Search for plain question marks inside words. If letters were replaced by question marks, the damage is permanent and the file needs to be recreated.
  5. Play the file in a second player. If one shows it correctly and another does not, set the second player's subtitle encoding or font.
  6. If only the first cue fails, check whether a byte order mark is present and whether that player handles it.

Re-save a subtitle file as UTF-8

Once the editor shows the text correctly, save a UTF-8 copy:

  • VS Code: click the encoding in the status bar, choose Reopen with Encoding and pick the correct one, then click it again and choose Save with Encoding, UTF-8.
  • Notepad on Windows: Save As, and choose UTF-8 in the Encoding list.
  • Notepad++: use the Encoding menu to convert to UTF-8, then save.
  • Subtitle editors such as Subtitle Edit let you pick the encoding when saving.

From a terminal, iconv converts in one step. For a Western European file, then for an older Arabic one:

iconv -f WINDOWS-1252 -t UTF-8 old.srt > fixed.srt

iconv -f WINDOWS-1256 -t UTF-8 arabic-old.srt > arabic.srt

If a particular player still misreads the UTF-8 file, try saving it with a byte order mark, which VS Code and Notepad++ offer as UTF-8 with BOM. If a strict uploader rejects the first cue, try saving without one.

A worked example: Hindi subtitles on a smart TV

Boxes on the TV, fine on the laptop

Suppose a school plays a translated science video with Hindi subtitles from a USB stick on a classroom smart TV, and every subtitle appears as a row of boxes. On a laptop, the same SRT opens correctly in VS Code, which reports UTF-8, and plays correctly in VLC. The file is fine; the TV's font has no Devanagari glyphs. Changing the encoding would not help. The teacher plays the video from the laptop through the TV instead, and for future lessons keeps a version with the subtitles burned into the picture for that room.

The diagnosis took two minutes because the first check, opening the file in an editor, separated a font problem from an encoding problem.

Pitfalls and limits

  • Question-mark damage cannot be undone by re-saving. Recreate the file from a source that still has the letters.
  • Converting a file that was already UTF-8 as if it were legacy produces double garbling, such as é patterns. Always confirm the current encoding first.
  • Automatic encoding detection is a guess, and short files give it little to work with.
  • Some devices simply cannot render certain scripts; no file change will fix a missing font.
  • Burning subtitles into the picture solves display on any screen, but they can no longer be switched off or corrected without re-exporting.

What mydubly's subtitle files use

mydubly saves its SRT and VTT downloads as UTF-8 text, which can hold all 21 supported languages, including Arabic, Hebrew, Hindi, Russian, Japanese, Korean and Chinese. Chinese output is Simplified; if your audience reads Traditional characters, see translating video for regional audiences. Each job produces one subtitle track, in the spoken language or a chosen target language, and subtitles are delivered as files rather than burned into the picture.

So when mydubly subtitles show strange characters, look at the display chain first: the player's encoding setting, its font, or a later re-save in another program that switched the encoding. Editing subtitle files safely, without changing their encoding by accident, is covered in how to edit an SRT file. If a file will not load at all, the checklist in SRT file not working is the better starting point.

When another workflow fits better

  • For screens without the right fonts, burn subtitles into a copy of the video with a video editor.
  • For right-to-left layout and punctuation, follow the dedicated right-to-left guidance rather than encoding fixes.
  • For a platform that rejects a file it reads incorrectly, check its documentation for the encoding it expects.

Next step: generate a fresh UTF-8 file

If a subtitle file has lost letters to question marks, recreate it rather than repairing it: run the original video through the subtitle generator, or the video translator for another language, and attach the new file using the steps in adding subtitles to a video.

Frequently asked questions

Why do my subtitles show é instead of é?

The file is UTF-8, but the player is reading it as a Western legacy encoding such as Windows-1252, so each two-byte accented letter appears as two symbols. Set the player's subtitle encoding to UTF-8, or save the file as UTF-8 with a byte order mark so the player detects it.

Should an SRT file be saved with or without a BOM?

Either can be right. A byte order mark helps some Windows software recognize UTF-8, while a few strict parsers misread it as text in the first cue. Start without one if the destination documents UTF-8 support; add it if a specific player keeps misreading accented letters.

Why do Arabic subtitles show separate letters instead of joined words?

The letters are correct, but the player or device is not applying Arabic shaping, which joins letters and changes their forms by position. Try a different player or a font with full Arabic support. If the device cannot shape Arabic, burned-in subtitles are the reliable fallback.

Can I fix subtitles where letters became question marks?

Not from that file. When text is saved in an encoding that cannot represent a letter, the letter is replaced by a question mark and the original is gone. Go back to an earlier copy, the source file, or regenerate the subtitles from the video.

Do VTT files have the same encoding problems as SRT?

Less often. The WebVTT format specifies UTF-8, so browsers and web players expect it. Problems still appear when a VTT file is edited and re-saved in another encoding by a desktop program, or when a player's font lacks the script.