> text fix-characters
Fix garbled characters
“Müller” instead of “Müller”, “–” instead of “–”, question marks in diamonds: paste text, and the broken spots are detected and converted back. Also: clean up disruptive invisible characters and look up any character.
Paste from an email, a CSV or an export, or open a file. Up to 200,000 characters, the repair runs as you type, above that only when you press the button.
Before
After
The counters show how often each kind occurs, before you change anything. Turn off what should stay.
Opening or closing follows from the position in the sentence. An apostrophe in the middle of a word (it’s) stays an apostrophe.
Your text with tags for everything you cannot normally see
NBSP = non-breaking space, ZWSP = zero-width space, SHY = soft hyphen, BOM = byte order mark, TAB and CR = tab and carriage return. Red tags are control characters. From 100,000 characters, only the beginning is shown.
Type or paste a character, also one you copied from a text. It also works with the number: U+00E4, ä or ä
More characters from your input
More for text: Count and clean up · Compare texts
Your text stays on this device. Opened files are also only read here, nothing is saved or transmitted.
How it works
- 01Repair: Paste text, done. Every broken spot is highlighted before and after, and the list above says what was converted and how often.
- 02Clean up: turn each kind of disruptive character on or off, with a counter. “Show invisibles” makes visible what is in the text, without changing anything.
- 03Look up: type a character and find out what it is, which number and bytes it has and how to write it in HTML.
Where the mess comes from: A “ü” in UTF-8, today’s usual encoding, is the byte sequence C3 BC. If a program reads these bytes as Windows-1252 (the old default encoding of Windows), it shows “ü”. The tool takes every character back to its byte and puts the bytes together again as UTF-8. This also works if the text was read as Latin-1 or broke twice in a row (“Müller”).
What does not work: Texts read the other way around, where a “ü” became a “�” or a “?”, cannot be saved: The character is gone, and nothing records which one it was. The tool tells you how many such spots there are, but does not guess. The repair only applies where a valid byte sequence stands: “SÃO PAULO” is left alone because no matching character follows the Ã. For Greek, Cyrillic, Hebrew and Arabic, text is only repaired if it has several such spots, so that a single “×” before a non-breaking space does not suddenly become a Hebrew letter. Where two readings are possible, the page shows both.
Files: “Open file” reads a text file (TXT, CSV, …) right here in the browser. The encoding is guessed (marker, valid UTF-8, otherwise Windows-1252) and can be changed. “Save as UTF-8 file” writes the result anew; the option “with BOM” sets the marker that Excel on Windows needs to open a CSV as UTF-8. Up to 5 MB.