What this tool shows

Paste any string and it is split the way a reader sees it — into grapheme clusters — and then into the code points behind each one. Every code point is shown with its name, its Unicode block and general category, and the exact UTF-8 and UTF-16 bytes it encodes to.

The same string is re-encoded into legacy encodings alongside Unicode: Shift_JIS and CP932, EUC-JP, ISO-2022-JP, Big5, GBK, GB18030, EUC-KR and over twenty more. Characters that cannot survive the round trip are called out, which is usually where mojibake starts.

For CJK text it resolves IRG source references, shows which Ideographic Variation Sequence selects which shape, and reports whether the fonts on the page can actually draw it. The four normalization forms — NFC, NFD, NFKC and NFKD — sit side by side so you can see what each one changes.

Everything runs in your browser. The text you paste is never uploaded.

Guides

Longer explanations of what the tool is showing you. Read them in Japanese

All guides