What this tool shows
Paste any string and it is split the way a reader sees it — into grapheme clusters — and then into the code points behind each one. Every code point is shown with its name, its Unicode block and general category, and the exact UTF-8 and UTF-16 bytes it encodes to.
The same string is re-encoded into legacy encodings alongside Unicode: Shift_JIS and CP932, EUC-JP, ISO-2022-JP, Big5, GBK, GB18030, EUC-KR and over twenty more. Characters that cannot survive the round trip are called out, which is usually where mojibake starts.
For CJK text it resolves IRG source references, shows which Ideographic Variation Sequence selects which shape, and reports whether the fonts on the page can actually draw it. The four normalization forms — NFC, NFD, NFKC and NFKD — sit side by side so you can see what each one changes.
Everything runs in your browser. The text you paste is never uploaded.
Guides
Longer explanations of what the tool is showing you. Read them in Japanese
- Characters Are a Lie: Understanding Grapheme ClustersWhy string.length gives wrong answers, what grapheme clusters really are, and how Intl.Segmenter fixes everything.
- UTF-8 Byte by Byte: How Characters Become BytesA visual, byte-level walkthrough of UTF-8 encoding showing exactly how code points map to 1-4 bytes.
- Unicode Normalization: NFC, NFD, NFKC, NFKD DemystifiedWhy the same-looking text can have different bytes, when each normalization form matters, and how to see the differences visually.
- Shift_JIS vs CP932: The Encoding Everyone ConfusesThe precise technical differences between Shift_JIS and CP932 (Windows-31J), with byte-level evidence.
- IVS: How Unicode Represents 47 Versions of the Same KanjiUnderstanding Ideographic Variation Sequences and Standardized Variation Sequences, with live font rendering of all registered variants.
- Unicode Homoglyph Attacks: When Characters Lie About Who They AreHow visually identical characters from different scripts enable phishing and spoofing, and how to detect them.