Unicode normalization: why two identical-looking strings differ
An e-acute can be one code point or a base letter plus a combining mark. Normalize to NFC before comparing, storing or indexing, or duplicates quietly slip through.
Two strings that look identical can compare unequal. It is not a browser bug: one character has several legal encodings in Unicode.
- U+00E9 (a single code point)
- U+0065 U+0301 (base letter plus a combining acute accent)
They render the same, but the bytes differ and so does length (1 versus 2).
The four normalization forms
| Form | What it does | Typical use |
|---|---|---|
| NFC | compose into the fewest code points | default for storage and comparison |
| NFD | decompose into base plus marks | text processing, stripping accents |
| NFKC | compatibility decompose, then compose | identifiers, search |
| NFKD | compatibility decompose | same, kept decomposed |
The rule is short: NFC for comparing and storing, NFD when you need to strip accents.
'é'.normalize('NFC').length // 1
'é'.normalize('NFD').length // 2
What goes wrong without it
| Scenario | Symptom |
|---|---|
| Username uniqueness | two “same” names both register |
| Login check | a password fails from another keyboard |
| Unique index | the database sees two distinct rows |
| Sorting | order does not match what the eye sees |
| Search | present data cannot be found |
Login is the nastiest one. A phone keyboard may emit NFC while a desktop IME emits NFD. Password hashing is byte-sensitive, so the user gets “wrong password”.
Where to normalize
Normalize at the boundary, trust it afterward. Do one normalize('NFC') where input enters, and everything written, cached or compared is canonical.
const normalizeKey = (s) => s.normalize('NFC').toLowerCase();
If existing data mixes both forms, do not just add a unique index. Normalize everything first, or the constraint fails the moment you create it.
Compatibility decomposition is different
NFKC is more aggressive: it turns ① into 1 and ㎏ into kg. Fine for identifiers and search, wrong for display because it changes what the user meant.
'①'.normalize('NFKC') // '1'
'①'.normalize('NFC') // '①'
Never apply NFKC to passwords. It merges visually distinct passwords into one and shrinks the key space.
Homoglyphs are another problem
Normalization solves different encodings of one character. It does not solve different characters that look alike: Cyrillic U+0430 and Latin U+0061 are indistinguishable in most fonts. That needs allowlists or confusable detection.
Byte-level equality always means equality after normalization. Do it once at the boundary and the rest is free.

Comments
…