Unicode normalization: why two identical-looking strings differ

An e-acute can be one code point or a base letter plus a combining mark. Normalize to NFC before comparing, storing or indexing, or duplicates quietly slip through.

Two strings that look identical can compare unequal. It is not a browser bug: one character has several legal encodings in Unicode.

  • U+00E9 (a single code point)
  • U+0065 U+0301 (base letter plus a combining acute accent)

They render the same, but the bytes differ and so does length (1 versus 2).

The four normalization forms

Form What it does Typical use
NFC compose into the fewest code points default for storage and comparison
NFD decompose into base plus marks text processing, stripping accents
NFKC compatibility decompose, then compose identifiers, search
NFKD compatibility decompose same, kept decomposed

The rule is short: NFC for comparing and storing, NFD when you need to strip accents.

'é'.normalize('NFC').length   // 1
'é'.normalize('NFD').length   // 2

What goes wrong without it

Scenario Symptom
Username uniqueness two “same” names both register
Login check a password fails from another keyboard
Unique index the database sees two distinct rows
Sorting order does not match what the eye sees
Search present data cannot be found

Login is the nastiest one. A phone keyboard may emit NFC while a desktop IME emits NFD. Password hashing is byte-sensitive, so the user gets “wrong password”.

Where to normalize

Normalize at the boundary, trust it afterward. Do one normalize('NFC') where input enters, and everything written, cached or compared is canonical.

const normalizeKey = (s) => s.normalize('NFC').toLowerCase();

If existing data mixes both forms, do not just add a unique index. Normalize everything first, or the constraint fails the moment you create it.

Compatibility decomposition is different

NFKC is more aggressive: it turns ① into 1 and ㎏ into kg. Fine for identifiers and search, wrong for display because it changes what the user meant.

'①'.normalize('NFKC')   // '1'
'①'.normalize('NFC')    // '①'

Never apply NFKC to passwords. It merges visually distinct passwords into one and shrinks the key space.

Homoglyphs are another problem

Normalization solves different encodings of one character. It does not solve different characters that look alike: Cyrillic U+0430 and Latin U+0061 are indistinguishable in most fonts. That needs allowlists or confusable detection.

Byte-level equality always means equality after normalization. Do it once at the boundary and the rest is free.

← Back to all posts

Comments

…