URL encoding: spaces, plus signs and double encoding

A plus is a space in the query string and a literal plus in the path; the two positions encode the same character differently; encoding twice yields a plausible value that never matches.

URL encoding has one rule (a percent sign and two hex digits), but three places follow different conventions. Mixing them is where bugs live.

Trap one: plus means space only in the query

?a=b+c       -> value is "b c"      (form encoding; + is a space)
/a+b         -> path is "/a+b"      (in a path, + is a plus)

The asymmetry is historical: query strings inherited application/x-www-form-urlencoded, paths did not. So:

Position Space Plus sign
Query string + or %20 must be %2B
Path must be %20 literal + is fine

A real plus in a query string must be encoded as %2B, or the other end decodes it as a space. This is the classic encoding bug.

Trap two: path and query have different reserved characters

encodeURIComponent encodes / ? & = #; encodeURI does not. Using the wrong one:

// encoding a single path segment
encodeURIComponent('a/b');   // 'a%2Fb'  <- one segment
encodeURI('a/b');            // 'a/b'    <- now two segments

// encoding a query value
encodeURIComponent('a&b');   // 'a%26b'  <- correct
encodeURI('a&b');            // 'a&b'    <- read as two parameters

The rule is simple: encodeURIComponent for a component, encodeURI for a whole URL. When building a query string, always the former.

The standard also requires encoding ! ' ( ) *, which encodeURIComponent leaves alone:

function strictEncode(text) {
  return encodeURIComponent(text)
    .replace(/[!'()*]/g, (ch) => '%' + ch.charCodeAt(0).toString(16).toUpperCase());
}

Most cases work without it, but signature schemes such as OAuth disagree over exactly these characters.

Trap three: double encoding

Encode an already-encoded string and % becomes %25:

original     a b
once         a%20b      <- correct
twice        a%2520b    <- one decode gives "a%20b", not "a b"

Double encoding never errors; it just yields the wrong value. Usual cause: a framework encoded once and the code encoded again.

Look at the density of % in the request: %25 almost always means double encoding. On the decode side, do not “decode until it stops changing” — if the value legitimately contains %, one decode too many destroys it.

Two practical rules

One: encode at the very last step of assembly.

const q = params.map(([k, v]) =>
  strictEncode(k) + '=' + strictEncode(v)
).join('&');

Assemble first and encode after, and you encode the separators too. URLSearchParams already handles this:

new URLSearchParams([['a', 'b c'], ['d', '1+2']]).toString();
// 'a=b+c&d=1%2B2'

Two: never trust a decoded value for path building.

%2E%2E%2F decodes to ../. A server that decodes and then joins paths has handed out directory traversal. The correct order is split the path into segments, validate each, then decode, or reject any decoded result containing / or ...

URL encoding is not merely swapping special characters for percent signs. Position sets the rules, path and query differ, and timing decides whether it works.

← Back to all posts

Comments

…