Base64: it is neither encryption nor compression
Base64 turns 3 bytes into 4 printable characters, growing the payload by a third. It solves carrying binary through a text channel, not keeping anything secret.
Base64 does exactly one job: map arbitrary bytes onto 64 printable characters. It hides nothing, shrinks nothing, and verifies nothing.
Why it exists
Some channels accept text only: mail bodies, JSON fields, URLs, HTTP headers. Raw binary gets truncated or mangled. Base64 splits 8-bit bytes into 6-bit groups and maps each to a character.
input "Man" → 4D 61 6E → 010011 010110 000101 101110 → "TWFu"
3 bytes in, 4 characters out. About 33 percent larger.
Common misconceptions
| Misconception | Reality |
|---|---|
| It is encryption | anyone can decode it and read the plaintext |
| It compresses | it always grows the data by a third |
| It is a secure transport | it fixes the alphabet, not secrecy or tampering |
| It can replace a hash | it is fully reversible, useless as a fingerprint |
Base64-encoding sensitive data into a URL parameter or cookie is plaintext. A great many token leaks start exactly there.
Variants and traps
There are several incompatible variants:
| Variant | Difference | Used in |
|---|---|---|
| Standard | + /, = padding |
email, JSON |
| URL-safe | - _, usually unpadded |
URLs, filenames, JWT |
| Unpadded | = dropped |
length-sensitive places |
Putting a standard-variant string in a URL breaks it: + decodes to a space in a query string and / can be swallowed as a path separator. JWT uses the URL-safe variant without padding for exactly this reason.
The padding rule
When the input length is not a multiple of 3, = is appended:
1 byte → 2 chars + "=="
2 bytes → 3 chars + "="
3 bytes → 4 chars
A decoder must accept both padded and unpadded forms, or integrations fail for no visible reason.
When not to use it
- You want secrecy → encrypt (AES-GCM and friends), not base64
- You want to save space → compress, or use a binary field
- You want verification → a hash or a checksummed encoding
The only good fit is data that must cross a text-only protocol and will be reliably decoded at the other end.
When you see a long run of
A-Za-z0-9+/ending in=, assume it is plaintext.

Comments
…