home / blog / base64 encoding

What Is Base64 Encoding? A Plain-English Guide

It's not encryption, it's not compression, and it always makes files bigger. Here's what Base64 is actually for.

Base64 shows up everywhere in web development: inline images in CSS, API tokens, email attachments, data embedded in URLs. Yet its purpose gets misunderstood constantly. People mistake it for encryption ("I base64'd the password, so it's secure") or for compression ("I'll base64 this to make it smaller"). It's neither. Base64 does exactly one thing: it turns binary data into plain text, so that systems which only trust text can carry it without corrupting it.

That single job sounds trivial, but it's the reason Base64 has survived, mostly unchanged, since it was formalized for email in the early 1990s (RFC 1421, later RFC 4648). Nearly every language runtime, browser, and command-line toolchain ships a built-in encoder/decoder, because so much plumbing still depends on it: HTTP headers, JSON payloads, certificate files, JWTs, and more.

The actual problem Base64 solves

A huge number of systems (email over SMTP, URLs, JSON, XML, and older text-only protocols) were designed to carry text, not arbitrary binary data. Binary data can contain byte values that these systems interpret as control characters, line breaks, or invalid encoding, which corrupts the data in transit.

Concretely: a raw binary file can contain any of the 256 possible byte values (0–255), including things like null bytes (0x00), the byte for a carriage return, or byte sequences that aren't valid UTF-8 at all. Drop that into a JSON string, an XML document, or an SMTP message body, and you'll get a parse error, silent truncation, or mangled bytes. Those formats only guarantee safe passage for a much smaller set of printable characters.

Base64 solves this by mapping every 3 bytes of binary input to 4 printable ASCII characters, drawn from a fixed alphabet of 64 characters (A-Z, a-z, 0-9, +, /, with = used for padding). The result is guaranteed to be plain, printable text that any text-safe system can carry without corruption, at the cost of making the data about 33% larger.

How the encoding actually works

Base64 works in fixed groups of 3 input bytes (24 bits total). Those 24 bits get split into four 6-bit chunks, and each 6-bit chunk (a number from 0–63) is looked up in the 64-character alphabet to produce one output character. That's the whole algorithm: no compression, no substitution cipher, just a straightforward regrouping of bits.

A worked example: the text Man is the three bytes 77, 97, 110, or in binary, 01001101 01100001 01101110. Regroup those 24 bits into four groups of 6: 010011 010110 000101 101110. Converted to decimal that's 19, 22, 5, 46, and looking those up in the Base64 alphabet gives T, W, F, u. So Man encodes to TWFu. Every Base64 encoder, in every language, is doing exactly this, byte group by byte group.

This 3-bytes-in, 4-characters-out ratio is also why the ~33% size increase is not an approximation you can improve on. It's a fixed mathematical consequence of packing 8-bit bytes into 6-bit output characters. 3 bytes is 24 bits either way; you can't represent 24 bits in fewer than four 6-bit symbols. There's no smarter Base64 implementation that avoids this; the overhead is baked into the format itself.

Where you'll actually run into it

Email attachments. SMTP was originally 7-bit text only. Attaching a binary file (an image, a PDF) means encoding it as Base64 text inside the message body, which is why email attachments have always inflated file size somewhat.

Data URIs in CSS/HTML. background-image: url(data:image/png;base64,...) embeds an image directly inside a stylesheet or HTML file instead of a separate request. It's useful for small icons where saving an extra HTTP request outweighs the size penalty.

Basic Auth and API tokens. The HTTP Authorization: Basic header carries username:password Base64-encoded, not for security (it's trivially reversible), but because HTTP headers must be text, and this is a compact, standard way to pack two text fields into one text value.

JWTs (JSON Web Tokens). Each part of a JWT is a JSON object, Base64-URL-encoded (a URL-safe variant that swaps +// for -/_) so the token can be safely passed in a URL or header.

Certificates and keys. PEM-formatted TLS certificates and private keys (the files that start with -----BEGIN CERTIFICATE-----) are Base64-encoded DER binary data, wrapped at 64 characters per line and framed with those header/footer markers. It's a direct descendant of the same line-wrapping convention Base64 originally used for email.

Standard Base64 vs. URL-safe Base64

The standard alphabet's last two characters, + and /, both carry special meaning in URLs: + often means "space" in query strings, and / is a path separator. Put standard Base64 output straight into a URL and, depending on the system, it may get mangled or need percent-escaping, which defeats the point of using a compact text encoding in the first place.

URL-safe Base64 (defined in RFC 4648 §5) fixes this by substituting - for + and _ for /, two characters with no special meaning in URLs, filenames, or cookies. Everything else about the algorithm is identical; it's the same 3-bytes-to-4-characters mapping, just with two swapped symbols. This is the variant JWTs, most OAuth tokens, and many REST APIs use by default.

One quirk worth knowing: standard and URL-safe Base64 are not interchangeable without translation. Feed a URL-safe string into a decoder expecting the standard alphabet (or vice versa) and it will fail or silently produce garbage on any character that happens to be +, /, -, or _. If a token "mostly" decodes but corrupts a few characters, this mismatch is the first thing to check.

Padding: what the = signs are for, and when they disappear

Because Base64 processes input 3 bytes at a time, input lengths that aren't a clean multiple of 3 leave a partial group at the end. Padding fills that gap so the output is always a multiple of 4 characters: one trailing = if the input had 2 leftover bytes, two trailing = if it had 1 leftover byte, and no padding at all if the input divided evenly by 3.

Padding matters for concatenation and fixed-length parsing. Some decoders use it to detect where a Base64 block ends. But plenty of real-world implementations drop it anyway: JWTs and most URL-safe Base64 usage strip trailing = characters entirely, because a JSON Web Token's structure is delimited by . characters, not by relying on padding, and removing it saves a couple of bytes on every token. If you're writing a decoder, don't assume padding will be present. RFC 4648 explicitly allows omitting it, so a compliant decoder needs to handle unpadded input too.

A gotcha worth knowing before it bites you

The 33% overhead isn't just a number on a spec page. It has a habit of quietly pushing requests past size limits that were sized around the original file. A 1.5 MB image, once Base64-encoded and dropped into a JSON body, becomes roughly 2 MB of text before you've added any JSON structure around it. That's enough to clear a 2 MB request-body cap on plenty of API gateways and CDNs, producing a 413 Payload Too Large that looks unrelated to the image itself, because on disk, before encoding, the file was well under the limit. When a request size limit and a Base64-encoded payload are both in play, budget for the encoded size, not the original file size, or you'll chase a phantom bug.

What Base64 is not

It is not encryption. Anyone can decode Base64 instantly. It's a public, fixed, reversible mapping with no key involved. If you see credentials or sensitive data Base64-encoded and nothing else, treat it as plain text, because functionally it is.

It does not compress data. It's the opposite: Base64 output is always roughly 33% larger than the original binary input, because it's trading 8 bits of information density per byte for 6 (since each output character only encodes 6 bits from the 64-character alphabet).

If you need actual security, use real encryption (like AES) and only Base64-encode the encrypted output afterward, if you need it in text form. If you need actual size reduction, use a real compression algorithm (like gzip) before, not instead of, any Base64 step required for transport.

Try it

GlaeKit's Base64 Encoder/Decoder handles UTF-8 text and files entirely in your browser — nothing is sent to a server.

Frequently asked questions

Is Base64 the same as encryption?

No. Base64 is a reversible encoding with no key or secret involved — anyone can decode it instantly using any standard tool. It provides zero confidentiality.

Why does Base64 output end with = or == sometimes?

Base64 processes input in groups of 3 bytes at a time. When the input length isn't a multiple of 3, the last group is padded with = characters to keep the output a valid, fixed-length block — one = if there's one leftover byte, two if there are two.

Why is my encoded output bigger than the original file?

By design. Every 3 bytes of input become 4 bytes of output, a fixed ~33% expansion — this is the tradeoff for making binary data safe to carry over text-only systems.

What's the difference between standard Base64 and Base64URL?

Standard Base64 uses + and / in its alphabet, both of which have special meaning inside a URL. Base64URL swaps those for - and _ so the encoded string can be used directly inside a URL or filename without escaping.

Can I use a standard Base64 decoder on a URL-safe Base64 string?

Not reliably. The two alphabets only differ in two characters (+// vs -/_), but a decoder built for one will fail or produce corrupted output on any character from the other. If a token decodes almost correctly but a few characters look wrong, check whether you're mixing the two variants.

Why did my API request fail with a 413 error after I Base64-encoded a file?

The ~33% size increase from Base64 encoding can push a payload past a request-size limit that was set based on the original file size. A file well under the limit on disk can exceed it once encoded and wrapped in JSON — budget for the encoded size, not the raw file size, when working against a size cap.