If you've ever tried to email a PDF and hit a size limit, you've probably noticed something odd: two documents that look identical on screen can differ in size by 10x or more. That's because "PDF" isn't one format. It's a container that can hold wildly different things inside, and how it was created determines almost everything about its file size.
The confusing part is that file size has almost nothing to do with how much content is on the page. A 40-page contract with dense paragraphs of text can weigh 300KB. A single scanned receipt can weigh 4MB. Once you understand what's actually stored inside a PDF, that stops being surprising, and you can start making deliberate choices about what to strip out instead of just mashing "compress" and hoping for the best.
What actually takes up space in a PDF
People assume a bloated PDF is a text problem or a "the format is just inefficient" problem. Almost never. In the overwhelming majority of oversized PDFs, one thing accounts for 80-95% of the file: embedded raster images at full resolution and full bit depth.
The size math explains why. Vector text (the actual letterforms you're reading right now) is stored as a small set of drawing instructions plus character codes. A page of body text, even a dense one, is typically 2-10KB once the font program is shared across the document. A single embedded color photo at 300 DPI on a US Letter page, by contrast, is roughly 2550×3300 pixels. Stored raw, that's over 25 million pixels at 24-bit color: around 75MB uncompressed before any image compression is even applied. JPEG encoding brings that down enormously. But it's still routinely the single biggest object in the file, by two or three orders of magnitude over the text layer.
Vector graphics (logos, charts, line art drawn as paths rather than pixels) sit in between: they scale losslessly and stay small, usually a few KB to a few hundred KB even for complex illustrations, because they're still just drawing instructions, not pixel grids. So when you're hunting for what to shrink, the order of priority is almost always: images first, then embedded fonts, then everything else is rounding error.
Why PDFs get bloated in the first place
Most oversized PDFs fall into one of three categories:
Scanned documents. A "scan to PDF" from a printer or phone app usually embeds each page as a full-resolution, uncompressed (or lightly compressed) image, often 300 DPI or higher, even if the page is just black text on white. A single scanned page can easily be 2-5MB; a 20-page scanned contract can hit 80MB.
Exports from design tools. PDFs exported from InDesign, PowerPoint, or Keynote often embed images at their original resolution and color depth, plus fonts, vector art, and sometimes unused embedded assets left over from editing. It's common for a slide deck exported straight from PowerPoint to carry the original, uncropped photo behind a cropped placeholder. The software keeps the full image data even though only a corner of it is ever displayed, because crop is applied as a rendering instruction, not a destructive edit.
Redundant embedded fonts. Every font used in the document (including every weight and style) can get fully embedded, sometimes duplicated across pages if the PDF was assembled by merging several source files. A PDF stitched together from ten separately-exported chapter files can end up with ten near-identical copies of the same font subset, one per source file, because nothing deduplicated them at merge time.
Lossy vs. lossless: two different kinds of compression
"Compress the PDF" actually bundles together two very different techniques, and it's worth knowing which one you're using, because only one of them can ever make your images look worse.
Lossless techniques remove data that was never visible in the first place, so there's no quality tradeoff at all. This includes stripping unused objects (form fields nobody fills in, orphaned annotations), deduplicating embedded font subsets that appear more than once, discarding metadata (author, editing history, thumbnail previews, XMP packets that can run to tens of KB per file), and re-encoding the PDF's internal cross-reference structure more compactly. None of this touches a single pixel. If a compressor claims a big size reduction with "no quality loss," this is almost always where the savings came from. Or it's about to move into the second category without telling you.
Lossy techniques throw away information that was visible, betting you won't notice. The two main levers are recompression (re-encoding an embedded JPEG at a lower quality factor, which discards fine color and detail information permanently) and downsampling (reducing the pixel dimensions of an embedded image, e.g. from 300 DPI to 150 DPI, which throws away actual pixels rather than just re-encoding the ones that remain). Both are genuinely useful (most images are wildly over-resolved for how they'll ever be viewed), but both are one-way trips. Once a JPEG has been re-encoded at quality 40, you can't get the quality-90 version back by "decompressing" it; the discarded detail is gone for good.
Text itself is not compressed in the lossy sense — actual text characters are already tiny compared to images, so quality loss in a compressed PDF almost always shows up in images, not in the crispness of your words.
Why scanned PDFs compress so much better than text-heavy ones
This is the pattern that trips people up: run the same compressor on two different PDFs and get wildly different results — a scanned document might shrink 85%, while a report written in Word and exported to PDF barely moves. That's not the compressor being inconsistent; it's a direct consequence of what's inside each file.
A scanned page is, structurally, one giant photograph of a mostly-white rectangle with some dark marks on it. It's an enormous amount of pixel data (millions of pixels) representing very little actual visual information (black ink on white paper, high contrast, low color variation). That's exactly the profile compression algorithms are best at collapsing — huge redundancy, low entropy. Downsample a 300 DPI black-and-white scan to 150 DPI and re-encode it, and you can lose 80-90% of the file size with the page remaining perfectly legible, because you were storing far more resolution than a page of printed text ever needed.
A text-native PDF, by contrast, is already close to its minimum size. The text is vector data, not pixels, so there's no resolution to throw away. If the document has few or no embedded images, there simply isn't much lossy compression can do — you're left with the lossless gains (metadata, redundant fonts, structural cleanup), which are real but modest, typically single-digit to low-double-digit percentage reductions.
How to compress without visible quality loss
A few rules of thumb that avoid the "my PDF looks compressed" problem:
1. Start with structural compression first. Removing unused objects and redundant fonts costs zero visual quality — it's pure waste removal. If that alone gets you under your size limit, stop there.
2. Only downsample images if the document is screen-bound. If nobody is going to print the PDF at high resolution, dropping embedded images to 150 DPI is invisible on any modern screen but can cut file size dramatically.
3. Avoid re-compressing an already-compressed JPEG repeatedly. Every re-save of a JPEG at a lossy quality setting compounds artifacts. If your PDF is going to be compressed more than once (e.g., re-uploaded to different systems), it's worth compressing the source images once at a sane quality before ever putting them in a PDF.
4. For scanned documents, consider OCR + re-export instead of raw compression. A scanned page is fundamentally a photo of text. Running OCR (optical character recognition) and re-exporting as a proper text-based PDF can shrink a multi-megabyte scan down to a few hundred kilobytes, because you're storing actual characters instead of pixels.
One concrete data point worth having in your head: a typical 300 DPI color scan of a printed page, saved as JPEG quality 85 inside a PDF, sits around 400-600KB. Drop it to 150 DPI at quality 60 and it usually lands in the 60-100KB range (a roughly 5-8x reduction) while still looking clean at normal reading zoom on a laptop screen. Push past quality 40 or below 100 DPI and JPEG blocking artifacts start showing up around text edges and fine lines, which is the point where "compressed" starts to look like "compressed."
When not to over-compress
Aggressive compression is the right call for most day-to-day sharing, but there are cases where you should leave more headroom, or skip lossy compression entirely:
Anything headed for professional printing. Print shops typically want 300 DPI minimum, and some large-format or fine-detail print jobs want more. Downsampling to 150 DPI for email-friendly size will look fine on screen and noticeably soft on paper. If a file needs to serve both purposes, keep an uncompressed master and generate a compressed copy for distribution rather than compressing the only copy you have.
Legal and regulatory documents. Contracts, court filings, medical records, and anything that might later need to be authenticated or examined closely (signatures, seals, fine print, watermarks) should generally avoid lossy image recompression. Some jurisdictions and record-retention policies specifically require preserving the document as originally filed — lossless cleanup (metadata, structure) is usually fine, but re-encoding embedded images is not, since it technically alters the file's rendered content.
Documents with fine technical detail. Engineering drawings, maps, and anything with small dense text or thin hairlines embedded as an image (rather than as vector data) can lose readability fast under downsampling, even if the overall picture looks acceptable at a glance. Zoom in before you commit to a compression level, not just eyeball the thumbnail.
Archival copies. If a PDF is going into long-term storage rather than being actively shared, the incentive to shrink it is much weaker — storage is cheap, and you may not get a second chance to go back to a higher-quality source later. Compress the copy you're distributing, not necessarily the one you're keeping.
Try it
GlaeKit's PDF Compressor runs entirely in your browser — it re-structures and optimizes the file locally, so nothing is ever uploaded to a server. Drop in a file and see the size difference before you download.
Frequently asked questions
Will compressing a PDF make the text blurry?
No — text in a PDF is stored as vector character data, not pixels, so it stays sharp at any zoom level regardless of compression. Blurriness after compression almost always comes from downsampled images, not text.
How much smaller can a PDF realistically get?
It depends entirely on the source. A well-optimized PDF exported directly from a text editor might only shrink 5-10%. A scanned document full of high-resolution images can often shrink 70-90% with no visible quality loss on screen.
Is it safe to compress a PDF containing sensitive documents?
It depends on the tool. Many online compressors upload your file to a server to process it, which means your document leaves your device. Browser-based tools that process everything locally (like GlaeKit's) never transmit the file anywhere.
Why does my PDF get bigger after I edit and re-save it?
Most PDF editors append changes rather than fully rewriting the file, to save time. Each edit adds new objects without removing the old ones, so a PDF that's been edited many times can accumulate a lot of dead weight — compressing it forces a full rewrite that discards the unused data.
What actually makes a PDF so large — is it the text or the images?
Almost always the images. Vector text is stored as compact drawing instructions and character codes, typically a few KB even for a dense page. A single embedded photo at print resolution can be tens of megabytes uncompressed, and even after JPEG encoding is usually the largest object in the file by a wide margin. When a PDF is unexpectedly huge, check the embedded images first.
Should I ever avoid compressing a PDF?
Yes — hold off on lossy compression for files headed to professional printing, legal or regulatory documents that must preserve their originally-filed content, and anything with fine technical detail like engineering drawings or maps. Lossless cleanup (removing metadata, deduplicating fonts) is safe in all of these cases; downsampling or recompressing images is not.