What compressing a PDF actually does
A PDF is a container, not an image. Whether it can be made smaller, and by how much, depends entirely on what is inside it.
Published Updated 4 min read
A PDF is a container of objects
A PDF file is a collection of numbered objects — dictionaries, arrays, numbers, strings and streams — plus a cross-reference table that records where each object begins. Pages are objects that point at content streams, fonts, images and other resources. There is no single compression setting to turn down, because there is no single thing being compressed.
This is why the same request, "make this PDF smaller", produces wildly different results on different files. The honest answer always depends on where the bytes currently are.
Find the bytes before compressing anything
In practice a PDF’s size comes from one of three places. Scanned documents are photographs in a PDF wrapper: effectively all the weight is in embedded images. Documents exported from a word processor or a design tool are mostly text drawing operators plus embedded font programs. Documents that have been edited many times, passed between tools, or assembled from other documents accumulate structural overhead: orphaned objects, incremental update sections, duplicated resources and an inefficient cross-reference layout.
Each of those calls for a different technique, and a tool that only implements one of them will appear to work brilliantly on some files and do nothing at all on others.
Structural rewriting: lossless and safe
The safest technique is to parse the document and write a clean copy. Objects that nothing references are dropped. Incremental update sections, which append revisions to the end of a file rather than modifying it in place, are collapsed into a single current state. Small objects are packed into object streams, and the cross-reference table is written as a compressed cross-reference stream rather than a plain-text table. Content streams are compressed with Flate.
Nothing visible changes, because no image is re-encoded and no font is altered. On a document that has accumulated bloat, the saving can be substantial. On a freshly exported, already well-formed PDF, the saving may be close to zero — and occasionally the rewritten file is marginally larger, because the original was already optimal and the new structure has its own overhead.
This is exactly what the Compress PDF tool here does. It is a structural rewrite, so it will never soften an image or substitute a font, and it will never make a scanned document dramatically smaller.
Image re-encoding: where the big numbers come from
When a service advertises a very large reduction on a scanned document, it is almost always re-encoding the images. That means two things happening together: downsampling, which reduces the stored pixel dimensions to something closer to the resolution the page is actually viewed or printed at, and re-encoding, which re-compresses the pixels with a lossier setting than before.
Both are destructive. Downsampling a 600 DPI scan to 150 DPI discards three-quarters of the samples in each direction, and the detail does not come back. Re-encoding runs the image through a lossy codec a second time, adding a new generation of artefacts on top of whatever the scanner produced. For a document destined to be read on screen and then discarded, that is a good trade. For an archival copy, a legal exhibit or anything that may be enlarged, it is not.
If a tool offers you a percentage but no control over target resolution or quality, be cautious about what it is deciding on your behalf.
Fonts, and what subsetting can and cannot do
Embedded fonts are complete programs, and a single family with several weights can occupy a meaningful share of a small document. Subsetting keeps only the glyphs the document actually uses and discards the rest. It is lossless for the document as it currently reads, but it does make later editing harder: if you subsequently type a character whose glyph was removed, the viewer has to substitute another font.
Merging a document made from several sources frequently leaves multiple near-identical copies of the same font embedded. Deduplicating those is a genuine, safe saving — but it requires the tool to be confident that the font programs really are equivalent, which is not always decidable.
Reasonable expectations
A text-heavy PDF that has been edited repeatedly often responds well to a structural rewrite. A clean export from a modern application usually does not, because the exporter already did the same work. A PDF that is mostly high-resolution photographs will barely move under any lossless technique, because JPEG data inside the container is already compressed and Flate cannot compress it further.
If a scanned document must get much smaller, the only real lever is resolution. Decide deliberately what resolution the document needs, apply it once, and keep the original. And before reaching for compression at all, check the obvious: a 40-megabyte PDF is quite often a 40-megabyte photograph that someone placed on a page at full size.