Scanned documents are the most compressible PDFs of all — because every page is a photograph of paper. That also makes them the easiest to ruin if you over-compress. Here's how to get a small, still-readable file.
Why scanned PDFs are so big
A scanner captures each page as a high-resolution image. A colour scan at 600 DPI can be several megabytes per page. The text isn't 'text' to the computer — it's pixels — which is why scanned PDFs aren't searchable until you run OCR.
How to shrink them safely
Downsample the page images to around 150 DPI and re-encode them at a moderate JPEG quality. This usually keeps text clearly legible while cutting size enormously. Convert to grayscale if colour isn't needed.
Make it searchable too
If you need to search or copy text from a scan, run OCR (optical character recognition) to add a text layer. This is separate from compression but often wanted at the same time.
Step by step
- Confirm it is a scan by trying to select text — if nothing highlights, it is images.
- Convert to grayscale first if the colour carries no meaning. This is the single biggest saving.
- Compress to the target size you need.
- Run OCR afterwards if you want the document to be searchable.
What usually goes wrong
Grayscale before compression, not after
Removing two of three colour channels first gives the compressor much less to work with, and the result is better at the same size.
Scans have a floor
Below a certain size a scan becomes blocky and the text starts to break up. If you are fighting it, the document needs fewer pages, not more compression.
Rescanning often beats compressing
A 600 DPI colour scan compressed hard looks worse than a fresh 300 DPI grayscale scan, and takes about the same time.
Frequently asked questions
Can I compress a scan to 100KB?
Often yes for one or two pages in grayscale. Many pages or full colour may not reach 100KB while staying readable.
Why are scanned PDFs so large?
Because every page is a photograph rather than text. A page of text is a few kilobytes; a scan of the same page can be hundreds.
What is the best order: grayscale, compress, or OCR?
Grayscale, then compress, then OCR. Recognition works on the final images, so running it last means it matches the file you are actually sending.