hddn Hide what matters.

Guides07 of 07

How to redact a scanned PDF

A scan has no text to select, so the usual advice fails. What actually works on a bitmap page, why flattening matters, and the hidden layer that catches people.

Last updated: September 18, 2026

07

You have a photographed contract, or a passport copy, or a payslip somebody scanned at the office. One name has to go. Every guide says the same thing: select the text, apply redaction. You drag across the name and nothing highlights. The cursor doesn’t even change shape.

Nothing is broken. There is no text there to select.

01

The page is a photograph

A native PDF stores text as instructions, one per glyph, with coordinates and a font. A scan stores one image per page and nothing else. The name you are looking at is a cluster of dark pixels inside a JPEG. As far as the file is concerned, it is the same kind of object as the shadow of the scanner lid.

So selection has nothing to grab. Selection tools read the content stream, and on a scanned page the content stream says roughly: draw this image, full width, full height. That is the whole page. Search finds nothing and extraction comes back empty.

02

Here, the black box can actually work

This is the one case where the obvious approach is salvageable, which is not what people expect after being told a hundred times that rectangles don’t redact.

On a genuine scan there is no text hiding underneath, because there is no text anywhere. If you paint a filled rectangle over the pixels and then flatten the page, the pixels under the box are overwritten by the box. They are not covered. They stop existing.

The condition is flattening, and it is not optional. Your export has to render each page to a new bitmap with the rectangle already burned in, then rebuild the PDF from those images. If the tool keeps the original scan and stacks a shape above it, you now have a black rectangle sitting on top of an intact photograph of the name, and pulling the embedded images out of a PDF is a single command for anyone who thinks to try.

03

Check whether your scan is really a scan

Plenty of documents that look like scans are not pure scans any more. Scanner software, document management systems, and anything with a “searchable PDF” option will run OCR and write the recognised words back into the file as an invisible text layer, positioned to sit exactly over the image. The page looks identical. The file now holds both a picture of the name and a real, extractable string of it.

Draw a box on the image and you have covered the pixels for the eye. The text layer is untouched and comes straight out with copy-paste, which is the black rectangle problem with a photograph laid over it.

Test it before you do anything else. Open the file and drag across a line of body text. If a highlight appears, or Ctrl+F finds a word you can see on the page, there is a text layer and you must remove that content as well as the pixels. Check page by page, because mixed documents are common: a typed cover sheet, then twelve scanned pages, then a signed appendix.

Then run that same check on the exported file, not the one you were editing.

04

An annotation is not a flattened box

A rectangle saved as a PDF annotation is a separate object in the file, with its own coordinates, colour and entry in the page’s annotation list. Readers let you click it, drag it, delete it. Even where the interface won’t, the object is sitting in the file structure and stripping annotations is trivial.

Flattening ends that argument. The shape stops being an object and becomes dark pixels in the page image, with nothing left to select or delete.

One thing trips people up here: printing to PDF is not flattening. On most systems the print pipeline re-emits text as text and keeps annotations as annotations. Look for wording about rasterising, flattening, or exporting as image, then verify the result rather than trusting the label.

05

Resolution, and the copy you forgot about

Some export paths downsample. Privacy is fine either way, but a redacted scan at low DPI can end up unreadable for the recipient, so look at the output before you send it rather than after they ask again.

The privacy problem is elsewhere in the file. Scanner software and office suites sometimes embed a thumbnail or preview image of each page, and a file saved incrementally can keep earlier revisions inside it. If that preview was generated before you drew the box, it is a small copy of the unredacted page riding along in the same document. Export to a fresh file instead of saving over the original, and open the document properties once before it leaves your machine.

06

OCR, used deliberately

Running OCR yourself is what makes this bearable. Recognise the text first and a tool can read the page well enough to propose candidates: names, ID numbers, dates of birth, account numbers. The boxes land on the right coordinates because recognition returns a position for every word, not just a transcript.

OCR misreads, though. Faint fax copies, skewed pages, handwriting in the margin, a 5 read as an S. Treat the candidates as a starting point and read the page yourself, because nothing downstream catches what the recogniser missed. And the export still has to flatten. Finding the name accurately does nothing if the box is only stacked on top of it.

hddn works this way for scanned pages: OCR runs in the browser, detections come back as suggestions you approve or reject, and the flattened export path renders each page to a bitmap with the confirmed boxes burned in. It refuses to edit text on OCR pages rather than pretending it can.

A scan feels safer than a text PDF because nothing selects and nothing searches. That feeling is the risk. Drag across one line before you draw a single box, and find out whether the file has been quietly transcribing itself this whole time.

How it works

Four steps. You stay in charge of every one.

  1. 01

    Open your PDF

    Drop the file onto the page. Your browser opens it right on your device.

  2. 02

    We mark candidates

    Names, emails, IDs, account numbers, and dates that may need to be hidden get highlighted.

  3. 03

    You review everything

    Confirm the marks, reject the ones you disagree with, or add your own by blacking out any text or area. On digital PDFs you can even edit the words. Nothing is hidden until you approve it.

  4. 04

    Export a clean PDF

    The confirmed data is permanently removed, not just covered, and saved back to your device.