[nonimo]
EN
Download

Redact a PDF before AI reads it

· Updated · Written and maintained by Nonimo

A black rectangle on a PDF looks final. On the screen the client’s name has gone, the printout agrees, and the file goes into ChatGPT for a summary. If the box was drawn rather than applied, the name went with it, because the rectangle was painted over the characters and every one of them is still inside the document.

This guide explains how to redact a PDF with the characters removed rather than hidden, using a test file we built and then pulled apart. It covers Acrobat and Preview, the three checks worth doing before any upload, and the scanned document, which fails in the opposite direction.

The Australian angle is simple. The OAIC treats typing personal information into a public chatbot as a disclosure, and a hidden name you did not mean to send is still sent. That makes a bad redaction a privacy problem rather than a formatting slip.

How to redact a PDF: four rules that hold

The short version fits in four lines, and everything else on this page explains why each line is there.

  1. Use the redaction function, never a pen or shape. A tool built for redaction deletes the characters under the mark. A rectangle, a black highlight or a shape only covers them.
  2. Apply the marks and save under a new name. In most tools the mark is a proposal until you confirm it, and the words stay until then.
  3. Test what you saved. Try to copy the black area, search for a word you removed, and select all the text on the page.
  4. Treat scans on their own terms. A scanned page has no text to delete, so the same tools and the same tests behave differently.

When the destination is a chatbot rather than a registry, a fifth route skips the file altogether: give the model only the text it needs, with identifiers already replaced. Redacting the file properly still matters on that route, because the source document stays on your system and gets shared in other ways later.

Which identifiers to take out is a separate question from how to take them out. For a progress note, the NDIS walkthrough sets out what comes out and in which order, and the global guide to masking a document covers the same ground for any file.

The black box and the text underneath it

A PDF page is closer to a stack of instructions than a picture. One instruction writes “Tamsin Achterberg” at a position on the page. Another fills a black rectangle at the same position. The viewer carries out both, in order, and you see black.

Nothing in that sequence deletes the first instruction. Anything that parses the document instead of displaying it, a search box, a copy and paste, a text extractor, an AI upload, goes to the text instruction and finds the name exactly where it was. A black rectangle laid over a name on a slide fails in exactly this way, and a PowerPoint test deck still gave the name back from under it.

A drawn box sits above the text layer Two layers of the same page. The top layer is a black rectangle that the screen shows. The bottom layer is the text Client: Tamsin Achterberg, which a text extractor and an AI upload still read. Client: Client: Tamsin Achterberg What the screen shows the drawing on top What an upload reads the text layer, name included
One page, two layers. Invented test letter, built with ReportLab and read with pypdf, 28 September 2026

The test file we built

To check this rather than repeat it, we made a one page letter from an invented tax practice, Wattle Creek Tax Advisers, to an invented client. It carries the usual Australian details: client name, birth date, street address, tax file number and mobile. The numbers are test values in valid formats, chosen so that they do not belong to anyone.

We then produced three versions. In the first we drew a black box over each of the five details, the way a markup tool does. In the second we used a real redaction function and applied it. The third is the original page turned into a picture, which is what a scanner produces.

What came out of each version

We ran each file through pypdf, a standard Python library that pulls the text out of a PDF. That is the same job any tool does when it reads a document rather than displaying it. This is the text it returned from the version with the black boxes:

Wattle Creek Tax Advisers
Private and confidential
Client: Tamsin Achterberg
Date of birth: 14/07/1986
Address: 22 Banksia Rise, Wirrimbah VIC 3999
Tax file number: 000 000 000
Mobile: 0491 570 156

Every covered detail came back, in order, with nothing to show that it had ever been covered. The applied redaction returned the labels and blank space after them. The scanned version returned no text at all, which is its own problem and has its own section below.

5 of 5covered details extracted from the box version
0 of 5after the redaction was applied
0characters of text in the scanned version
Our test letter, pypdf 6.10.2, 28 September 2026

The client’s name, date of birth and tax file number are exactly what a tax practice is not meant to hand to an AI tool. The black boxes made the letter look safe and changed nothing about what a machine reads.

What Australian courts and the LPLC already tell practitioners

None of this is new to the courts. The Federal Court of Australia publishes a Guide to redacting documents in electronic form for practitioners filing documents, and its warnings are the same failure we reproduced, described from the receiving end.

The clearest summary for Australian readers comes from the Legal Practitioners’ Liability Committee, the Victorian professional indemnity insurer for lawyers. In its article Write tech, wrong text, last updated on 11 June 2026, it says that practitioners’ failure to properly redact electronic court documents “has resulted in third parties uncovering and reading text that should have been deleted”.

What the guide tells you to do, and not to do

The LPLC lists the Federal Court’s key points. Put side by side, they read as a short rule book for anyone preparing a file for someone else.

the guide sayswhy
Replace the text with [Redacted]the words are gone, not covered
No black shading over the textthe text is not removed
No white text colourit can still be copied and searched
No black boxes from Adobe’s annotation toolthe text can still be found
Or delete the text in a copy and check it in Notepadproves nothing is left
Redact scans with software made for ita picture needs a different tool

Source: LPLC, Write tech, wrong text, summarising the Federal Court of Australia’s guide.

The article adds a point that matters for anyone who scans paperwork, a NSW driver licence included. Where a document “is converted to a searchable PDF, the shaded text can still be searched”. A scanner with text recognition switched on rebuilds the very layer a black box fails to reach.

The guide was written for documents going to a court registry and then possibly to the public. A chatbot upload reaches a smaller audience with a longer memory, and the same mechanics apply to it without any change. Courts are also writing rules about what goes into generative AI at all, and the practice note from the NSW Supreme Court is where most practitioners meet them first.

What ChatGPT and Claude receive from an uploaded PDF

What happens inside a model is the provider’s business, and we will not guess at it. What the providers say publicly about how they take a PDF in is enough to answer the question that matters here, which is whether the covered words arrive.

ChatGPT: the digital text, and on Enterprise the page as well

OpenAI’s help article on file uploads is direct about it. “ChatGPT Enterprise supports Visual Retrieval for PDF files.” For everything else it says: “All other plans and document files only support text-based retrieval. This means that ChatGPT will extract digital text from the file and discard any images.”

Extracting the digital text is exactly what pypdf did to our letter. On those plans the drawn box is thrown away with the rest of the drawing, and the name is part of the text that is kept. Retention after the upload is its own subject, and the ChatGPT data guide tracks what OpenAI keeps and deletes.

Claude: a picture of each page, plus its words

Anthropic’s developer documentation on PDF support describes the process in two steps. “The system converts each page of the document into an image. The text from each page is extracted and provided alongside each page’s image.”

That documentation covers PDFs sent through Anthropic’s developer platform, and it is the most detailed public description the company gives. Read literally, the picture of the page would show black where the box is, while the extracted words sent next to it would still carry the name. Two inputs arrive, and the redaction covers just one.

What ChatGPT and Claude receive from a PDF with a drawn black box For ChatGPT outside Enterprise, OpenAI says the digital text is extracted and images discarded, so the name arrives as text. For Claude through Anthropic's developer platform, each page goes in as an image, where the box shows black, and as extracted text, where the name is still present. The PDF black box drawn over the name Extracted text Client: Tamsin Achterberg Page image Client: [black] ChatGPT other plans: text only Claude API: image and text
Both routes carry the text layer. OpenAI File Uploads FAQ; Anthropic, PDF support

We did not find an equivalent public description for Microsoft Copilot, so we make no claim about it here. The safe assumption for any tool is the one the test file supports: if the text is in the PDF, the tool can read it. What Microsoft does publish about data handling is collected in the Copilot guide.

Sending the passage instead of the file

Most questions about a document do not need the whole document. A summary of one clause, a reply to a letter or a check on a calculation needs that passage, not the letterhead, the signature block and the footer carrying the client number. Copying the relevant paragraphs into the prompt, with the identifiers in them replaced, leaves the rest of the file on your side, including the page you forgot was there.

It also takes the layers out of the question. Pasted text is exactly the text you can see, with nothing drawn on top and nothing hiding underneath, so what you checked is what goes. The cost is effort: for a forty page contract, a proper redaction of the file may be quicker than choosing passages by hand.

Acrobat Pro in six steps: how to redact a PDF

Adobe documents redaction as a feature of Acrobat Pro, and its Australian guide walks through six steps. The names below are the ones Adobe uses on that page. Menus move between versions, so if a label differs on your screen, search the tools list for “Redact”.

  1. Keep the original. Adobe’s own advice is to save a copy before you start, because redactions are irreversible once applied. Work on the copy.
  2. Open the redaction tool. From the Edit menu choose Redact a PDF, or right click and choose Redact.
  3. Set how the marks look, if you want to. Set properties in the Redact toolset changes how the marks look, for instance the fill colour or overlay text such as [Redacted], the wording the Federal Court prefers.
  4. Mark the content. Double click a word, or drag across a block of text or an image. Every name, number and address you want gone needs its own mark.
  5. Repeat across pages where needed. A right click option applies the same mark on other pages, which helps with letterheads and repeated reference numbers.
  6. Apply, then save. Choose Apply or Apply & Save. Until you do, the marks are only proposals and the text is still underneath.

The step that gets skipped

Step six is the easy one to miss. A marked document looks finished: the marks are visible, the colour is right, and it can be saved and sent. Only after Apply does Acrobat actually take the characters out. Word has the same trap with its markup, because a clean view of tracked changes still leaves them in the file.

The other common failure is using the Comment tools instead. A rectangle or a highlight from the commenting panel is an annotation, the “annotation tool” the Federal Court guide tells practitioners not to use. It sits on top of the page like a sticky note. For practices that redact every week, such as law firms preparing court bundles, having a second person open the saved file is a cheap safeguard.

Search and redact for repeated details

A client’s name can appear in a letterhead, a salutation, a table and a footer. Acrobat can search the document for a word or phrase and mark every match for redaction, which is safer than hunting by eye. It will still miss a name split across two lines or written differently, “T. Achterberg” or “Tamsin”, so search for each form.

A redacted PDF can still describe itself in its properties, the author for instance, and that is a separate job from the one on this page.

If the PDF started life as a Word file, that file has its own hiding places, and redacting the PDF does nothing to them.

Preview on a Mac, and the tools that only draw

Mac users have a built in option. Apple’s Australian guide to annotating a PDF in Preview describes a Redaction Selection tool: “Select text to permanently remove it from view. You can change the redaction as you edit but once you close the document, the redaction becomes permanent.”

Two details in that sentence matter. The redaction is editable while the file is open, and it becomes permanent on closing. So the test belongs after you close the document and open it again, not while you are still editing. “Remove it from view” is Apple’s phrase, and the only way to know what happened to the characters is the copy test below.

The tools that look like redaction

Most PDF software can draw a black rectangle, and most offer a black highlighter. Neither is redaction. The table sorts the common options by what they do to the text, not by what they are called.

what you usedwhat happens to the wordssafe for an upload
Acrobat Pro, Redact a PDF, then Applyremovedyes, after testing
Preview, Redaction Selection, file closedApple: permanently removed from viewtest it
Rectangle or shape from the comment toolscovered, still in the fileno
Black highlightercovered, still in the fileno
White text or white boxinvisible, still in the fileno
Online redaction servicethe full file is uploaded first, and what comes back variescheck what it keeps and who can see it, then test

Sources: Adobe Acrobat Australia; Apple Support Australia; LPLC summary of the Federal Court guide.

Any online redaction service receives the whole file before it covers anything, so three things matter: what is uploaded, whether it is kept, and who can see it. The tool above this guide covers the document and gives you the result, and the document is discarded as soon as you get it. For a privileged client document or a file carrying a tax file number, the Nonimo app does the covering locally, and the file stays on your computer from start to finish.

Three tests for a redacted file before the upload

No redaction is finished until the saved file has been tested, and the tests take less than a minute. Do them on the file you are about to upload, not the one you worked on, and after closing and reopening it.

testhowa failure looks like
Copy and pasteselect across the black area, paste into Notepad or TextEditthe covered words appear
Searchsearch the PDF for a surname or a number you removedthe viewer finds a match
Select allselect all text on the page and paste itthe full letter comes back

The first test is the one the Federal Court guide describes, as summarised by the LPLC.

Plain text editors matter for the first test. Pasting into Word or an email can carry formatting across and make a result harder to read, while Notepad shows only characters. If the black area pastes as nothing, or as the word [Redacted], the text is gone. A workbook needs a different test, because Excel never lists a very hidden sheet, and checking a spreadsheet from outside Excel is the one that finds it.

What the tests cannot see

The tests check the text layer, and that is where uploads look first. They do not check a picture inside the PDF, so a photographed driver’s licence or a signature image is outside their reach. They also do not tell you whether a sentence that names nobody still identifies the client, which is a judgement no tool makes for you.

For health and disability documents that second limit is the one that usually bites. A diagnosis, a town and a month can point at one person with every name removed, as our guide to sensitive information explains with the categories the Act treats differently.

A scan is a photograph of paper saved inside a PDF, like a scanned drivers licence. There are no characters in it, only pixels, so a text extractor finds nothing. Our scanned version of the test letter returned zero characters of text, and the three tests above all come back empty, which looks like success and proves nothing.

The information is still there, visible to anyone and anything that looks at the picture. Whether an AI tool reads that picture depends on how the provider takes the file in.

When a picture is read as a page

Anthropic’s documentation, quoted above, turns every PDF page into an image and gives the model both. On that route the scanned letter is read from the picture, and every detail on the paper is in front of the model. OpenAI, for its part, limits visual reading of PDFs to ChatGPT Enterprise, with other plans keeping the digital text and dropping images.

So the same scan can arrive in full in one tool and arrive empty in another. Clinics meet this every day with scanned referrals and pathology reports, and the Ahpra rules for clinicians using AI apply to those same files. For a document you cannot afford to leak, the only planning assumption that holds across tools is that the picture will be read. The same goes for a phone photo of the page, and what each tool keeps of an uploaded photo is covered separately.

Scanned PDF: why the redaction tests come back empty A scanned page is a picture. A text extractor returns zero characters, so copy, search and select all find nothing. A tool that reads the page as an image still sees every detail on the paper. Scanned page pixels, no text Text extractor 0 characters Reading the image every detail visible empty tests look like success and prove nothing
Our scanned test letter returned no text. Anthropic, PDF support; OpenAI File Uploads FAQ

Redacting the picture itself

For a scan, or a photographed NSW licence, the redaction has to change the pixels. In Acrobat a redaction mark can be dragged over an area of an image, and applying it removes what sits under the mark. The Federal Court’s advice, through the LPLC, is to redact a copy with software built for that job. Then look at the result at full zoom, because the tests that work on text tell you nothing here.

Be careful with text recognition. Many scanners and PDF programs can make a scan searchable, which adds an invisible layer of recognised text over the picture. That turns a scan back into the problem this guide started with: a black box painted on the image can leave the recognised words untouched underneath, which is the point the LPLC makes about searchable PDFs.

One missed tax file number is still a disclosure

The OAIC settled what an upload is in its October 2024 guidance on off the shelf AI products, revised in January 2025. Its worked example has insurance staff typing a claim, health details included, into a public chatbot, and the regulator’s verdict is that the company “is disclosing the information to the owners of the chatbot”.

That disclosure has to satisfy APP 6. Its best practice advice goes further, and says organisations are better off entering no personal information at all, “and particularly sensitive information”, in generative AI tools open to the public. A name left under a black box is personal information that you entered, even though you believed you had not.

Why a missed TFN costs more than a missed name

Tax file numbers sit under their own rules, and a business that is otherwise exempt from the Privacy Act can still be caught by them. The Privacy Act and AI tools guide explains who counts as a file number recipient, and why holding TFNs brings an otherwise exempt practice under the notification rules.

Whether a particular slip has to be reported is its own assessment with a clock attached, and our breach guide walks through that test.

If the file has already gone

Stop, then establish what went. Open the copy you actually uploaded, not the original on your desk, and run the three tests on it: the result tells you which details the tool received. Deleting the conversation tidies your own history, but what the provider holds after a deletion depends on its terms, and the guide to Anthropic’s data rules shows how much those terms vary by plan.

If the details are personal information, the notifiable data breach scheme may be in play. Section 26WH of the Privacy Act expects an entity that suspects an eligible data breach to assess it promptly and reasonably, and to do what it reasonably can to complete that within 30 days. Note what went, when and to which tool while you still remember, because that note is where the assessment starts.

The same client letter, with labels instead of boxes

Nonimo works from the other end. Rather than editing layers, you hand it the document and receive two things: the text with identifiers turned into labels, ready to paste into a prompt, plus a covered version of the file. When the input is a PDF, what you get back is a clean PDF of the text, with the original’s metadata stripped out.

This is the text of the test letter, as Nonimo returned it:

BEFORE                                          AFTER (Nonimo)
Wattle Creek Tax Advisers                       Wattle Creek Tax Advisers
Client: Tamsin Achterberg                       Client: [PERSON_1]
Date of birth: 14/07/1986                       Date of birth: [BIRTH_DATE_1]
Address: 22 Banksia Rise, Wirrimbah VIC 3999    Address: [RECORD_FIELD_1]
Tax file number: 000 000 000                    Tax file number: [TFN_1]
Mobile: 0491 570 156                            Mobile: [PHONE_1]
Refund of $2,418                                Refund of $2,418

The labels go to the model, and the Nonimo app puts the real values back into the reply on your side. Because the link back to the client survives on your side, this is pseudonymisation; the comparison of the two terms explains why that matters in Australia. The firm name and the refund stayed, because the model needs them to answer.

The Nonimo app reads a fully scanned PDF on the computer, on Mac and Windows, and gives you its text, covered, ready for a prompt. A scan you send as a file still needs the picture treatment above. What the app stores on your computer, and the daily usage count it sends us, are described on the security page.

Sources

Nonimo is the software that does this on your own computer: it masks client names and IDs before your text reaches ChatGPT. No account needed, and the app does it without your files leaving your machine.

Common questions

Which redaction method actually strips the text out of a PDF?

The short answer to how to redact a PDF: pick the redaction function of your PDF software, never its drawing or highlighting tools. In Acrobat Pro you mark each detail and confirm with Apply, which Adobe describes as irreversible. Reopen the saved file and paste the black area into Notepad to prove nothing came across. Nonimo works from the other end, handing back labelled text.

Can ChatGPT see words hidden behind a black rectangle?

Yes, when the rectangle was merely painted over them. OpenAI states that on plans other than Enterprise, ChatGPT pulls the digital text out of an uploaded file, and a painted shape carries no text of its own, so it screens nothing from that step. Delete the characters first, or give the model Nonimo's labelled version in place of the original.

Does a black highlight or a filled shape remove the words?

No. A black highlight or a filled shape changes what you see on the screen and on paper, while the words stay in the file. The Federal Court's guide, as summarised by the LPLC, warns against black shading and white text for exactly that reason. Real redaction removes the characters, and a copy and paste test shows which one you did.

Can you redact in the free Adobe Acrobat Reader?

Adobe documents redaction as an Acrobat Pro feature, under the Redact a PDF tool, so the free Reader is not where it lives. On a Mac, Preview has its own redaction selection. Whatever you use, open the saved file afterwards and try to copy the covered area, because the tool's name tells you nothing about what it did.

What is the quickest way to check a redacted PDF?

Three checks, all on the file you saved. Paste the black area into a plain text editor, search for a surname you removed, and select everything on the page. If the name surfaces anywhere, the redaction failed. On a scanned page all three come back empty whatever is visible, because a picture holds no characters, so look at a scan instead of searching it.

Can an AI read a scanned PDF?

That varies by provider. Anthropic's developer documentation says every PDF page becomes an image that the model reads together with any text, so a scan is understood from the picture. OpenAI reserves visual reading of PDFs for ChatGPT Enterprise, and other plans keep only digital text. The Nonimo app reads a fully scanned PDF on the computer, on Mac and Windows, and returns its text with the details covered.

What does the Federal Court say about redacting electronic documents?

Its guide to redacting documents in electronic form, summarised by Victoria's Legal Practitioners' Liability Committee, says to replace the text with [Redacted], not to use black shading or white text, and not to draw black boxes with Adobe's annotation tool. For scans it says to redact a copy with software built for that purpose.

Is a leaked name under a black box a privacy incident in Australia?

It can be. In the OAIC's view, personal information typed into a public chatbot has been handed to the company that runs it, and APP 6 then decides whether that was allowed. A name you thought was hidden has travelled all the same. Whether it must be reported is a separate test. Nonimo replaces recognised identifiers with labels before the chatbot sees the text.