· Updated · Written and maintained by Nonimo
A black rectangle on a PDF looks final. On the screen the client’s name has gone, the printout agrees, and the file goes into ChatGPT for a summary. If the box was drawn rather than applied, the name went with it, because the rectangle was painted over the characters and every one of them is still inside the document.
This guide explains how to redact a PDF with the characters removed rather than hidden, using a test file we built and then pulled apart. It covers Acrobat and Preview, the three checks worth doing before any upload, and the scanned document, which fails in the opposite direction.
The Australian angle is simple. The OAIC treats typing personal information into a public chatbot as a disclosure, and a hidden name you did not mean to send is still sent. That makes a bad redaction a privacy problem rather than a formatting slip.
How to redact a PDF: four rules that hold
The short version fits in four lines, and everything else on this page explains why each line is there.
- Use the redaction function, never a pen or shape. A tool built for redaction deletes the characters under the mark. A rectangle, a black highlight or a shape only covers them.
- Apply the marks and save under a new name. In most tools the mark is a proposal until you confirm it, and the words stay until then.
- Test what you saved. Try to copy the black area, search for a word you removed, and select all the text on the page.
- Treat scans on their own terms. A scanned page has no text to delete, so the same tools and the same tests behave differently.
When the destination is a chatbot rather than a registry, a fifth route skips the file altogether: give the model only the text it needs, with identifiers already replaced. Redacting the file properly still matters on that route, because the source document stays on your system and gets shared in other ways later.
Which identifiers to take out is a separate question from how to take them out. For a progress note, the NDIS walkthrough sets out what comes out and in which order, and the global guide to masking a document covers the same ground for any file.
The black box and the text underneath it
A PDF page is closer to a stack of instructions than a picture. One instruction writes “Tamsin Achterberg” at a position on the page. Another fills a black rectangle at the same position. The viewer carries out both, in order, and you see black.
Nothing in that sequence deletes the first instruction. Anything that parses the document instead of displaying it, a search box, a copy and paste, a text extractor, an AI upload, goes to the text instruction and finds the name exactly where it was. A black rectangle laid over a name on a slide fails in exactly this way, and a PowerPoint test deck still gave the name back from under it.
The test file we built
To check this rather than repeat it, we made a one page letter from an invented tax practice, Wattle Creek Tax Advisers, to an invented client. It carries the usual Australian details: client name, birth date, street address, tax file number and mobile. The numbers are test values in valid formats, chosen so that they do not belong to anyone.
We then produced three versions. In the first we drew a black box over each of the five details, the way a markup tool does. In the second we used a real redaction function and applied it. The third is the original page turned into a picture, which is what a scanner produces.
What came out of each version
We ran each file through pypdf, a standard Python library that pulls the text out of a PDF. That is the same job any tool does when it reads a document rather than displaying it. This is the text it returned from the version with the black boxes:
Wattle Creek Tax Advisers
Private and confidential
Client: Tamsin Achterberg
Date of birth: 14/07/1986
Address: 22 Banksia Rise, Wirrimbah VIC 3999
Tax file number: 000 000 000
Mobile: 0491 570 156
Every covered detail came back, in order, with nothing to show that it had ever been covered. The applied redaction returned the labels and blank space after them. The scanned version returned no text at all, which is its own problem and has its own section below.
The client’s name, date of birth and tax file number are exactly what a tax practice is not meant to hand to an AI tool. The black boxes made the letter look safe and changed nothing about what a machine reads.
What Australian courts and the LPLC already tell practitioners
None of this is new to the courts. The Federal Court of Australia publishes a Guide to redacting documents in electronic form for practitioners filing documents, and its warnings are the same failure we reproduced, described from the receiving end.
The clearest summary for Australian readers comes from the Legal Practitioners’ Liability Committee, the Victorian professional indemnity insurer for lawyers. In its article Write tech, wrong text, last updated on 11 June 2026, it says that practitioners’ failure to properly redact electronic court documents “has resulted in third parties uncovering and reading text that should have been deleted”.
What the guide tells you to do, and not to do
The LPLC lists the Federal Court’s key points. Put side by side, they read as a short rule book for anyone preparing a file for someone else.
| the guide says | why |
|---|---|
| Replace the text with [Redacted] | the words are gone, not covered |
| No black shading over the text | the text is not removed |
| No white text colour | it can still be copied and searched |
| No black boxes from Adobe’s annotation tool | the text can still be found |
| Or delete the text in a copy and check it in Notepad | proves nothing is left |
| Redact scans with software made for it | a picture needs a different tool |
Source: LPLC, Write tech, wrong text, summarising the Federal Court of Australia’s guide.
The article adds a point that matters for anyone who scans paperwork, a NSW driver licence included. Where a document “is converted to a searchable PDF, the shaded text can still be searched”. A scanner with text recognition switched on rebuilds the very layer a black box fails to reach.
The guide was written for documents going to a court registry and then possibly to the public. A chatbot upload reaches a smaller audience with a longer memory, and the same mechanics apply to it without any change. Courts are also writing rules about what goes into generative AI at all, and the practice note from the NSW Supreme Court is where most practitioners meet them first.
What ChatGPT and Claude receive from an uploaded PDF
What happens inside a model is the provider’s business, and we will not guess at it. What the providers say publicly about how they take a PDF in is enough to answer the question that matters here, which is whether the covered words arrive.
ChatGPT: the digital text, and on Enterprise the page as well
OpenAI’s help article on file uploads is direct about it. “ChatGPT Enterprise supports Visual Retrieval for PDF files.” For everything else it says: “All other plans and document files only support text-based retrieval. This means that ChatGPT will extract digital text from the file and discard any images.”
Extracting the digital text is exactly what pypdf did to our letter. On those plans the drawn box is thrown away with the rest of the drawing, and the name is part of the text that is kept. Retention after the upload is its own subject, and the ChatGPT data guide tracks what OpenAI keeps and deletes.
Claude: a picture of each page, plus its words
Anthropic’s developer documentation on PDF support describes the process in two steps. “The system converts each page of the document into an image. The text from each page is extracted and provided alongside each page’s image.”
That documentation covers PDFs sent through Anthropic’s developer platform, and it is the most detailed public description the company gives. Read literally, the picture of the page would show black where the box is, while the extracted words sent next to it would still carry the name. Two inputs arrive, and the redaction covers just one.
We did not find an equivalent public description for Microsoft Copilot, so we make no claim about it here. The safe assumption for any tool is the one the test file supports: if the text is in the PDF, the tool can read it. What Microsoft does publish about data handling is collected in the Copilot guide.
Sending the passage instead of the file
Most questions about a document do not need the whole document. A summary of one clause, a reply to a letter or a check on a calculation needs that passage, not the letterhead, the signature block and the footer carrying the client number. Copying the relevant paragraphs into the prompt, with the identifiers in them replaced, leaves the rest of the file on your side, including the page you forgot was there.
It also takes the layers out of the question. Pasted text is exactly the text you can see, with nothing drawn on top and nothing hiding underneath, so what you checked is what goes. The cost is effort: for a forty page contract, a proper redaction of the file may be quicker than choosing passages by hand.
Acrobat Pro in six steps: how to redact a PDF
Adobe documents redaction as a feature of Acrobat Pro, and its Australian guide walks through six steps. The names below are the ones Adobe uses on that page. Menus move between versions, so if a label differs on your screen, search the tools list for “Redact”.
- Keep the original. Adobe’s own advice is to save a copy before you start, because redactions are irreversible once applied. Work on the copy.
- Open the redaction tool. From the Edit menu choose Redact a PDF, or right click and choose Redact.
- Set how the marks look, if you want to. Set properties in the Redact toolset changes how the marks look, for instance the fill colour or overlay text such as [Redacted], the wording the Federal Court prefers.
- Mark the content. Double click a word, or drag across a block of text or an image. Every name, number and address you want gone needs its own mark.
- Repeat across pages where needed. A right click option applies the same mark on other pages, which helps with letterheads and repeated reference numbers.
- Apply, then save. Choose Apply or Apply & Save. Until you do, the marks are only proposals and the text is still underneath.
The step that gets skipped
Step six is the easy one to miss. A marked document looks finished: the marks are visible, the colour is right, and it can be saved and sent. Only after Apply does Acrobat actually take the characters out. Word has the same trap with its markup, because a clean view of tracked changes still leaves them in the file.
The other common failure is using the Comment tools instead. A rectangle or a highlight from the commenting panel is an annotation, the “annotation tool” the Federal Court guide tells practitioners not to use. It sits on top of the page like a sticky note. For practices that redact every week, such as law firms preparing court bundles, having a second person open the saved file is a cheap safeguard.
Search and redact for repeated details
A client’s name can appear in a letterhead, a salutation, a table and a footer. Acrobat can search the document for a word or phrase and mark every match for redaction, which is safer than hunting by eye. It will still miss a name split across two lines or written differently, “T. Achterberg” or “Tamsin”, so search for each form.
A redacted PDF can still describe itself in its properties, the author for instance, and that is a separate job from the one on this page.
If the PDF started life as a Word file, that file has its own hiding places, and redacting the PDF does nothing to them.
Preview on a Mac, and the tools that only draw
Mac users have a built in option. Apple’s Australian guide to annotating a PDF in Preview describes a Redaction Selection tool: “Select text to permanently remove it from view. You can change the redaction as you edit but once you close the document, the redaction becomes permanent.”
Two details in that sentence matter. The redaction is editable while the file is open, and it becomes permanent on closing. So the test belongs after you close the document and open it again, not while you are still editing. “Remove it from view” is Apple’s phrase, and the only way to know what happened to the characters is the copy test below.
The tools that look like redaction
Most PDF software can draw a black rectangle, and most offer a black highlighter. Neither is redaction. The table sorts the common options by what they do to the text, not by what they are called.
| what you used | what happens to the words | safe for an upload |
|---|---|---|
| Acrobat Pro, Redact a PDF, then Apply | removed | yes, after testing |
| Preview, Redaction Selection, file closed | Apple: permanently removed from view | test it |
| Rectangle or shape from the comment tools | covered, still in the file | no |
| Black highlighter | covered, still in the file | no |
| White text or white box | invisible, still in the file | no |
| Online redaction service | the full file is uploaded first, and what comes back varies | check what it keeps and who can see it, then test |
Sources: Adobe Acrobat Australia; Apple Support Australia; LPLC summary of the Federal Court guide.
Any online redaction service receives the whole file before it covers anything, so three things matter: what is uploaded, whether it is kept, and who can see it. The tool above this guide covers the document and gives you the result, and the document is discarded as soon as you get it. For a privileged client document or a file carrying a tax file number, the Nonimo app does the covering locally, and the file stays on your computer from start to finish.
Three tests for a redacted file before the upload
No redaction is finished until the saved file has been tested, and the tests take less than a minute. Do them on the file you are about to upload, not the one you worked on, and after closing and reopening it.
| test | how | a failure looks like |
|---|---|---|
| Copy and paste | select across the black area, paste into Notepad or TextEdit | the covered words appear |
| Search | search the PDF for a surname or a number you removed | the viewer finds a match |
| Select all | select all text on the page and paste it | the full letter comes back |
The first test is the one the Federal Court guide describes, as summarised by the LPLC.
Plain text editors matter for the first test. Pasting into Word or an email can carry formatting across and make a result harder to read, while Notepad shows only characters. If the black area pastes as nothing, or as the word [Redacted], the text is gone. A workbook needs a different test, because Excel never lists a very hidden sheet, and checking a spreadsheet from outside Excel is the one that finds it.
What the tests cannot see
The tests check the text layer, and that is where uploads look first. They do not check a picture inside the PDF, so a photographed driver’s licence or a signature image is outside their reach. They also do not tell you whether a sentence that names nobody still identifies the client, which is a judgement no tool makes for you.
For health and disability documents that second limit is the one that usually bites. A diagnosis, a town and a month can point at one person with every name removed, as our guide to sensitive information explains with the categories the Act treats differently.
Scanned PDFs: no text layer, and nothing to search
A scan is a photograph of paper saved inside a PDF, like a scanned drivers licence. There are no characters in it, only pixels, so a text extractor finds nothing. Our scanned version of the test letter returned zero characters of text, and the three tests above all come back empty, which looks like success and proves nothing.
The information is still there, visible to anyone and anything that looks at the picture. Whether an AI tool reads that picture depends on how the provider takes the file in.
When a picture is read as a page
Anthropic’s documentation, quoted above, turns every PDF page into an image and gives the model both. On that route the scanned letter is read from the picture, and every detail on the paper is in front of the model. OpenAI, for its part, limits visual reading of PDFs to ChatGPT Enterprise, with other plans keeping the digital text and dropping images.
So the same scan can arrive in full in one tool and arrive empty in another. Clinics meet this every day with scanned referrals and pathology reports, and the Ahpra rules for clinicians using AI apply to those same files. For a document you cannot afford to leak, the only planning assumption that holds across tools is that the picture will be read. The same goes for a phone photo of the page, and what each tool keeps of an uploaded photo is covered separately.
Redacting the picture itself
For a scan, or a photographed NSW licence, the redaction has to change the pixels. In Acrobat a redaction mark can be dragged over an area of an image, and applying it removes what sits under the mark. The Federal Court’s advice, through the LPLC, is to redact a copy with software built for that job. Then look at the result at full zoom, because the tests that work on text tell you nothing here.
Be careful with text recognition. Many scanners and PDF programs can make a scan searchable, which adds an invisible layer of recognised text over the picture. That turns a scan back into the problem this guide started with: a black box painted on the image can leave the recognised words untouched underneath, which is the point the LPLC makes about searchable PDFs.
One missed tax file number is still a disclosure
The OAIC settled what an upload is in its October 2024 guidance on off the shelf AI products, revised in January 2025. Its worked example has insurance staff typing a claim, health details included, into a public chatbot, and the regulator’s verdict is that the company “is disclosing the information to the owners of the chatbot”.
That disclosure has to satisfy APP 6. Its best practice advice goes further, and says organisations are better off entering no personal information at all, “and particularly sensitive information”, in generative AI tools open to the public. A name left under a black box is personal information that you entered, even though you believed you had not.
Why a missed TFN costs more than a missed name
Tax file numbers sit under their own rules, and a business that is otherwise exempt from the Privacy Act can still be caught by them. The Privacy Act and AI tools guide explains who counts as a file number recipient, and why holding TFNs brings an otherwise exempt practice under the notification rules.
Whether a particular slip has to be reported is its own assessment with a clock attached, and our breach guide walks through that test.
If the file has already gone
Stop, then establish what went. Open the copy you actually uploaded, not the original on your desk, and run the three tests on it: the result tells you which details the tool received. Deleting the conversation tidies your own history, but what the provider holds after a deletion depends on its terms, and the guide to Anthropic’s data rules shows how much those terms vary by plan.
If the details are personal information, the notifiable data breach scheme may be in play. Section 26WH of the Privacy Act expects an entity that suspects an eligible data breach to assess it promptly and reasonably, and to do what it reasonably can to complete that within 30 days. Note what went, when and to which tool while you still remember, because that note is where the assessment starts.
The same client letter, with labels instead of boxes
Nonimo works from the other end. Rather than editing layers, you hand it the document and receive two things: the text with identifiers turned into labels, ready to paste into a prompt, plus a covered version of the file. When the input is a PDF, what you get back is a clean PDF of the text, with the original’s metadata stripped out.
This is the text of the test letter, as Nonimo returned it:
BEFORE AFTER (Nonimo)
Wattle Creek Tax Advisers Wattle Creek Tax Advisers
Client: Tamsin Achterberg Client: [PERSON_1]
Date of birth: 14/07/1986 Date of birth: [BIRTH_DATE_1]
Address: 22 Banksia Rise, Wirrimbah VIC 3999 Address: [RECORD_FIELD_1]
Tax file number: 000 000 000 Tax file number: [TFN_1]
Mobile: 0491 570 156 Mobile: [PHONE_1]
Refund of $2,418 Refund of $2,418
The labels go to the model, and the Nonimo app puts the real values back into the reply on your side. Because the link back to the client survives on your side, this is pseudonymisation; the comparison of the two terms explains why that matters in Australia. The firm name and the refund stayed, because the model needs them to answer.
The Nonimo app reads a fully scanned PDF on the computer, on Mac and Windows, and gives you its text, covered, ready for a prompt. A scan you send as a file still needs the picture treatment above. What the app stores on your computer, and the daily usage count it sends us, are described on the security page.
Sources
- OAIC, Guidance on privacy and the use of commercially available AI products, published 21 October 2024, updated 17 January 2025. The insurer worked example in which entering a claim into a public chatbot is a disclosure to its owners; APP 6; the best practice recommendation not to enter personal information into publicly available generative AI tools. oaic.gov.au
- Privacy Act 1988 (Cth), section 26WH. The assessment of a suspected eligible data breach and the 30 days to complete it. legislation.gov.au
- Legal Practitioners’ Liability Committee, Write tech, wrong text, last updated 11 June 2026. Failed redaction in electronic court documents read by third parties; black shading and white text; searchable PDFs; the summary of the Federal Court’s key points. lplc.com.au
- Federal Court of Australia, Guide to redacting documents in electronic form. The source of the points the LPLC summarises. fedcourt.gov.au
- OpenAI Help Center, File Uploads FAQ. Visual Retrieval for PDFs on ChatGPT Enterprise; text based retrieval on all other plans, extracting digital text and discarding images. help.openai.com
- Anthropic, PDF support, Claude developer documentation. Each page converted into an image, with the text of each page extracted and provided alongside it. platform.claude.com
- Adobe Acrobat, How to redact a PDF. Searching the document for words or phrases and marking every match for redaction. adobe.com
- Adobe Acrobat Australia, How to redact PDF. The six steps, Edit menu and Redact a PDF, Set properties, Apply and Apply & Save, and the advice to keep a copy because redactions are irreversible. adobe.com/au
- Apple Support Australia, Annotate a PDF in Preview on Mac. The Redaction Selection tool and when the redaction becomes permanent. support.apple.com/en-au
- Nonimo, test letter and extraction, 28 September 2026. An invented letter built with ReportLab in three versions (drawn boxes, applied redaction, scanned image) and read with pypdf 6.10.2; the Nonimo output shown above. Figures and code blocks on this page come from that run.
Nonimo is the software that does this on your own computer: it masks client names and IDs before your text reaches ChatGPT. No account needed, and the app does it without your files leaving your machine.
Common questions
Which redaction method actually strips the text out of a PDF?
The short answer to how to redact a PDF: pick the redaction function of your PDF software, never its drawing or highlighting tools. In Acrobat Pro you mark each detail and confirm with Apply, which Adobe describes as irreversible. Reopen the saved file and paste the black area into Notepad to prove nothing came across. Nonimo works from the other end, handing back labelled text.
Can ChatGPT see words hidden behind a black rectangle?
Yes, when the rectangle was merely painted over them. OpenAI states that on plans other than Enterprise, ChatGPT pulls the digital text out of an uploaded file, and a painted shape carries no text of its own, so it screens nothing from that step. Delete the characters first, or give the model Nonimo's labelled version in place of the original.
Does a black highlight or a filled shape remove the words?
No. A black highlight or a filled shape changes what you see on the screen and on paper, while the words stay in the file. The Federal Court's guide, as summarised by the LPLC, warns against black shading and white text for exactly that reason. Real redaction removes the characters, and a copy and paste test shows which one you did.
Can you redact in the free Adobe Acrobat Reader?
Adobe documents redaction as an Acrobat Pro feature, under the Redact a PDF tool, so the free Reader is not where it lives. On a Mac, Preview has its own redaction selection. Whatever you use, open the saved file afterwards and try to copy the covered area, because the tool's name tells you nothing about what it did.
What is the quickest way to check a redacted PDF?
Three checks, all on the file you saved. Paste the black area into a plain text editor, search for a surname you removed, and select everything on the page. If the name surfaces anywhere, the redaction failed. On a scanned page all three come back empty whatever is visible, because a picture holds no characters, so look at a scan instead of searching it.
Can an AI read a scanned PDF?
That varies by provider. Anthropic's developer documentation says every PDF page becomes an image that the model reads together with any text, so a scan is understood from the picture. OpenAI reserves visual reading of PDFs for ChatGPT Enterprise, and other plans keep only digital text. The Nonimo app reads a fully scanned PDF on the computer, on Mac and Windows, and returns its text with the details covered.
What does the Federal Court say about redacting electronic documents?
Its guide to redacting documents in electronic form, summarised by Victoria's Legal Practitioners' Liability Committee, says to replace the text with [Redacted], not to use black shading or white text, and not to draw black boxes with Adobe's annotation tool. For scans it says to redact a copy with software built for that purpose.
Is a leaked name under a black box a privacy incident in Australia?
It can be. In the OAIC's view, personal information typed into a public chatbot has been handed to the company that runs it, and APP 6 then decides whether that was allowed. A name you thought was hidden has travelled all the same. Whether it must be reported is a separate test. Nonimo replaces recognised identifiers with labels before the chatbot sees the text.