· Updated · Written and maintained by Nonimo
A black rectangle on a client’s name hides it from your eyes and from nobody else. The words sit one layer down, still inside the file, and on most plans the first thing ChatGPT does with an uploaded PDF is pull that layer out. Learning how to redact a PDF for an AI comes down to two moves: delete the text with a tool built for it, then prove on the saved copy that it is gone.
This guide does both with a test file you can rebuild yourself, and it stays on one question: what reaches the AI. What a redaction does to privilege or to a confidentiality rule is a legal question, answered in what redaction protects under US law.
How to redact a PDF: why a black box is not redaction
A PDF page is a list of drawing instructions. One writes a line of text in a given font at a given spot; another fills a rectangle with black. Pick a shape tool, a black highlighter or a marker in a free viewer, and the program adds one more instruction to the list. The one that wrote the name stays, and it still runs every time the page is drawn. The rectangle is simply painted afterward, on top.
That is why the covered text survives in three places at once. It is still selectable, so a copy and paste brings it back. It is still searchable, so the search box finds it behind the box. And it is still in the text layer that programs read without looking at the page at all, which is how assistants and extraction tools see a document.
Real redaction works on the list, not on the picture. A redaction tool finds the instructions that draw the covered words and deletes them, then paints a box where they were so the reader knows something was removed. After it runs, the black box is a record of a deletion rather than a lid over the text. In a spreadsheet, a black fill is always the lid, and the value sits unchanged beneath it.
So the practical answer comes in three parts, and the sections below go through them one at a time. Use a tool that deletes, apply it, and save a new copy. Search every covered value again in that copy, because a redaction pass removes what you marked and nothing else. Then look at what an extractor pulls from the saved file, since that is closer to what the AI gets than anything you see on the screen.
What an AI actually reads when you upload a PDF
You do not have to guess at this, because the two largest providers describe it themselves. OpenAI’s file uploads FAQ, updated this month, says that ChatGPT Enterprise supports visual retrieval for PDFs, and that all other plans and document files only support text based retrieval, which it spells out as extracting the digital text from the file and discarding any images.
Read that against the black box. On a Free, Plus or Business account, what the model works from is the digital text. The rectangle is not text, so nothing about it survives the trip, and the words it was covering are exactly the kind of content that does. What that account then keeps, and how long, is in what OpenAI says about your data.
| Tool, as its maker describes it | What it takes from a PDF | What a drawn black box does |
|---|---|---|
| ChatGPT, plans other than Enterprise | The digital text; images discarded | Nothing: the text under it is read |
| ChatGPT Enterprise, visual retrieval | The text plus the images on the page | Hides the words in the picture, not in the text |
| Claude, PDF support in Anthropic’s API | Each page as an image, with its extracted text alongside | Hides the words in the picture, not in the text |
Sources: OpenAI File Uploads FAQ and Visual Retrieval with PDFs FAQ; Anthropic, PDF support. Read September 28, 2026.
Seeing the box is not the same as not reading the words
Anthropic’s documentation is the more telling of the two. In its description, every page is turned into a picture, and the text pulled from that page travels with it. A model given both can see that a box is there, and it is also handed the sentence sitting under the box. Nothing in either provider’s help pages says the model skips text that happens to sit under a shape.
The same logic reaches past the two chat apps. Search tools, summarizers, document loaders and most programs that ingest PDFs start by extracting the text layer, because that is the cheap and reliable part. What happens to an upload after Anthropic receives it is set out in what Anthropic says about training on your data.
One memo, five files: what came out of each
To test this rather than argue it, we built a one page claim review memo with invented details: a claimant called Delia Brannock, a birth date, the SSN 000-47-2215 in an area Social Security never issues, an address at 1600 Example Road in Dayton, a phone number from the 555 block reserved for fiction, and a policy number. Two lines of narrative follow, and the second mentions Ms. Brannock again by surname.
Then we made five versions of the same page and pulled the text out of each with two free Python libraries, pypdf and PyMuPDF. The two libraries agreed on every file. The script, the five PDFs and the full output sit in the folder of working files behind this guide.
| Version of the memo | What extraction returned | Covered values that came out |
|---|---|---|
| The original, as typed | Everything | Not covered |
| Black boxes drawn into the page | Everything, word for word | 4 of 4 |
| Black boxes added as annotations | Everything, word for word | 4 of 4 |
| Real redaction applied to the same four fields | Labels left, values gone | 0 of 4, surname still in the narrative |
| A scan of the page, image only | Nothing at all | No text layer to read |
Test run September 28, 2026, on 5 versions of one page, with pypdf 6.10.2 and PyMuPDF 1.26.5.
The second row is the one to remember. That version is drawn the way a flattened box ends up, with the rectangle written into the page itself rather than sitting on it as a note, and it still gave up the name, the birth date, the SSN and the address. Flattening changes what the rectangle is. It does not touch the text underneath.
The fourth row is the quieter lesson. The redaction worked exactly as instructed: the four values we marked were deleted and extraction returned empty labels. But the narrative line still said Ms. Brannock, because nobody had marked it. A real tool removes what you select, which makes the search that comes after it the step that decides whether the file is clean.
How to black out text in a PDF so it is really gone
The phrase people type is how to black out text in a PDF, and the tools answer it in two very different ways. Some black it out by covering it, which is the failure described above. Others delete it and leave the black mark as a receipt. The label on the button does not always tell you which one you have, so the test in the next section is not optional.
| Tool | Where the real one is | What to avoid in the same program |
|---|---|---|
| Adobe Acrobat Pro | All tools, Redact a PDF, Redact text and images | Comment and drawing tools, a black highlight |
| Preview on a Mac | Redaction Selection in the markup tools | A filled rectangle from Shapes |
| Online redaction sites | Their redact function, if it deletes | Uploading before you know what the site keeps and who can see it |
Sources: Adobe Acrobat Pro help on redaction; Apple, Preview User Guide, Annotate a PDF. Read September 28, 2026.
Acrobat Pro: mark, then apply, then save a copy
Adobe’s help puts the tools under All tools and then Redact a PDF. Choose Redact text and images, drag across the words or the image you want gone, and select Apply. Acrobat then asks whether to sanitize and remove hidden information as well, and if you do, it asks for a name and a place for the sanitized copy. The redaction tools are in Acrobat Pro; the free Reader cannot do this.
Two details matter more than the menus. Marking is not redacting: until you apply, the marks are only a plan, and a file saved at that point still carries every word. And the sanitize step deals with the information a file carries about itself, which another guide in this series takes on.
Search the whole file, not just the page you are on
Acrobat Pro also has Find text and redact, which searches for words, phrases or patterns and lists every hit. Tick the ones you want, choose Mark checked results for redaction, and apply as before. This is the step that would have caught the surname in our memo. It only works on text that is text, so it will not find words in a scan.
Preview on a Mac: the redaction tool, not a rectangle
Preview has a Redaction Selection tool among its markup tools. Apple describes it as taking the selected text out of view; you can still adjust it during the session, and closing the document fixes it for good. Work on a duplicate, because after that there is no way back.
Apple’s same page lists the Shapes tool, and a black filled rectangle from there is an annotation, which Apple says stays editable when you save normally. It also warns that table of contents entries linking to redacted text are not removed. If the file has an outline, open it and check what the entries say.
Online redaction sites, and what happens to your file
Online redaction sites top the results for this search, and some may delete text correctly. To clean a client’s file on someone else’s server, you first send the uncleaned file there, so before the first upload, read the site’s terms for three answers: what gets uploaded, whether it is kept, and who can see it. The tool above this guide answers the second one plainly: it masks the PDF, gives you the result to copy or download, and discards the document as soon as you have it. If the document must not leave your computer at all, use a program that runs on your own machine, like the two programs above or the Nonimo app. The same reasoning is why a written AI policy usually names the tools staff may use.
How to check if a PDF is properly redacted before it leaves your desk
Test the file you saved. The version still open in your editor can show you something different. The four checks below take a few minutes and catch the two failures that matter: a box that covers without deleting, and a value you never marked.
- Search for every value. Use the viewer’s search box for each name, number and address you meant to remove. Search surnames on their own and first names on their own, since documents rarely repeat a full name.
- Copy under each box. Drag across a black mark, copy, and paste into a blank document. Anything that appears was never deleted.
- Read the text layer. Open the file with a free extractor, or save it as text, and read the result top to bottom. This is the view closest to what an assistant receives.
- Check the outline and attachments. Bookmarks, form fields and attached files have their own text, and a pass over the page does not reach them.
Why the search step is the one that pays
The copy test tells you whether the tool worked. The search tells you whether you did. In our memo the claimant line was cleaned perfectly and the narrative line underneath still named her.
Letters, emails exported to PDF and case notes are worse than a form, because the name turns up in the greeting, in the body and in the signature block. When the PDF began as a Word draft, clear the draft’s tracked changes and comments first, so the export starts from one settled version of the text.
A search also finds values in the places you do not look at, such as the running header on page 12. If the file started as a Word document, the easiest fix is often to clean the Word file and export it again, which is its own set of steps. Either way the check is the same: search the copy you will actually send, and whatever file is the one that goes out is the one you test.
Scanned PDFs are pictures, and pictures are a different problem
A scan has no text layer. Each page is a photograph of paper, and our fifth test file, an image only version of the same memo, returned no text at all from either library. That sounds like good news and is not quite, because whether an AI reads it depends on the tool and not on you.
On the OpenAI side, the file uploads FAQ says that plans other than Enterprise extract digital text and discard images, which leaves a pure scan with little to work from. Anthropic says the opposite path exists in its API: every page also goes in as an image, and a model that can read images can read what the scan shows. The safe rule is simple. On a scan, what you can still see is what the AI can see.
How long each provider keeps the picture afterward is set out in our guide to what AI assistants keep from your photos.
Black marker on paper works, if the marker is on the paper
The US Court of Federal Claims gives the plain method in its note on PDF redaction: for a photocopy or a scan, print it, black out the text with a marker and scan the paper back in. The marker becomes part of the picture, and there is nothing underneath it to recover.
What does not work is the digital version of the same idea, a black shape drawn over a scanned page, or over a photo of a driver’s license, in a viewer. The shape sits on top of an untouched photograph. Anything that reads the image without the shape, or with it moved, gets the original.
Mixed files, and the OCR layer you did not know was there
Many PDFs are both. A scanner that runs OCR stores the picture plus an invisible text layer so you can search, which means a scanned page can carry text after all. The tests from the previous section apply unchanged: search it, copy under the boxes, read the extracted text. If a search finds words on what looks like a photograph, there is a text layer, and it needs the same treatment as any other.
A federal court, the NSA and a defense filing in Washington
None of this is new, and American institutions have been warning about it for two decades. In December 2005 the National Security Agency published a paper called Redacting with Confidence. Its first listed mistake is covering text with black rectangles or highlighting it in black, which it says works for printed paper and does not work in a digital file.
The US Court of Federal Claims says the same in a short note for filers: the highlighter function in Adobe puts a black box over the data but merely hides it, and anyone can copy the box into a word processing document and see what was under it. Federal courts care because Rule 5.2 of the Federal Rules of Civil Procedure limits what a filing may show of an SSN, a birth date and an account number.
The filing that proved the point in public
In January 2019 Paul Manafort’s lawyers filed a response to allegations from the special counsel’s office. The sealed passages sat under black bars. As the ABA Journal reported on January 10, 2019, highlighting the bars let anyone copy the text into a new document, and that is how an allegation about sharing campaign polling data became public.
Nothing about that failure involved an AI. What has changed is the reader. A reporter had to think of highlighting the bars; an assistant fed the PDF gets the text by default and does not need to think of anything. If a file is headed for an AI tool from a firm, the questions a court asks about what left the building are covered in the guide on disclosing AI use in court filings.
For a practice, the same file raises a different question. A PDF with a patient’s details that reaches a chatbot plan without a business associate agreement is a HIPAA problem whatever the black boxes looked like, and the HIPAA compliant AI guide sets out which plans OpenAI will cover with one.
How to redact a PDF when all the AI needs is the text
Most of the time the assistant does not need the file. It needs the words, so it can summarize the letter, draft a reply or pull dates out of a report. That opens a simpler route than editing the PDF, and it avoids the layers, the outline and the attachments entirely. A slide deck has the same shortcut: copying from Outline view leaves the notes and comments behind.
- Take the text out. Select all and copy, or save the PDF as text, so you have plain words with no layers behind them.
- Clean it where you can see everything. In plain text, nothing is hidden: search each name, number and address and swap it for a label such as Claimant or Policy.
- Read it once, top to bottom. Look for the second mention, the signature line and the reference number in the footer.
- Paste the text, not the file. The assistant gets what you read, and nothing you did not.
This also sidesteps the information a file carries about itself, such as its author and history, because none of that travels with pasted text. Which details count as identifying in the first place is a longer list than names and SSNs, and it deserves its own checklist.
The difference between removing names and making a text truly anonymous is set out in deidentified vs anonymized.
Nonimo and the PDF: text out, a covered copy back
The Nonimo app runs on your computer, and the PDF you drop into it stays there. You get back the text with personal details replaced by placeholders, ready for the AI, plus a covered copy: a fresh PDF carrying the text alone, without the original’s layout and without its metadata. The tool on this page is the quick route for a one-off file: it masks the document, gives you the result, and does not keep it. Here is our test memo through the app.
Extracted from the PDF What goes to the AI
Claimant: Delia Brannock Claimant: [PERSON_1]
Date of birth: 03/14/1979 Date of birth: [RECORD_FIELD_1]
SSN: 000-47-2215 SSN: [REFERENCE_1]
Address: 1600 Example Road, Dayton, OH 45419 Address: [ADDRESS_1]
Phone: (937) 555-0148 Phone: [RECORD_FIELD_2]
Policy number: HX-2291-0457 Policy number: [REFERENCE_2]
Ms. Brannock reports a collision on I-75 Ms. [PERSON_2] reports a collision on I-75
on 06/02/2026. on 06/02/2026.
The surname in the narrative, the one a manual pass missed, was replaced as well. Paste the AI’s answer into the app and it restores the real details, on your machine. The collision date and the highway stay readable, which is what the AI needs to help.
A PDF that is a scan from start to finish is read right on the computer, on Mac and Windows, up to 50 pages, and comes back as masked text. The security page lists what the app stores on your computer and what leaves it, and the page for law firms shows how a firm sets it up.
Two pieces of advice about PDFs that do not hold up
The first is that flattening a PDF, or printing it to PDF, makes a black box safe. Flattening turns the box from a note into part of the page, and our second test file shows that a box written into the page still leaves every covered value in the text. The words were written into the page before the box ever was, and flattening leaves them where they were.
The second is that an AI looks at the page the way you do, so it cannot see what you covered. OpenAI’s help says most ChatGPT plans read the digital text and discard the images, and Anthropic says Claude gets the page image and its text together. Both hand the covered words to the model. If claim files are your daily PDFs, our page for insurance agencies shows a claim note before and after.
Sources
Checked September 28, 2026.
- OpenAI, File Uploads FAQ. ChatGPT Enterprise supports visual retrieval for PDFs; all other plans and document files use text based retrieval, extracting the digital text and discarding images.
- OpenAI, Visual Retrieval with PDFs FAQ. On Enterprise, ChatGPT reads both the text and the embedded images or diagrams of a PDF.
- Anthropic, PDF support. Each page is converted into an image, and the text extracted from each page is provided alongside it.
- Adobe, Redact sensitive content in PDFs in Acrobat Pro. Redact a PDF, Redact text and images, Apply, and the option to sanitize and remove hidden information.
- Adobe, Search and redact text in PDFs with Acrobat Pro. Find text and redact, patterns, and Mark checked results for redaction.
- Apple, Preview User Guide: Annotate a PDF. Redaction Selection, permanent once the document is closed; table of contents entries linking to redacted text are not removed; annotations saved normally stay editable.
- US Court of Federal Claims, PDF File Redaction Best Practices. The Adobe highlighter hides data without removing it; copy and paste reveals it; for scans, print, black out with a marker and scan again.
- National Security Agency, Redacting with Confidence, December 13, 2005. Copy hosted by the US Court of Appeals for the Seventh Circuit. Covering text with black rectangles or highlighting it in black is the first of three common mistakes.
- Federal Rules of Civil Procedure, Rule 5.2. What a court filing may show of a Social Security number, a birth date, a minor’s name and a financial account number.
- ABA Journal, January 10, 2019. The Manafort defense filing, whose blacked out text could be copied and pasted into a new document.
Nonimo is the software that does this on your own computer: it masks client names and IDs before your text reaches ChatGPT. No account needed, and the app does it without your files leaving your machine.
Common questions
What is the safe way to redact a PDF before uploading it to ChatGPT?
Knowing how to redact a PDF starts with a real redaction tool, not a drawing tool: Redact text and images in Acrobat Pro, then Apply, or Redaction Selection in Preview on a Mac. Then test the saved copy by searching for each name and copying the covered lines into a blank document. OpenAI says most ChatGPT plans read the digital text of a PDF and discard its images.
Does a black box over text in a PDF hide it from an AI?
Not when the box sits on top of the words. A shape or a black highlight hides the words on screen and leaves them in the text layer, and that layer is what text extraction returns. In our test file, all four covered values came out of both the drawn box and the box added as an annotation, word for word. Only a redaction tool that removes the text itself changes that result.
Does flattening or printing to PDF remove the text under a black box?
Flattening makes the box part of the page, but the text instructions are still there underneath it. Our second test file draws the box straight into the page content, which is roughly what a flattened box becomes, and the name, birth date, SSN and address all came out when the text was extracted. Printing to a PDF keeps the text as text, so it does not change this either.
How can I tell whether a PDF was redacted properly?
Three tests, all on the copy you will send. Search for every name and number you removed, including surnames on their own. Select the area under each black box, copy, and paste into a blank document. Then extract the text of the whole file with any free tool and read it. If any of the three shows a covered value, the redaction did not happen.
Can I redact a PDF for free on a Mac?
Yes. Preview has a Redaction Selection tool, and according to Apple the change becomes permanent at the moment the document is closed. Use a duplicate, because there is no undo after that. Do not use a filled rectangle from the Shapes tool instead: Apple describes those as annotations that stay editable. Apple also notes that table of contents entries linking to redacted text are not removed.
Does ChatGPT or Claude read a scanned PDF?
It depends on the tool and the plan. A scan has no text layer, and OpenAI says most ChatGPT plans extract digital text and discard images, which leaves a scan with little to read. Anthropic's documentation for Claude says every PDF page also goes in as an image, so anything visible on the scan can be read. On a scan, whatever you can still see is what the AI can see.
What went wrong with the Manafort filing in 2019?
In January 2019, Paul Manafort's lawyers filed a response to the special counsel's allegations with black bars over the sealed passages. As the ABA Journal reported on January 10, 2019, highlighting the bars let anyone copy the text and paste it into a new document. The bars covered the words on screen, and the words were still in the file.
What comes back from Nonimo when you drop in a PDF?
It cleans the text. You drop the PDF into the Nonimo app on your computer, where the file stays, and get the text back with personal details replaced by placeholders, ready for the AI, and you can download a covered copy. That copy is a clean new PDF with the text and none of the original's metadata. The Nonimo app reads scanned PDFs too, on Mac and Windows, on the computer itself, up to 50 pages. The tool on this page masks the PDF, gives you the result to copy or download, and discards the document as soon as you have it.