Remove metadata from PDF, Word and photos before AI reads them
· Written and maintained by Nonimo
A Word document or a PDF keeps a second record that never appears on the page: its author, its most recent editor, which company owns the software, when it was printed, its template and, for any photo inside, where that photo was taken. To remove metadata from PDF, Word and photo files before any assistant opens them, you clean that record on a copy, then open the copy and check that each field is empty.
This guide covers that record and nothing else. Why a black rectangle hides nothing in a PDF is explained in the PDF guide on black boxes that fail, and comments and tracked changes inside a Word file belong to a separate guide on redacting in Word. What happens to an upload once it reaches OpenAI is in OpenAI’s own account of how it uses data.
What a PDF or Word file says about you without showing it
Metadata is information stored alongside the content, in sections of the file that no page view displays. In Microsoft’s account, some properties are filled in by people, like a title or a subject line, and others are kept current by Office itself, like the name of whoever saved the document most recently and the day it was first created. You never type most of it. It accumulates as the file moves through an office.
Some of it is harmless. Much of it names people. An author field carries the name from the Office account of whoever created the file, which may be a colleague who left two years ago. The company field often holds the firm’s legal name. A template named for a client, or a title that someone once filled in with the client’s surname, turns a neutral memo into one that identifies the matter before anyone reads a line.
| Field | Where it lives | What it can reveal |
|---|---|---|
| Author, last saved by | Word core properties; PDF Info and XMP | Names of the people who worked on it |
| Company, manager | Word extended properties | The firm, and who the author reports to |
| Title, subject, keywords | Both formats | Often the client or matter name |
| Created, modified, printed | Both formats | When the work happened, sometimes a time zone |
| Template, hyperlink base | Word extended properties | Template names and network folder paths |
| Custom properties | Word; custom tab in Acrobat | Anything a document system wrote there |
| Photo EXIF | Pictures inside either format | Camera, time and GPS position |
Sources: Microsoft Support on the Document Inspector; our test files, September 28, 2026.
Why nobody notices it
Metadata is invisible in exactly the view where people check their work. You read the page, fix the page and send the page, and the property sheet rides along untouched. The American Bar Association made the same observation in 2006: many programs automatically embed the name of the computer’s owner, the date and time of creation and the identity of whoever saved it last, and that information might simply be a right click away.
For HR files there is an extra wrinkle: how PHI and PII differ in an employee file affects how sensitive a staff name in a property field turns out to be.
Where to look before you remove anything
Start by reading what your file says. Cleaning blind means you cannot tell afterward whether anything changed, and each program shows a different slice of the record.
In Word for Microsoft 365, the Info page under File shows a short list of properties down its right side. That list is partial. Microsoft’s instructions go two steps further: the link that expands it to every property, and then the Advanced Properties dialog, whose Summary tab holds 8 fields, from Title and Author through Company to Comments. The Custom tab beside it keeps whatever document systems have added.
In Acrobat and in Preview
In Acrobat, Adobe’s help puts the fields under Menu, then Document properties, then the Description tab: title, author, subject and keywords. The button Additional Metadata opens the full XMP record, which is a second copy of much of the same information, stored in a different part of the PDF.
On a Mac, Preview’s Inspector, opened from the toolbar, shows general information about a PDF, which Apple says includes the author name, and its More info tab shows where a photo was taken. It is a quick look, not a full inventory: the XMP record and custom fields need Acrobat or a script. For a law firm reviewing files at volume, our law firm page explains how the text side is handled.
| Program | Where the properties are | What it shows best |
|---|---|---|
| Word for Microsoft 365 | File, Info, Show All Properties, Advanced Properties | Summary, statistics and custom fields |
| Acrobat | Menu, Document properties, Description, Additional Metadata | Info fields and the XMP record |
| Preview on a Mac | Inspector button in the toolbar | Author, keywords, photo location |
Sources: Microsoft Support on file properties; Adobe help on PDF metadata; Apple’s Preview User Guide. Read September 28, 2026.
One claim memo, two files: what a script read from each
To show what a program gets, we built two test files around an invented claim memo at an invented insurance agency, Castellan & Rusk. Every name in them is made up and belongs to nobody: an author called Harlan Quisenberry, a colleague called Priya Vandermolen who saved it last, a manager called Ray Tesdahl, a claimant called Corliss Vandegrift and a policy number, PL-7730-1186. The photo inside was stamped with the position of a public lawn in Columbus, Ohio.
Replacing names in the body with labels is a different job, with a root guide of its own; here the body stays neutral on purpose.
The body of the memo is two neutral sentences about a collision on I-70. Everything that identifies anyone sits outside the page, in the metadata. Then we read both files back with ordinary free Python libraries, python-docx for the Word file and pypdf and PyMuPDF for the PDF, which is close to what any software does when it opens a document to work on it.
The Word file gave up every field
The script read back all of it. The author and the colleague who saved it last came out by name. So did the company, the manager, the template called Castellan Rusk Claims Memo, and a hyperlink base pointing at a claims folder on the agency’s file server. The custom property named the claimant and her policy number, and a separate part of the file held the reviewer’s work email.
The statistics were just as specific: 14 revisions, an editing clock of 312 minutes, a creation date in June, a last save in September and a print date the day before. Not one of those details is visible when the memo is open in Word. Every one was available to the first program that opened the archive.
The PDF kept two copies of the same story
The PDF told its story twice. Its Info dictionary returned the author, the title with the claimant’s surname, the subject with the policy number, the keywords and the software that made it. Its XMP record, a separate block of XML inside the same file, returned the author again, the creating tool and a unique document ID. Clearing one of the two does not clear the other.
Two details surprised us. The creation date carried a time zone offset, which tells a reader roughly where the computer was. And PyMuPDF and pypdf agreed on every field, so this is not the quirk of one library.
| What we stored | Word file | |
|---|---|---|
| Author and last person to save | Both read back | Author read back, twice |
| Company, manager, template | All read back | Not stored |
| Title, subject, keywords | All read back | All read back |
| Claimant in a custom property | Read back | Not stored |
| Photo with camera and GPS | Read back from the picture | No photo in this file |
Test run September 28, 2026, with pypdf 6.10.2, PyMuPDF 1.26.5 and python-docx 1.2.0. Invented data.
Why an upload carries the metadata and a paste does not
When you paste text into a chat window, the property sheet stays behind on your computer. When you upload the file, it travels with the file, because the properties are part of the file. The question is whether anything on the other side looks at them, and OpenAI answers that directly.
Its File Uploads FAQ gives examples of tasks ChatGPT can do with an uploaded document, and under extraction it lists this one: “Extract metadata (author, creation date, etc.) from a document.” That does not mean every conversation reads the author field. It means the fields are within reach of the tool, and a user or a prompt that asks for them will get them. The same FAQ lists all common file extensions for documents among the supported types.
The model is not the only reader
OpenAI’s page on data analysis adds that, for some jobs, ChatGPT produces Python and executes it in a notebook environment. A few lines of that code are enough to open a .docx or a PDF and print its properties, which is exactly what our test script did. Any assistant that can run code on a file you uploaded can do the same, whatever the chat interface happens to show. The same code reads a workbook’s hidden sheets just as easily.
Then there is everything that happens before and after the chat. Files uploaded through a firm’s own tools, agents or integrations are parsed by software that may index properties too, since Microsoft itself presents them as a way to search for documents. How long an upload is kept depends on the product and plan; for Anthropic’s side, see Anthropic’s statements on model training. The only version of the property sheet you control is the one on your own disk.
Word’s property sheet: a broad pass, then a close read
Word gives you two tools for this, one broad and one precise. Use both, on a copy, and in that order. The broad tool empties whole categories; the precise one lets you see what is left and type over what the broad tool leaves in place. Settle the tracked changes before either one, so that both tools work on a file that has already been through accepting or rejecting every change and deleting the comments.
Comments, tracked changes and hidden text are a separate category of the same inspection and a separate job, with their own traps, covered in a guide of their own. Here the focus is the property sheet: author, company, dates, template and custom fields.
The broad pass: Inspect Document
On your copy, Inspect Document sits under Check for Issues on the Info page of the File tab. The category that matters for this job is the one Microsoft labels “Document Properties and Personal Information”. Microsoft lists 10 kinds of item for it.
The first is the content of 3 tabs of the properties dialog, Custom, Statistics and Summary. The rest are email headers, routing slips, send for review details, document server properties, document management policy information, content type information, databinding links, a username and the template’s name. Choose Remove All beside it.
Microsoft also warns that some information the inspector removes cannot be restored, which is the reason for the copy. It also notes that the OpenDocument format needs the inspector run every time you save in it. If your firm keeps files in a document management system, expect server properties there as well, because Microsoft lists them separately. LibreOffice has its own save setting, yet a Writer .odt still keeps comment text and deleted lines after it runs.
The precise pass: Advanced Properties
Then open File, Info, Properties, Advanced Properties and read every tab. Check the Summary fields one by one, and the Custom tab for anything a document system or another program wrote there. Look at the template name especially: a firm template that bears a client’s or a partner’s surname puts that surname in plain sight. If a field survived, clear it by hand or type a neutral value such as the matter type, then save.
A useful habit for a team is to agree on neutral defaults, so that new documents do not start life carrying a client or an employee’s name. It sits well inside an AI acceptable use policy alongside the rules on what may be uploaded at all.
Acrobat and Preview: steps to remove metadata from PDF copies
A PDF needs its own pass, because it stores properties in two places and because many PDFs arrive from outside, already stamped with somebody else’s name, as any CPA firm collecting client paperwork sees every tax season. If you want to remove metadata from PDF copies reliably, clean both the Info fields and the XMP record, then verify with a reader that shows both.
In Acrobat Pro, Adobe documents a sanitize step under All tools, then Redact a PDF: a toggle for sanitizing the file and removing hidden information, followed by a prompt to save the sanitized copy under a new name.
Adobe lists metadata among what sanitizing removes, together with attachments, hidden layers and data from previous saves. The US District Court for the Southern District of Illinois gives the older version of the same advice, with Sanitize Document removing all hidden information and metadata automatically.
Editing single fields instead
If you only want to change one field, Acrobat lets you edit it under Menu, Document properties, Description. That changes the Info fields you can see there. Open Additional Metadata afterward to check the XMP record, because our test file showed how easily the two copies disagree. For a single sensitive field, sanitizing is the surer route; editing is for correcting a title, not for cleaning a file.
What a court means by basic metadata
The District of Maine’s metadata FAQ, compiled from a 2012 memorandum of the Administrative Office of the US Courts, makes a point worth keeping in mind: PDF files still retain some basic file description metadata, such as the file name and date. A PDF is often safer than a Word file for revision history, but it is never empty by default. On a Mac, Preview’s Inspector will show you the author; it does not show you the whole record.
Remove metadata from PDF exports before they leave Word
Many PDFs start as Word files, and the export is where property sheets quietly cross from one format to the other. The step people rely on to “flatten” a document carries the names across with it. It is an easy way for a client’s name to end up in the properties of a PDF that nobody ever opened in Acrobat.
Microsoft’s instructions for saving a Word document as a PDF go through File, Save a Copy, then PDF, then More Options and Options. One of the choices there reads Document properties, and Microsoft’s wording is plain: if you want to include document properties in the PDF, make sure it is selected. Clear it before you export a file that is headed anywhere outside the office.
What crossed in our test
We do not have Word on the test machine, so we exported our Word file with LibreOffice instead and read the result.
Four fields crossed into the PDF: the title with the claimant’s surname, the author, the subject with the policy number and the keywords. The company and the template did not. And the photo inside the exported PDF no longer carried EXIF when we checked it. Another program may behave differently, which is why the check on the finished PDF is the step that counts.
Printing to PDF is sometimes offered as the fix. It does rebuild the file, and a federal court FAQ calls it one of the ways to remove revision metadata, but the same FAQ notes that PDFs still keep basic description data such as the file name and date. A deck needs one more look before export, since PowerPoint’s PDF options can bring comments and hidden slides along.
Whatever route you take, open the result and read its properties before you upload it. If cyber insurance is part of the conversation at your firm, this is the kind of control an underwriter asks about.
Photos inside the document carry their own GPS and camera
A picture pasted into a report is a file inside a file, with its own metadata. Phone cameras normally write EXIF into each photo, a driver’s license copy included: the camera make and model, the time, and, when location is on, the latitude and longitude. The document’s properties can be spotless while the photo on page three still says where it was taken and on what phone.
Our test photo kept all of that inside the .docx. The script found the invented camera make and model, the moment the shutter fired, the claimant’s name stored as the photographer, and GPS coordinates that resolve to a lawn in Columbus. Microsoft’s inventory of what the inspector finds in Word covers properties, comments, headers, hidden text and custom XML, among others, and it does not mention metadata stored inside pictures.
What the Army learned from four helicopters
In March 2012 the US Army published a warning to soldiers called Geotagging poses security risks. It cites a case from 2007: soldiers photographed a new fleet of helicopters on the flightline at a base in Iraq, the photos were uploaded to the internet, and from them the enemy determined the exact location of the helicopters inside the compound and destroyed four AH-64 Apaches in a mortar attack.
An insurance claim photo is not a flightline, but the mechanism is the same: coordinates nobody meant to share, in a place nobody looked. A photo of a client’s damaged kitchen places the client’s home. For health practices, a patient photo with location is worth checking against what HIPAA asks of an AI tool before it goes near a chatbot.
Take the location out first
The cleaner order is to remove location from the photo before you insert it. Apple documents a switch for this on iPhone and iPad: in the Photos app, select the photos, tap the share button, then Options, and turn off Location; the sheet then shows Location Not Included. None of the big four assistants says whether it strips location from an uploaded photo, as our guide to what they keep from your photos found.
Apple also explains how to keep the Camera from saving location in the first place: in Location Services, under Privacy & Security in Settings, set Camera to Never. For photos already in a document, replace them with copies that were cleaned first. Our insurance agency page covers the text of claim files.
What US lawyers and federal courts already say about metadata
None of this is a new worry, and American legal institutions have been writing about it for two decades. The clearest statement is ABA Formal Opinion 06-442, issued on August 5, 2006. It dealt mainly with the lawyer who receives metadata, and concluded that the Model Rules do not specifically prohibit reviewing and using it.
The same opinion points the burden back at the sender. A lawyer who is concerned about sending a document that contains metadata, it says, may be able to limit the likelihood of its transmission by scrubbing it or by sending a different version without the embedded information.
Its footnote cites the Model Rule 1.6 comment on acting competently to safeguard client information. State bars have written their own opinions since, and court orders on AI add another layer, which our page on AI disclosure in court filings follows.
Courts that publish their own instructions
Federal courts publish practical instructions because filings land on public dockets. The Southern District of Illinois has a short guide on removing metadata with the Document Inspector and with Acrobat’s sanitize tool. The District of Maine’s FAQ explains why editable formats are where accidental disclosure usually happens, and why PDFs still keep basic description data.
An AI upload is not a filing, and no court rule governs what you give a chatbot. But the same logic applies.
A property sheet that should not reach opposing counsel should not reach a third party service either, and whether the upload itself becomes a problem for the firm comes up in the guide on breach notice after an upload. For privilege questions, redacting before AI sets out what a clean copy does and does not protect.
Nonimo and a file’s properties
The Nonimo app is installed locally. Drag a file onto it and you get two results: a version of its text where placeholders stand in for personal details, for pasting into the AI. You also get a copy of the file in which those details are covered. Below is what Nonimo made of the text of an invented claim memo.
What you typed What goes to the AI
Claim review: collision on I-70 Claim review: collision on I-70
Claimant: Corliss Vandegrift Claimant: [PERSON_1]
Insured: John Doe Insured: [PERSON_2]
Policy number: PL-7730-1186 Policy number: [REFERENCE_1]
Email: cvandegrift@example.com Email: [EMAIL_1]
Phone: (614) 555-0182 Phone: [RECORD_FIELD_1]
When the answer arrives, the app puts the original names back into it, on your own machine. In the Word copy the app hands back, the creator, the most recent saver, the company and the manager come back blank; placeholders replace the title, subject, keywords, description, category and custom fields; the record of who edited it, emails included, is blanked, the thumbnail is dropped, and pictures inside are stripped of EXIF.
For a PDF, the copy is a clean new file with the text and none of the original’s metadata. The data the app saves on your computer is described on the security page, and the license page lifts the monthly word limit.
Sources
Checked September 28, 2026.
- OpenAI, File Uploads FAQ. Under extraction, the example “Extract metadata (author, creation date, etc.) from a document”; all common file extensions for documents supported.
- OpenAI, Data analysis with ChatGPT. For some data analysis tasks, ChatGPT writes and runs Python code in a stateful Jupyter notebook environment.
- Microsoft Support, Remove hidden data and personal information by inspecting documents, presentations, or workbooks. What the Document Properties and Personal Information inspector covers in Word, including username and template name; run it on a copy; the OpenDocument caveat.
- Microsoft Support, View or change the properties for an Office file. File, Info, Show All Properties, Advanced Properties; the Summary tab fields; standard, automatic and custom properties.
- Microsoft Support, Save or convert to PDF. In Word, More Options, Options, and the Document properties choice that includes properties in the PDF.
- Adobe, Sanitize PDFs in Acrobat. All tools, Redact a PDF, the sanitize toggle and the Save Sanitized Document prompt; metadata among what is removed.
- Adobe, Edit document metadata in PDFs. Menu, Document properties, Description, and Additional Metadata for the XMP record.
- Apple, View information about PDFs and images in Preview on Mac. The Inspector shows file size and the author name; More info shows where a photo was taken.
- Apple, Manage location metadata in Photos. Share, Options, turn off Location; Settings, Privacy & Security, Location Services, Camera.
- US Army, Geotagging poses security risks, March 2012. The 2007 case in Iraq: uploaded photos revealed the helicopters’ position and four AH-64 Apaches were destroyed.
- ABA Formal Opinion 06-442, August 5, 2006. Programs embed the owner’s name, creation time and last saver; a sender may limit transmission by scrubbing metadata or sending a different version.
- US District Court, Southern District of Illinois, Removing metadata and hidden information. Document Inspector steps for Word, and Acrobat’s Remove Hidden Information and Sanitize Document.
- US District Court, District of Maine, Metadata FAQ. From a 2012 Administrative Office memorandum: printing to PDF removes revision metadata, and PDFs still keep basic file description metadata such as file name and date.
- Nonimo test files, September 28, 2026. An invented .docx and PDF built with python-docx, reportlab and PyMuPDF, read with python-docx 1.2.0, pypdf 6.10.2 and PyMuPDF 1.26.5, and exported with LibreOffice.
Nonimo is the software that does this on your own computer: it masks client names and IDs before your text reaches ChatGPT. No account needed, and the app does it without your files leaving your machine.
Common questions
Which tool should I use to remove metadata from PDF files before an upload?
In Acrobat Pro, the dependable one is the sanitize option under All tools, then Redact a PDF, which saves a cleaned copy; Adobe lists metadata among what it takes out. Reopen that copy and read Document properties and Additional Metadata. OpenAI names pulling a document's author and creation date as one task ChatGPT can do on request, so expect those fields to be readable.
Which details does a .docx store about the people who made it?
More than the author. Our invented test file held its creator, the colleague who saved it most recently, the agency's name, a manager, the template, a folder on a file server, fourteen revisions, 312 minutes of editing, a print date and a reviewer's email. Microsoft notes that Office keeps some of this up to date on its own, such as who saved the file most recently and when it was created.
Can ChatGPT see who wrote a document you give it?
It can if you ask. OpenAI's help center counts getting a document's author and the date it was made among its extraction examples, and its data analysis page says ChatGPT can write Python and execute it for some jobs, which is exactly how a program reads a file's property sheet. Treat every field as readable, and clean the copy before it leaves your computer.
Does exporting a Word document to PDF keep the author?
It can. Word's PDF export has a Document properties choice under More Options, then Options, and with it selected the fields move into the PDF. Our own export, done with LibreOffice because Word was not on the test machine, carried the title, author, subject and keywords across. Clear that choice or clean the Word properties first, and then read the finished PDF anyway.
Does the Document Inspector strip the location from photos in a document?
The Word inventory Microsoft publishes for the Document Inspector says nothing about metadata stored inside pictures. In our test, the photo in the .docx still carried its camera model and GPS position when a script read it. The safer order is to remove location from the photo before inserting it; on an iPhone, the share sheet has an Options switch for Location.
How do I see who wrote a PDF on a Mac?
Open the PDF in Preview and click the Inspector button in the toolbar. Apple says the Inspector shows information about a document such as file size and the author name, and its More info tab shows where a photo was taken. Acrobat shows more: Document properties, Description, then Additional Metadata, which opens the XMP record that Preview does not display in full.
Is a file with no metadata safe to give an AI?
No, it is only a file that no longer describes its makers. Metadata is the layer outside the page; the names, numbers and addresses in the body are a separate problem, and so are comments and tracked changes. A clean property sheet on a letter that still names the client in its first line protects the author and nobody else. Check both layers before the upload.
How does Nonimo treat the author and properties of a Word file?
It gives you the text with personal details covered, ready to paste, plus a copy of the document with those details covered. For a .docx, the creator, the most recent saver, the company and the manager come back blank, while the title, keywords, subject and custom fields carry placeholders.