[nonimo]
EN
Download

Remove metadata from PDF, Word and photos before AI reads them

· Written and maintained by Nonimo

A Word document or a PDF keeps a second record that never appears on the page: its author, its most recent editor, which company owns the software, when it was printed, its template and, for any photo inside, where that photo was taken. To remove metadata from PDF, Word and photo files before any assistant opens them, you clean that record on a copy, then open the copy and check that each field is empty.

This guide covers that record and nothing else. Why a black rectangle hides nothing in a PDF is explained in the PDF guide on black boxes that fail, and comments and tracked changes inside a Word file belong to a separate guide on redacting in Word. What happens to an upload once it reaches OpenAI is in OpenAI’s own account of how it uses data.

What a PDF or Word file says about you without showing it

Metadata is information stored alongside the content, in sections of the file that no page view displays. In Microsoft’s account, some properties are filled in by people, like a title or a subject line, and others are kept current by Office itself, like the name of whoever saved the document most recently and the day it was first created. You never type most of it. It accumulates as the file moves through an office.

Some of it is harmless. Much of it names people. An author field carries the name from the Office account of whoever created the file, which may be a colleague who left two years ago. The company field often holds the firm’s legal name. A template named for a client, or a title that someone once filled in with the client’s surname, turns a neutral memo into one that identifies the matter before anyone reads a line.

FieldWhere it livesWhat it can reveal
Author, last saved byWord core properties; PDF Info and XMPNames of the people who worked on it
Company, managerWord extended propertiesThe firm, and who the author reports to
Title, subject, keywordsBoth formatsOften the client or matter name
Created, modified, printedBoth formatsWhen the work happened, sometimes a time zone
Template, hyperlink baseWord extended propertiesTemplate names and network folder paths
Custom propertiesWord; custom tab in AcrobatAnything a document system wrote there
Photo EXIFPictures inside either formatCamera, time and GPS position

Sources: Microsoft Support on the Document Inspector; our test files, September 28, 2026.

Why nobody notices it

Metadata is invisible in exactly the view where people check their work. You read the page, fix the page and send the page, and the property sheet rides along untouched. The American Bar Association made the same observation in 2006: many programs automatically embed the name of the computer’s owner, the date and time of creation and the identity of whoever saved it last, and that information might simply be a right click away.

For HR files there is an extra wrinkle: how PHI and PII differ in an employee file affects how sensitive a staff name in a property field turns out to be.

Where to look before you remove anything

Start by reading what your file says. Cleaning blind means you cannot tell afterward whether anything changed, and each program shows a different slice of the record.

In Word for Microsoft 365, the Info page under File shows a short list of properties down its right side. That list is partial. Microsoft’s instructions go two steps further: the link that expands it to every property, and then the Advanced Properties dialog, whose Summary tab holds 8 fields, from Title and Author through Company to Comments. The Custom tab beside it keeps whatever document systems have added.

In Acrobat and in Preview

In Acrobat, Adobe’s help puts the fields under Menu, then Document properties, then the Description tab: title, author, subject and keywords. The button Additional Metadata opens the full XMP record, which is a second copy of much of the same information, stored in a different part of the PDF.

On a Mac, Preview’s Inspector, opened from the toolbar, shows general information about a PDF, which Apple says includes the author name, and its More info tab shows where a photo was taken. It is a quick look, not a full inventory: the XMP record and custom fields need Acrobat or a script. For a law firm reviewing files at volume, our law firm page explains how the text side is handled.

ProgramWhere the properties areWhat it shows best
Word for Microsoft 365File, Info, Show All Properties, Advanced PropertiesSummary, statistics and custom fields
AcrobatMenu, Document properties, Description, Additional MetadataInfo fields and the XMP record
Preview on a MacInspector button in the toolbarAuthor, keywords, photo location

Sources: Microsoft Support on file properties; Adobe help on PDF metadata; Apple’s Preview User Guide. Read September 28, 2026.

One claim memo, two files: what a script read from each

To show what a program gets, we built two test files around an invented claim memo at an invented insurance agency, Castellan & Rusk. Every name in them is made up and belongs to nobody: an author called Harlan Quisenberry, a colleague called Priya Vandermolen who saved it last, a manager called Ray Tesdahl, a claimant called Corliss Vandegrift and a policy number, PL-7730-1186. The photo inside was stamped with the position of a public lawn in Columbus, Ohio.

Replacing names in the body with labels is a different job, with a root guide of its own; here the body stays neutral on purpose.

The body of the memo is two neutral sentences about a collision on I-70. Everything that identifies anyone sits outside the page, in the metadata. Then we read both files back with ordinary free Python libraries, python-docx for the Word file and pypdf and PyMuPDF for the PDF, which is close to what any software does when it opens a document to work on it.

Where our invented .docx kept its properties, part by part The test document shown as a folder. The page holds two neutral sentences and a photo. Beside it, four separate parts hold the metadata: core properties with the author, last saver, title and dates; extended properties with company, manager, template and a network path; a custom property with the claimant's name and policy number; and a people part with a reviewer's email. The photo carries its own camera model and GPS position. claim memo.docx What the page shows Claim review, collision on I-70 Struck from behind, in June A photo of the damage (camera model and GPS inside) No names on the page Core: author, last saved by, title, dates Harlan Quisenberry · Priya Vandermolen Extended: company, manager, template Castellan & Rusk · a claims folder path Custom property: ClientMatter Vandegrift, Corliss (policy PL-7730-1186) People: the reviewer's account pvandermolen@example.com
The property parts of our invented test file, unzipped and read with a script on September 28, 2026

The Word file gave up every field

The script read back all of it. The author and the colleague who saved it last came out by name. So did the company, the manager, the template called Castellan Rusk Claims Memo, and a hyperlink base pointing at a claims folder on the agency’s file server. The custom property named the claimant and her policy number, and a separate part of the file held the reviewer’s work email.

The statistics were just as specific: 14 revisions, an editing clock of 312 minutes, a creation date in June, a last save in September and a print date the day before. Not one of those details is visible when the memo is open in Word. Every one was available to the first program that opened the archive.

The PDF kept two copies of the same story

The PDF told its story twice. Its Info dictionary returned the author, the title with the claimant’s surname, the subject with the policy number, the keywords and the software that made it. Its XMP record, a separate block of XML inside the same file, returned the author again, the creating tool and a unique document ID. Clearing one of the two does not clear the other.

Two details surprised us. The creation date carried a time zone offset, which tells a reader roughly where the computer was. And PyMuPDF and pypdf agreed on every field, so this is not the quirk of one library.

What we storedWord filePDF
Author and last person to saveBoth read backAuthor read back, twice
Company, manager, templateAll read backNot stored
Title, subject, keywordsAll read backAll read back
Claimant in a custom propertyRead backNot stored
Photo with camera and GPSRead back from the pictureNo photo in this file

Test run September 28, 2026, with pypdf 6.10.2, PyMuPDF 1.26.5 and python-docx 1.2.0. Invented data.

Why an upload carries the metadata and a paste does not

When you paste text into a chat window, the property sheet stays behind on your computer. When you upload the file, it travels with the file, because the properties are part of the file. The question is whether anything on the other side looks at them, and OpenAI answers that directly.

Its File Uploads FAQ gives examples of tasks ChatGPT can do with an uploaded document, and under extraction it lists this one: “Extract metadata (author, creation date, etc.) from a document.” That does not mean every conversation reads the author field. It means the fields are within reach of the tool, and a user or a prompt that asks for them will get them. The same FAQ lists all common file extensions for documents among the supported types.

312
minutes of editing time recorded inside our invented test memo, next to 14 revisions

The model is not the only reader

OpenAI’s page on data analysis adds that, for some jobs, ChatGPT produces Python and executes it in a notebook environment. A few lines of that code are enough to open a .docx or a PDF and print its properties, which is exactly what our test script did. Any assistant that can run code on a file you uploaded can do the same, whatever the chat interface happens to show. The same code reads a workbook’s hidden sheets just as easily.

Then there is everything that happens before and after the chat. Files uploaded through a firm’s own tools, agents or integrations are parsed by software that may index properties too, since Microsoft itself presents them as a way to search for documents. How long an upload is kept depends on the product and plan; for Anthropic’s side, see Anthropic’s statements on model training. The only version of the property sheet you control is the one on your own disk.

Word’s property sheet: a broad pass, then a close read

Word gives you two tools for this, one broad and one precise. Use both, on a copy, and in that order. The broad tool empties whole categories; the precise one lets you see what is left and type over what the broad tool leaves in place. Settle the tracked changes before either one, so that both tools work on a file that has already been through accepting or rejecting every change and deleting the comments.

Comments, tracked changes and hidden text are a separate category of the same inspection and a separate job, with their own traps, covered in a guide of their own. Here the focus is the property sheet: author, company, dates, template and custom fields.

The broad pass: Inspect Document

On your copy, Inspect Document sits under Check for Issues on the Info page of the File tab. The category that matters for this job is the one Microsoft labels “Document Properties and Personal Information”. Microsoft lists 10 kinds of item for it.

The first is the content of 3 tabs of the properties dialog, Custom, Statistics and Summary. The rest are email headers, routing slips, send for review details, document server properties, document management policy information, content type information, databinding links, a username and the template’s name. Choose Remove All beside it.

10kinds of item in the properties category of the inspector
3property tabs the inspector empties
8fields on the Summary tab to read afterward
Microsoft Support, Document Inspector and file properties pages, read September 28, 2026

Microsoft also warns that some information the inspector removes cannot be restored, which is the reason for the copy. It also notes that the OpenDocument format needs the inspector run every time you save in it. If your firm keeps files in a document management system, expect server properties there as well, because Microsoft lists them separately. LibreOffice has its own save setting, yet a Writer .odt still keeps comment text and deleted lines after it runs.

The precise pass: Advanced Properties

Then open File, Info, Properties, Advanced Properties and read every tab. Check the Summary fields one by one, and the Custom tab for anything a document system or another program wrote there. Look at the template name especially: a firm template that bears a client’s or a partner’s surname puts that surname in plain sight. If a field survived, clear it by hand or type a neutral value such as the matter type, then save.

A useful habit for a team is to agree on neutral defaults, so that new documents do not start life carrying a client or an employee’s name. It sits well inside an AI acceptable use policy alongside the rules on what may be uploaded at all.

Acrobat and Preview: steps to remove metadata from PDF copies

A PDF needs its own pass, because it stores properties in two places and because many PDFs arrive from outside, already stamped with somebody else’s name, as any CPA firm collecting client paperwork sees every tax season. If you want to remove metadata from PDF copies reliably, clean both the Info fields and the XMP record, then verify with a reader that shows both.

In Acrobat Pro, Adobe documents a sanitize step under All tools, then Redact a PDF: a toggle for sanitizing the file and removing hidden information, followed by a prompt to save the sanitized copy under a new name.

Adobe lists metadata among what sanitizing removes, together with attachments, hidden layers and data from previous saves. The US District Court for the Southern District of Illinois gives the older version of the same advice, with Sanitize Document removing all hidden information and metadata automatically.

Steps to remove metadata from PDF copies: the Info fields and the XMP record are two places A PDF drawn with two separate boxes inside. The Info fields box lists title, author, subject and keywords, and says Description tab edits these. The XMP record box lists the author again, the creating tool and a document ID, and says Additional Metadata shows this. A line underneath says sanitizing clears both. Info fields Title, author, subject, keywords Edited on the Description tab XMP record Author again, creating tool, ID Shown by Additional Metadata Editing one leaves the other; sanitizing is meant to clear both
The two places our test PDF kept its author. Adobe help on PDF metadata and sanitizing, read September 28, 2026

Editing single fields instead

If you only want to change one field, Acrobat lets you edit it under Menu, Document properties, Description. That changes the Info fields you can see there. Open Additional Metadata afterward to check the XMP record, because our test file showed how easily the two copies disagree. For a single sensitive field, sanitizing is the surer route; editing is for correcting a title, not for cleaning a file.

What a court means by basic metadata

The District of Maine’s metadata FAQ, compiled from a 2012 memorandum of the Administrative Office of the US Courts, makes a point worth keeping in mind: PDF files still retain some basic file description metadata, such as the file name and date. A PDF is often safer than a Word file for revision history, but it is never empty by default. On a Mac, Preview’s Inspector will show you the author; it does not show you the whole record.

Remove metadata from PDF exports before they leave Word

Many PDFs start as Word files, and the export is where property sheets quietly cross from one format to the other. The step people rely on to “flatten” a document carries the names across with it. It is an easy way for a client’s name to end up in the properties of a PDF that nobody ever opened in Acrobat.

Microsoft’s instructions for saving a Word document as a PDF go through File, Save a Copy, then PDF, then More Options and Options. One of the choices there reads Document properties, and Microsoft’s wording is plain: if you want to include document properties in the PDF, make sure it is selected. Clear it before you export a file that is headed anywhere outside the office.

Remove metadata from PDF exports: which fields crossed from our Word file into the PDF Two boxes joined by an arrow labeled export to PDF. The Word file on the left lists title, author, subject, keywords, company, template and the photo's GPS. The PDF on the right lists the four fields that arrived: title, author, subject and keywords. Company and template did not arrive, and the photo's EXIF was not found in the exported image. The Word file Title, author, subject, keywords Company and manager Template name Photo with GPS export to PDF The exported PDF Title, author, subject, keywords Company: not carried Template: not carried Photo EXIF: not found
Our test .docx exported to PDF with LibreOffice and read with pypdf, September 28, 2026

What crossed in our test

We do not have Word on the test machine, so we exported our Word file with LibreOffice instead and read the result.

Four fields crossed into the PDF: the title with the claimant’s surname, the author, the subject with the policy number and the keywords. The company and the template did not. And the photo inside the exported PDF no longer carried EXIF when we checked it. Another program may behave differently, which is why the check on the finished PDF is the step that counts.

Printing to PDF is sometimes offered as the fix. It does rebuild the file, and a federal court FAQ calls it one of the ways to remove revision metadata, but the same FAQ notes that PDFs still keep basic description data such as the file name and date. A deck needs one more look before export, since PowerPoint’s PDF options can bring comments and hidden slides along.

Whatever route you take, open the result and read its properties before you upload it. If cyber insurance is part of the conversation at your firm, this is the kind of control an underwriter asks about.

Photos inside the document carry their own GPS and camera

A picture pasted into a report is a file inside a file, with its own metadata. Phone cameras normally write EXIF into each photo, a driver’s license copy included: the camera make and model, the time, and, when location is on, the latitude and longitude. The document’s properties can be spotless while the photo on page three still says where it was taken and on what phone.

Our test photo kept all of that inside the .docx. The script found the invented camera make and model, the moment the shutter fired, the claimant’s name stored as the photographer, and GPS coordinates that resolve to a lawn in Columbus. Microsoft’s inventory of what the inspector finds in Word covers properties, comments, headers, hidden text and custom XML, among others, and it does not mention metadata stored inside pictures.

What the Army learned from four helicopters

In March 2012 the US Army published a warning to soldiers called Geotagging poses security risks. It cites a case from 2007: soldiers photographed a new fleet of helicopters on the flightline at a base in Iraq, the photos were uploaded to the internet, and from them the enemy determined the exact location of the helicopters inside the compound and destroyed four AH-64 Apaches in a mortar attack.

2007
the year geotagged photos from a base in Iraq gave away where four Apache helicopters were parked. US Army, 2012

An insurance claim photo is not a flightline, but the mechanism is the same: coordinates nobody meant to share, in a place nobody looked. A photo of a client’s damaged kitchen places the client’s home. For health practices, a patient photo with location is worth checking against what HIPAA asks of an AI tool before it goes near a chatbot.

Take the location out first

The cleaner order is to remove location from the photo before you insert it. Apple documents a switch for this on iPhone and iPad: in the Photos app, select the photos, tap the share button, then Options, and turn off Location; the sheet then shows Location Not Included. None of the big four assistants says whether it strips location from an uploaded photo, as our guide to what they keep from your photos found.

Apple also explains how to keep the Camera from saving location in the first place: in Location Services, under Privacy & Security in Settings, set Camera to Never. For photos already in a document, replace them with copies that were cleaned first. Our insurance agency page covers the text of claim files.

What US lawyers and federal courts already say about metadata

None of this is a new worry, and American legal institutions have been writing about it for two decades. The clearest statement is ABA Formal Opinion 06-442, issued on August 5, 2006. It dealt mainly with the lawyer who receives metadata, and concluded that the Model Rules do not specifically prohibit reviewing and using it.

The same opinion points the burden back at the sender. A lawyer who is concerned about sending a document that contains metadata, it says, may be able to limit the likelihood of its transmission by scrubbing it or by sending a different version without the embedded information.

Its footnote cites the Model Rule 1.6 comment on acting competently to safeguard client information. State bars have written their own opinions since, and court orders on AI add another layer, which our page on AI disclosure in court filings follows.

Courts that publish their own instructions

Federal courts publish practical instructions because filings land on public dockets. The Southern District of Illinois has a short guide on removing metadata with the Document Inspector and with Acrobat’s sanitize tool. The District of Maine’s FAQ explains why editable formats are where accidental disclosure usually happens, and why PDFs still keep basic description data.

An AI upload is not a filing, and no court rule governs what you give a chatbot. But the same logic applies.

A property sheet that should not reach opposing counsel should not reach a third party service either, and whether the upload itself becomes a problem for the firm comes up in the guide on breach notice after an upload. For privilege questions, redacting before AI sets out what a clean copy does and does not protect.

Nonimo and a file’s properties

The Nonimo app is installed locally. Drag a file onto it and you get two results: a version of its text where placeholders stand in for personal details, for pasting into the AI. You also get a copy of the file in which those details are covered. Below is what Nonimo made of the text of an invented claim memo.

What you typed                           What goes to the AI
Claim review: collision on I-70          Claim review: collision on I-70
Claimant: Corliss Vandegrift             Claimant: [PERSON_1]
Insured: John Doe                        Insured: [PERSON_2]
Policy number: PL-7730-1186              Policy number: [REFERENCE_1]
Email: cvandegrift@example.com           Email: [EMAIL_1]
Phone: (614) 555-0182                    Phone: [RECORD_FIELD_1]

When the answer arrives, the app puts the original names back into it, on your own machine. In the Word copy the app hands back, the creator, the most recent saver, the company and the manager come back blank; placeholders replace the title, subject, keywords, description, category and custom fields; the record of who edited it, emails included, is blanked, the thumbnail is dropped, and pictures inside are stripped of EXIF.

For a PDF, the copy is a clean new file with the text and none of the original’s metadata. The data the app saves on your computer is described on the security page, and the license page lifts the monthly word limit.

Sources

Checked September 28, 2026.

Nonimo is the software that does this on your own computer: it masks client names and IDs before your text reaches ChatGPT. No account needed, and the app does it without your files leaving your machine.

Common questions

Which tool should I use to remove metadata from PDF files before an upload?

In Acrobat Pro, the dependable one is the sanitize option under All tools, then Redact a PDF, which saves a cleaned copy; Adobe lists metadata among what it takes out. Reopen that copy and read Document properties and Additional Metadata. OpenAI names pulling a document's author and creation date as one task ChatGPT can do on request, so expect those fields to be readable.

Which details does a .docx store about the people who made it?

More than the author. Our invented test file held its creator, the colleague who saved it most recently, the agency's name, a manager, the template, a folder on a file server, fourteen revisions, 312 minutes of editing, a print date and a reviewer's email. Microsoft notes that Office keeps some of this up to date on its own, such as who saved the file most recently and when it was created.

Can ChatGPT see who wrote a document you give it?

It can if you ask. OpenAI's help center counts getting a document's author and the date it was made among its extraction examples, and its data analysis page says ChatGPT can write Python and execute it for some jobs, which is exactly how a program reads a file's property sheet. Treat every field as readable, and clean the copy before it leaves your computer.

Does exporting a Word document to PDF keep the author?

It can. Word's PDF export has a Document properties choice under More Options, then Options, and with it selected the fields move into the PDF. Our own export, done with LibreOffice because Word was not on the test machine, carried the title, author, subject and keywords across. Clear that choice or clean the Word properties first, and then read the finished PDF anyway.

Does the Document Inspector strip the location from photos in a document?

The Word inventory Microsoft publishes for the Document Inspector says nothing about metadata stored inside pictures. In our test, the photo in the .docx still carried its camera model and GPS position when a script read it. The safer order is to remove location from the photo before inserting it; on an iPhone, the share sheet has an Options switch for Location.

How do I see who wrote a PDF on a Mac?

Open the PDF in Preview and click the Inspector button in the toolbar. Apple says the Inspector shows information about a document such as file size and the author name, and its More info tab shows where a photo was taken. Acrobat shows more: Document properties, Description, then Additional Metadata, which opens the XMP record that Preview does not display in full.

Is a file with no metadata safe to give an AI?

No, it is only a file that no longer describes its makers. Metadata is the layer outside the page; the names, numbers and addresses in the body are a separate problem, and so are comments and tracked changes. A clean property sheet on a letter that still names the client in its first line protects the author and nobody else. Check both layers before the upload.

How does Nonimo treat the author and properties of a Word file?

It gives you the text with personal details covered, ready to paste, plus a copy of the document with those details covered. For a .docx, the creator, the most recent saver, the company and the manager come back blank, while the title, keywords, subject and custom fields carry placeholders.