Remove metadata from Word and PDF files before any AI upload
· Written and maintained by Nonimo
A Word file or a PDF says things about itself that never appear on the page: who wrote it, who saved it last, which firm it came from, the template behind it, when it was started and how long someone spent on it. To remove metadata from Word and PDF files before an AI upload, you clear the properties in each program, deal with any photos inside, and then read the properties back on the copy you actually send.
That matters more with AI tools than with email, because an upload hands over the whole file and the tool can read all of it. We built a test letter to show how much a clean looking page can carry, and this guide works through Word, embedded photos and PDFs in turn, with the menus Microsoft, Adobe and Apple publish for Australian users.
What the page itself says is a separate job. Hiding names in the visible text of a PDF is dealt with in the Australian guide to redacting a PDF, while tracked changes, comments and hidden text inside a Word document are covered in another guide in this series. This one is about the properties.
Remove metadata from Word and PDF: the order that works
Doing it in this sequence saves you cleaning twice. Work from a copy, clear the file’s description, deal with the pictures, and only then check.
| Step | Word | |
|---|---|---|
| Keep the original | File > Save As, new name | duplicate the file first |
| See the properties | File > Info, Show All Properties | Menu or File > Document Properties |
| Clear them | Inspect Document, Document Properties and Personal Information | Redact a PDF, Sanitize document |
| Photos | remove location on the phone first | same |
| Read back | File > Info on the saved copy | Document Properties on the saved copy |
Sources: Microsoft Support (Australia); Adobe Acrobat Australia; Adobe Acrobat; Apple Support (Australia).
The table is deliberately short, and the detail sits in the sections below.
The habit that runs through all of it is reading the properties back on the saved file, not trusting the dialog you just clicked through. For practices that send client files out every day, such as accountants handling TFNs, that check takes a minute and it is the step people skip. When you remove metadata from Word and PDF versions of the same document, check both, because each keeps its own set.
The properties behind a clean page: our test letter
To see how much a file reveals beyond its page, we built a short letter of advice from an invented Melbourne law practice, Karrawong Legal, to an invented client, Mirela Quennell. Every name, number and path in it is made up. The page had been cleaned: the heading reads “Settlement advice: [CLIENT]”, and nothing in the visible text names the client or the staff who worked on it.
We then filled in the properties the way they fill up in real offices, where a template sets some, Word sets others, and a colleague types a few into the dialog to help with filing. We saved it as a .docx with a phone photo inside, and made a PDF of the same letter.
What two standard libraries read out of it
We read the files back with python-docx and pypdf, two ordinary Python libraries, plus Pillow for the photo. None of this needs special software: the Word properties are small XML files inside the .docx, and the PDF keeps its description in plain fields. This is what came back from the Word file, beyond the page itself.
| Where in the file | Field | What our test file said |
|---|---|---|
| Core properties | Last saved by · Author | Anika Borrowdale · Tobias Rendle |
| Core properties | Title · Subject | Quennell v Stonebark Joinery · Mirela Quennell, unfair dismissal |
| Core properties | Keywords · Comments | Quennell, settlement, KL-2026-0417 · client prefers calls after 4 pm |
| Extended properties | Company · Manager | Karrawong Legal · Tobias Rendle |
| Extended properties | Template · Link base | the practice’s employment template · a server path ending in Quennell |
| Custom properties | MatterNumber · ClientName | KL-2026-0417 · Mirela Quennell |
| The photo inside | EXIF | phone make and model, date and time, GPS coordinates |
Our test file, read with python-docx 1.2.0 and Pillow 11.3.0.
Nothing in the file was unusual. Each field is one a busy office fills without thinking, and none of them shows when the document is printed or viewed on screen. The same is true outside Microsoft Office, where a LibreOffice Writer .odt file keeps its author and title in a small file of its own.
The comment field deserves a second look. “Client prefers calls after 4 pm” names nobody, and it is still information about the client, typed by someone who never imagined it leaving the building. If your engagement letters tell clients how their files are handled with AI, the properties are part of that promise, and the engagement letter AI clause has wording for it.
Why a file’s metadata matters once an AI tool has the whole file
When you paste a paragraph into a chatbot, you choose every word that goes. When you upload a file, the tool receives the file, and what it does with the parts you cannot see depends on the tool and on what you ask of it. A slide deck makes the point plainly, since its speaker notes and skipped slides are parts of the file the slideshow never displays.
OpenAI’s File Uploads FAQ is direct about it. One of the extraction examples it gives is pulling a document’s metadata, such as who wrote it and when it was created, and it explains that for some analysis ChatGPT can write and run its own Python code. Reading the properties of our test letter took a few lines of that kind of code.
We make no claim about what happens inside any model. The point is simpler: the properties are readable, and nothing about an upload hides them.
The fields an assistant might quote back
A request as ordinary as “summarise this and tell me who it’s from” invites a tool to look beyond the body. If the author field names your senior associate and the title names the client, an answer can repeat both, into a chat history that OpenAI keeps on its own schedule. The same applies to any provider, and Copilot’s handling of documents has a page of its own.
There is a quieter risk as well. A file often goes to an AI tool precisely because it has been cleaned for that purpose, and the person uploading it believes the job is done. The properties are the part of the file nobody looked at while cleaning.
Word document properties: author, last saved by, company
Microsoft describes document properties, or metadata, as details like author, subject and title, together with fields Word keeps up to date by itself, such as the creation date and the name on the most recent save. Some are typed by people; many are written by the software without anyone noticing.
That is how our test letter came to carry two staff names: one person started it as author, and another saved it last. Neither name appears anywhere on the page, and neither person typed it into the file on purpose.
Where Word shows them
Open File > Info. The column on the right lists the main properties, and a link beneath them, Show All Properties, expands the list. According to Microsoft’s properties page, some metadata, “such as Author”, has to be removed by right clicking the property and choosing Remove.
For the full set, choose Properties at the top of that column and then Advanced Properties. The Author, Manager and Company fields sit on the Summary tab next to Title, Subject, Category, Keywords and Comments.
The Custom tab holds anything a matter system or a colleague added, and in our test file it was where the client’s name sat in full. Custom fields are common in firms that file documents by matter, and law practices tend to have the most of them.
Dates, editing time and the revision count
Word keeps the date a file was created and last saved, a revision number, and the total editing time. In our test file those said the letter was started on 3 August, saved for the fourteenth time on 25 September, and worked on for 187 minutes.
None of that is a name, but it can still identify a matter to someone who knows the practice, and it can contradict what a letter says about when advice was given. Microsoft lists the Statistics tab among the things the Document Properties and Personal Information inspector removes.
The template name and the network path
Two fields slip past because no one types them. The template name records which .dotm or .dotx file the document was built on, and practices often name templates after a client, a partner or a type of matter. The hyperlink base can hold a folder path on your server, and in our file it ended in the client’s surname.
Microsoft’s table for that inspector includes “Template name” and “Username” among the items it finds and removes, alongside email headers and routing slips. That makes the inspector the most reliable way to deal with fields you cannot see in the Info column.
Clearing them with the Document Inspector
The route is File > Info > Check for Issues > Inspect Document. Select the Document Properties and Personal Information box, and Custom XML Data too, run Inspect, and use Remove All for each item it reports. Microsoft advises working on a copy, since data the inspector removes cannot always be brought back. Accept the tracked edits and delete the comments before this pass; taking track changes out of Word permanently covers that step.
| Inspector | What Microsoft says it removes |
|---|---|
| Document Properties and Personal Information | all three properties tabs, Custom included; email headers; routing slips; username; template name |
| Custom XML Data | custom XML stored within the document |
| Comments, Revisions, Versions, and Annotations | reviewer notes, tracked edits, version details, ink |
Source: Microsoft Support (Australia), the Document Inspector page.
Excel and PowerPoint keep the same fields
The same inspector exists in Excel and PowerPoint, and Microsoft’s page lists a few extras for workbooks. Beyond author, title and the name of the last person to save, an Excel file can hold printer properties, such as a printer path, and file path information for publishing web pages. The cells have hiding places of their own, such as hidden sheets and values formatted to look blank, which is why redacting a spreadsheet starts by unhiding everything.
That matters because spreadsheets are among the files most often handed to an AI tool for analysis: a payroll export, a debtor list, a tenancy ledger. The rows need their own cleaning, and the properties sit on top of them. Run the inspector in whichever program made the file, not only in Word.
The last row belongs to a different job, the content inside the document, which is where a colleague’s name tends to appear next to a comment. On a Mac, Word’s menus are laid out differently, so follow Microsoft’s page for the version you have.
Photos inside the document: camera, time and GPS
Most phone cameras write their own metadata, called EXIF, into every picture, a photo of your drivers licence included, before it goes anywhere near a document. Apple’s Australian guide explains that when Location Services is on for the Camera, the coordinates where the photo was taken “are embedded into each photo and video”. It adds that people you share them with “may be able to access the location metadata and learn where it was taken”.
The photo in our test file was a typical example. It still carried the make and model, Apple and iPhone 15, the time it was taken, 14 August 2026 at 16:42, and the coordinates -37.8183 and 144.9671, which we pointed at a public place in central Melbourne rather than anyone’s address.
Photos end up in Word files all the time: a damaged workbench in an insurance claim, a mouldy wall in a tenancy dispute, a wound in an incident report, a vehicle in a claim an insurance broker is preparing. The photo is there for what it shows, and the location it quietly carries is often a client’s home.
Removing location before the photo goes in
The reliable fix is at the source. Apple’s guide gives three routes on an iPhone: tap Options in the share sheet and turn Location off before sharing, adjust a single photo to No Location in Photos, or stop the Camera recording location at Settings > Privacy & Security > Location Services > Camera, set to Never.
We built our test file with a library that keeps pictures exactly as they are. We did not test whether a given version of Word keeps or drops EXIF when you insert or compress a picture, so treat any photo you did not clean, a NSW licence photo included, as carrying its location. The chatbots are no clearer: none of the four big ones says whether it strips that data from an uploaded photo, as our guide to what AI tools keep from your photos found.
For disability and health providers the stakes are higher, because a photo from a home visit can place a participant, and the NDIS case note guide covers what else should come out of those files.
PDF metadata: the Info dictionary and the XMP copy
A PDF describes itself in two places. The older one is a set of document information fields: Title, Author, Subject, Keywords, Creator, Producer and the creation and modification dates. The newer one is an XMP packet, a small block of XML inside the file that can repeat the same details and add more.
Our PDF of the test letter held both, and pypdf read both without effort. The Info fields named Tobias Rendle as author and the Quennell matter as title and subject; the XMP packet repeated the author and the title.
Clearing one copy leaves the other
To see why the two matter, we cleared only the document information fields, the way a program does when it edits only the visible properties, and read the file again. The fields came back empty. The XMP copy still named the author and the client matter.
That is why a PDF check has to look for both, and why a tool built to remove everything is safer than editing fields one at a time.
The date stamp and its time zone
The creation and modification dates in a PDF are stamped with more than a day and a time. In our test file the creation stamp ended in +02’00’, the offset from UTC of the computer that produced it, because that is how the software recorded local time.
An offset on its own identifies nobody. Alongside a creation time, a template name and an author, it helps place a document in an office and a working day, and that is the kind of combination that turns scattered details into a person.
Document Properties and Additional Metadata in Acrobat
Adobe’s four step page on PDF metadata goes like this: open the PDF, choose Menu (Windows) or File (macOS) > Document Properties, change or clear the metadata there, including what sits under Additional Metadata, then press OK and save. Adobe warns that once you delete PDF metadata “it can’t be undone”. It also notes that permission settings on a shared PDF may stop you deleting it at all.
Additional Metadata is where the XMP details show up, so it is the part of that dialog to open every time. If the PDF came from Word, look for the Word title and author there, and remember that the Word file you kept is a separate job.
The Sanitize document option in Acrobat Pro
For a thorough clean, Adobe’s Australian page on sanitising and redacting sets out three steps: make a copy of your PDF, select Redact a PDF from the tools, then choose whether to selectively remove or remove all hidden data. The Sanitize document option sits at the bottom of the redaction options. Adobe says sanitising removes “hidden PDF data and metadata” so that sensitive information is not passed on when you share or publish the file.
Sanitising deals with the file’s description of itself. Covering what the page says is redaction, and the PDF redaction guide covers the tests that prove it worked.
APP 11, the OAIC and personal information embedded in files
The Privacy Act reaches metadata through its definition of personal information in section 6(1), which covers information or opinions about a person who is identified, or who can reasonably be identified. The OAIC’s guide What is personal information? adds that this can include information about a person’s business or work activities. An author field naming a staff member and a title naming a client both fit.
The question the OAIC already asks
The OAIC’s Guide to securing personal information, published on 5 June 2018 for APP 11, asks two questions that read as if they were written for this guide. One, under web security: are there clear policies and procedures for “the identification and removal of embedded personal information from files before they are published online”? The other, under staff training: are staff trained to identify and remove embedded personal information not intended for public release?
The guide was written with publishing in mind, and a chatbot upload is not a web page. The logic carries across all the same: if a practice has to think about names embedded in files it posts online, it has the same reason to think about files it hands to an outside service.
For a small practice without a privacy officer, the practical version of that question is one line in the office procedure: properties are cleared before any file leaves, whether it is going to a court, a client or an AI tool.
When the upload itself is the disclosure
The OAIC’s AI guidance for businesses, first issued in October 2024, applies APP 6 to anything entered into an AI system, which confines use and disclosure to the purpose the information was collected for. A client’s name in a title field is entered just as surely as a name in the body. For everyday tools, that test is applied step by step on the page about which AI tools the Act allows.
Whether an upload like that must be notified is its own assessment, set out in the guide to chatbot uploads and data breaches. This guide describes the law in general terms, not your file.
After you remove metadata from Word and PDF, read the copy back
The last step is the one that proves the others worked. Open the saved copy, not the file you were editing, and look again: Word’s File > Info view with every property showing, Document Properties and Additional Metadata in Acrobat.
Look at the file name while you are there. It travels with every upload, it is often the first thing a tool repeats back, and practices routinely name files after the client and the matter. Renaming the copy costs nothing.
Check four things in particular. The author and last saved by fields should be empty or generic. The title and subject should not name a client, even if the file name does not. The custom tab should hold nothing a stranger could use. And any photo should have been cleaned before it went in, because nothing on this list reads EXIF.
What a clean properties panel does not prove
Clearing the properties says nothing about the text. A letter with empty metadata can still name the client in its second paragraph, and a table of staff can identify everyone in it without a single property set. The body needs its own pass, and where it holds health or union details, the Act’s rules on sensitive information raise the bar for it.
The opposite also holds. Replacing names in the body leaves the properties untouched, which is how our test letter came to have a clean page and five properties naming the client. Labels in place of names are pseudonymisation, not de-identification, as de-identified vs pseudonymised explains.
What Nonimo hands back from a Word or PDF file
Drop a document into the Nonimo app and you receive the text with identifiers covered, ready for the AI tool’s prompt, and a copy of the file with the same details covered.
Inside a returned Word file, the fields for who wrote it, who saved it last, the company and the manager are emptied. Custom properties, keywords, category, title, subject and description are covered like the text, and embedded images lose their EXIF. A PDF returns as a clean new PDF of the text, with the original’s metadata stripped out.
This is how Nonimo handled text taken from the test letter, the same kind of detail that sits in its properties:
BEFORE AFTER (Nonimo)
Karrawong Legal: settlement advice Karrawong Legal: settlement advice
Client: Mirela Quennell Client: [PERSON_1]
Prepared by: Tobias Rendle Prepared by: [RECORD_FIELD_1]
Matter number: KL-2026-0417 Matter number: [REFERENCE_1]
Client mobile: 0491 570 156 Client mobile: [PHONE_1]
Client email: mirela.quennell@example.com Client email: [EMAIL_1]
The app puts the real details back into the AI tool’s reply on your screen. For a Word file you will send as it is, a quick pass with the Document Inspector tidies the template name, the link base and the revision dates as well. The security page explains what the app keeps on your computer, and why the only thing it sends us is a daily count with no text in it.
Sources
- Microsoft Support (Australia), Remove hidden data and personal information by inspecting documents, presentations, or workbooks. Document properties as metadata, including who most recently saved the file; the Document Inspector route; the items the Document Properties and Personal Information inspector finds and removes, including username and template name; the advice to work on a copy. support.microsoft.com
- Microsoft Support (Australia), View or change the properties for an Office file. File > Info, Show All Properties, removing Author by right click, Advanced Properties and the Summary tab fields. support.microsoft.com
- Adobe Acrobat, Remove metadata from a PDF in 4 steps. Document Properties, Additional Metadata, and the warnings that deletion cannot be undone and that permissions can block it. adobe.com
- Adobe Acrobat Australia, Remove sensitive content and information by sanitising and redacting. Make a copy, Redact a PDF, the Sanitize document option, and removal of hidden PDF data and metadata. adobe.com/au
- Apple Support (Australia), Manage location metadata in Photos, published October 2025. Coordinates embedded in each photo, what recipients can learn, and the Options, No Location and Camera settings. support.apple.com/en-au
- OpenAI Help Center, File Uploads FAQ. Extracting metadata (author, creation date) listed among extraction examples; ChatGPT writing and running Python code for some data analysis. help.openai.com
- OAIC, Guide to securing personal information, published 5 June 2018. The APP 11 questions on removing embedded personal information from files and on staff training. oaic.gov.au
- OAIC, What is personal information?, published 5 May 2017. Personal information can include information about a person’s business or work activities. oaic.gov.au
- OAIC, Guidance on privacy and the use of commercially available AI products, 21 October 2024, updated 17 January 2025. APP 6 applied to personal information input into an AI system. oaic.gov.au
- Privacy Act 1988 (Cth), section 6(1). The definition of personal information. legislation.gov.au
- Nonimo, test letter, 28 September 2026. An invented letter built as a .docx with properties and a photo carrying EXIF, and as a PDF with document information fields and XMP; read with python-docx 1.2.0, pypdf 6.10.2 and Pillow 11.3.0; the Nonimo output shown above. Every name, number, path and coordinate is invented or points to a public place.
Nonimo is the software that does this on your own computer: it masks client names and IDs before your text reaches ChatGPT. No account needed, and the app does it without your files leaving your machine.
Common questions
How do you remove metadata from Word and PDF files?
Start from a copy. In Word, choose Inspect Document from File > Info > Check for Issues and clear Document Properties and Personal Information, then look at File > Info again. In Acrobat Pro, open Redact a PDF and use Sanitize document, or empty the Document Properties fields. Strip location from phone photos before they go in. In the Word copy the Nonimo app returns, the author and company fields are emptied too.
Does an uploaded file tell the chatbot who wrote it?
It can, because it has the file. OpenAI's File Uploads FAQ gives metadata extraction, author and creation date included, as one of the jobs ChatGPT will take on for an uploaded file, and adds that ChatGPT can write and run its own Python code for some analysis. Our test pulled every property out with two ordinary Python libraries in a few lines. Nonimo covers the details in pasted text and in the copy it returns.
What is the fastest way to see a Word file's metadata?
Open File > Info. The column on the right lists the main properties, and a link under it expands the full set. Microsoft notes that a few fields, Author among them, are removed by right clicking the property and choosing Remove. Properties > Advanced Properties opens the dialog with the Summary, Statistics and Custom tabs, and the Custom tab is where client names added by filing systems tend to sit.
Can a photo inside a Word document show where it was taken?
Yes, when the photo carries location metadata. Apple explains that with Location Services on for the Camera, coordinates are written into each photo. The phone photo in our test file still held the camera model, the time and GPS coordinates to four decimal places. Switch Location off under Options in the share sheet before the photo goes in. The Nonimo app's Word copy drops the EXIF from embedded images.
What is XMP metadata in a PDF?
It is a second description of the document, written as XML inside the PDF, next to the older document information fields. Both can hold the author, the title and dates. In our test, emptying the older fields left the XMP copy naming the author and the client matter, so any check has to cover both. Acrobat's Sanitize document option is designed to take out hidden data and metadata in one pass.
Does Acrobat's Sanitize document option remove metadata?
Adobe says sanitising takes hidden PDF data and metadata out of the file so it is not passed on when the PDF is shared or published. Its Australian page places the option under Redact a PDF, with a choice between removing items selectively and removing all of them. Work on a copy, because Adobe warns deleted metadata cannot be brought back. Nonimo works differently: a PDF returns as a new file with only the text.
Is an author name in the file properties personal information?
Usually, yes. Section 6(1) of the Privacy Act covers information about a person who is identified or can reasonably be identified, and the OAIC says that can include information about someone's work. A colleague's name in the author field, or a client's name in the title, is personal information that leaves with the file. The Nonimo app empties the author and last saved by fields in its copy of a Word file.
Does Nonimo strip the metadata from a PDF?
A PDF dropped into Nonimo gives you that text with identifiers covered for pasting, and a fresh PDF to download holding only the text, with the original's metadata stripped out. Scans are pictures, and the Nonimo app, on Mac and Windows, reads a fully scanned PDF on the computer and returns its text with the details covered. For a PDF you plan to send in its original form, sanitise a copy in Acrobat Pro.