[nonimo]
EN
Download

Remove personal data from a CSV before AI sees it

· Written and maintained by Nonimo

To upload a CSV to ChatGPT, Claude or Copilot without the personal data, clean the file before it leaves your computer. Export only the fields the task needs, open a copy, and delete or code every column that names someone. Empty the notes column, save as CSV UTF-8, then read the file in Notepad or TextEdit before you attach it.

The same care applies to plain text notes, LibreOffice Writer documents and RTF files, each with its own places to look.

These are the formats of everyday office exports: the client list from a practice management system, the tenant file from a lettings package, a case note saved as text, a letter written in LibreOffice by a charity that does not pay for Microsoft 365. This guide covers them one by one. Spreadsheets saved as .xlsx have their own guide, and so do PowerPoint presentations.

How to upload a CSV to ChatGPT without the personal data

A CSV is the easiest file there is to check, and the easiest to send with far too much in it. Exports are generous by default: a contacts download can run to dozens of columns, from date of birth to bank details, and the AI needs perhaps four of them. Work in this order:

  1. Export less. Most systems let you pick the fields before the download. Untick everything the task will not use.
  2. Open a copy. Give it a name that marks it as the version for the AI, and never overwrite the original export.
  3. Delete, code or generalise. Remove the columns that identify people, replace names with codes, and blur dates and postcodes the task still needs.
  4. Deal with free text. A notes or comments column goes, or gets read line by line.
  5. Save as CSV UTF-8. The encoding matters for accented names, for the £ sign and for any tool that reads the file afterwards.
  6. Read it as text. Open the saved file in Notepad or TextEdit and look at what will actually be sent.

Where a file goes after you press send is another matter, covered for ChatGPT in the UK and for Microsoft Copilot. This guide is about what you hand over in the first place.

What a CSV carries, and why that is good news

A comma separated file is text. Each line is a row, commas or semicolons split the cells, and that is the whole format: no second sheet, no comment boxes, no formatting that turns a number white. Nothing in a CSV can be switched off from view and left in the file.

The ICO says as much in its guidance on hidden information in documents. That guidance is dated 31 July 2025. Among its steps for checking a file before disclosure, it suggests converting complex files to simpler formats, naming txt and csv, to reveal all the information the document can display. In other words, the regulator uses the format of this guide as a test for all the others.

FormatBesides the visible text, the file can holdFirst thing to do
.csv, .tsvEvery column in the export, however wideRead the header row
.txt, .mdNothing, though a link can hide an addressSearch for names and @
.odtComments, tracked changes, hidden text, header, propertiesAccept changes, delete comments
.rtfComments, deleted text, hidden text, authorOpen it in Notepad and search

Sources: ICO guidance on hidden information in documents; LibreOffice Help; our test files of 28 September 2026.

Every column goes, including the ones nobody reads

The risk in a CSV is not what hides but what sprawls. An export built for a mail merge or a migration includes whatever the database holds, and when you open it in Excel the screen shows the first ten or so columns. National Insurance numbers in column R, a bank sort code in column W and a free text history in column AF are all in the file, and all of them go to the AI.

Excel can also mislead you about the values themselves. Microsoft’s page on keeping leading zeros and large numbers explains that by default Excel removes leading zeros from numerical text and converts it to a number, so a mobile number without spaces can appear with its first zero missing. The file still holds what the export wrote. Reading it in a text editor shows you the characters the AI will receive, not Excel’s interpretation of them.

A CSV saved from Excel holds a single sheet

If the CSV starts life as a workbook, Microsoft’s text import and export page notes that Excel warns you that only the current worksheet will be saved to the new file. The other sheets are left behind, which sounds reassuring and is not a check: you still have to look at the sheet you are saving.

That sheet can have hidden rows. In a test with LibreOffice 26.8, a spreadsheet with a hidden row saved to CSV wrote that row out with the rest, as ordinary data. Hiding was never removal, and a plain text file has no way of hiding anything again. Everything particular to workbooks, from very hidden sheets to pivot caches, is in the Excel guide.

Tabs, semicolons and the .tsv file

Some systems export with tabs between the cells, as a .tsv or a .txt. Excel writes the same thing when you choose Text (Tab delimited), which, as Microsoft’s list of text formats notes, saves only the active sheet. Nothing about the advice changes: a tab separated file is still one grid of text, every column still travels, and the header row is still where you start. Semicolons are the same story with a different mark between the cells.

Deciding what each column needs to be

With the header row in front of you, go through the columns one at a time and give each a verdict: keep, delete, code or blur. The UK GDPR’s minimisation principle says personal data should go no further than the purpose calls for, and a column is a very literal way of applying it. A question about which services earn the most needs services and fees. It does not need a single name.

Column in a typical exportWhat the AI usually needsSend
Client nameSomething to tell rows apartA code such as C01
NI number, UTRNothingDelete the column
Sort code and accountNothingDelete the column
Mobile, emailNothingDelete the column
Full postcodeRough location, if anythingThe postcode district
NotesRarely anythingDelete, or clean by hand

The suggestions are ours. UTR is the Unique Taxpayer Reference HMRC issues for Self Assessment.

Delete columns in the copy rather than clearing them, so the header row no longer promises data. If you work in Excel, remember that deleting a column on one sheet does nothing to the same column elsewhere in the workbook; in a CSV there is only the one sheet, which is one less thing to miss.

Codes in the file, the key somewhere else

When the AI must keep clients apart, say to work out who owes what, give each a code in place of the name and keep the list of codes in a separate file, in a different folder. The ICO’s pseudonymisation guidance explains why: anyone who also has the list can turn coded rows back into people, so they remain personal data and the list needs separate storage and proper security.

Our explainer on pseudonymisation under UK law goes further into what that changes.

Two habits make codes work better. Avoid initials: in a practice of forty clients, MT is Morwenna Treloar to anyone who knows the list. And if the AI will see more than one export, give each client the same code in all of them, or it cannot connect an invoice in one file to a fee in another. A code that means something to you and nothing to an outsider is the aim.

A column without a name can still point at one person: a date of birth next to a full postcode, a job title in a team of four. Which combinations to watch for, with the UK identifiers, is covered in the checklist of what personal data to remove. For a CSV the answer is usually to blur rather than delete, so an age band stands in for the date and the district for the postcode.

Save it as CSV UTF-8, not plain CSV

A CSV has no font and no settings, so the only thing that tells a program how to read its characters is the encoding. UTF-8 covers every alphabet, which matters more than it sounds in a UK client list: Welsh names, Irish names, a client who writes her name with an accent, and the pound sign in every fee column.

Older Windows character sets, such as Windows-1252, hold no more than 256 characters each. That is too few for every alphabet, and a £ sign or a name like Siân or Zoë written in one can reach another tool as stray symbols. Microsoft’s advice on importing contacts from a CSV puts it simply: for best results the CSV should have UTF-8 encoding, because it works for all languages and alphabets.

UTF-8
the encoding Microsoft advises for CSV imports, as it works for all languages and alphabets. Microsoft Support, Outlook contacts

From Excel

Excel offers two CSV formats in the Save As list, and they sit close together. Choose File, Save As, then in Save as type pick CSV UTF-8 (Comma delimited) (*.csv), the file type Microsoft names in that same guide. Next to it in the list sits CSV (Comma delimited), the older option: Microsoft’s list of text formats Excel can save describes it as a file for use on another Windows operating system, and it does not write UTF-8.

Save the CSV last. If the workbook still carries other sheets, hidden rows or notes, deal with those first, as the Excel guide sets out, and only then save the CSV from it.

Here is one row of an invented client list saved in UTF-8 and in Windows-1252, the older Windows character set, then read back by a program that expects UTF-8:

Saved as UTF-8
Client,Fee (£)
Siân Pryce,420

Saved as Windows-1252
Client,Fee (�)
Si�n Pryce,420

The name and the currency are the first casualties, and a tool that cannot read a character may stop reading the file altogether.

Fixing an export already saved the old way

If you already have an export in the old encoding, Microsoft’s guide shows a way round it: in a new workbook choose Data, From Text/CSV, pick the encoding under File Origin that makes the characters display properly, load it, and save the result as CSV UTF-8.

From LibreOffice Calc

In Calc, choose File, Save As, set the file type to Text CSV and tick Edit filter settings before you press Save. The Export text files dialogue that follows lets you set the character set, where Unicode (UTF-8) is the one to pick, and the field delimiter. A comma is what most tools expect; a semicolon is common in files from continental colleagues.

Text cells that contain the delimiter get wrapped in quotation marks, which is where a column of “Surname, Forename” comes from. That is correct CSV, but it is worth noticing, because a name split across a comma is easy to miss when you search the file by eye.

Text files and notes: .txt and .md

A plain text file is the one format where what you see is exactly what there is. No layers, no properties inside the file, no markup that the screen leaves out. The difficulty moves elsewhere: text files are where people write in sentences, and a name in a sentence is not in a column you can delete.

Case notes exported from a practice management system, a transcript of a phone call, minutes typed up after a tenancy inspection: each tends to mention a client by name, then by first name, then as “Mrs T”. A search for the surname finds the first of those and misses the other two, so search for every form, and then for the patterns that give away the details around them:

Notepad and TextEdit both search with Ctrl+F or Command+F.

Email threads and their signatures

A thread saved as text carries more people than the one you meant to ask about. Every reply quotes the one before it, so a single file can hold the client, their partner, a solicitor on the other side and three colleagues. Signatures are the worst part: a name, a job title, a direct line and a mobile, repeated at the foot of every message.

Cut the thread down to the message you need before you do anything else, then deal with the names left in it.

A .md file adds one wrinkle. A link written in Markdown shows its words when rendered and keeps its target in the source, so “email the client” can carry a full address that nobody reading the rendered notes will see. Open the file as text, not in a preview. For the NHS version of this problem, where a clinical note is also special category data, see the guide for NHS staff and ChatGPT.

Emails are worth a mention here because the ICO mentions them: the same guidance suggests saving an email as a .txt file as one way of removing its metadata. The body still needs the same read as any other note, and the UK list of details to strip out is a good prompt for what to look for.

LibreOffice Writer files: the parts of an .odt you never see

A LibreOffice Writer document is a zip archive of XML files, and the page you read is drawn from only some of what is inside. To show how much sits elsewhere, we built a short tenancy review as an .odt with invented names, in the structure LibreOffice writes, and listed the text of every part.

7people named inside the file
1named in the body you read
3XML parts holding names
Our test .odt with invented names; every part listed by script, 28 September 2026
Where a name satPart of the .odtShown on the page?
A comment and its authorcontent.xmlIn the margin, if comments show
Text deleted under track changescontent.xmlNot once changes are hidden
Text formatted as hiddencontent.xmlNo
The page header and footerstyles.xmlOnly in page view or print
Author, last editor, titlemeta.xmlOnly in File, Properties

Source: our test file. Part names follow the OpenDocument format LibreOffice uses.

Seven people, one in the body text, spread over 3 of the XML parts. An upload sends the file, so all seven go. What each place needs is below.

Comments and the names on them

Comments are where a colleague writes “arrears since June, speak to his brother first”. Each one carries its author and a date. LibreOffice’s comments menu offers Delete Comment, Delete Comment Thread, Delete Comment By, which removes everything one author wrote, and Delete All Comments. On the version you are preparing for the AI, use the last one: Edit, Comment, Delete All Comments.

Resolving a comment is not deleting it. LibreOffice marks it Resolved under the date and keeps it in the document.

Tracked changes, hidden text and the page header

A deleted sentence under track changes is still in the file until the change is accepted. Open Edit, Track Changes and, as LibreOffice’s help lists, choose Accept All, which settles every recorded change at once. Do it after you have read them, because the deleted text is often the part that names a previous tenant or client.

Hidden text is formatting, not removal. To see it, tick Hidden characters under Tools, Options, LibreOffice Writer, Formatting Aids, then switch on View, Formatting Marks; the Formatting Aids page describes both. Hidden paragraphs have their own switch, View, Field Hidden Paragraphs.

Headers and footers sit in a different part of the file and are easy to forget; look at them in page view. Footnotes deserve the same read, because a note at the foot of the page is a natural place to name a source or a witness: delete the ones the AI does not need.

Remove personal information on saving, and its limits

LibreOffice has a setting that sounds like the answer. Under Tools, Options, LibreOffice, Security, Options (LibreOffice, Preferences on a Mac), the security options include Remove personal information on saving. It swaps the names of authors on comments and changes for generic ones such as Author1 and Author2, resets their time stamps, and strips user data from the file properties.

What it leaves alone is the content. Every comment and every tracked change stays, with its text, just signed by Author1 instead of a colleague. A comment that names a client names them just as clearly. In a .docx the same markup has to be settled by hand, and our guide to removing track changes and comments in Word gives the menus for it.

Use the setting, and delete the comments and accept the changes as well. The document’s own properties, under File, Properties, General, have a Reset Properties button, and unticking Apply user data stops your name being saved with the file. Properties in Word and PDF files are the subject of our metadata guide.

Saving as plain text is a check, not a cure

Following the ICO’s suggestion, you might save the document as text to see what it holds. As a check it works, and the result surprises people. In a test with LibreOffice 26.8, a Writer document exported as text kept the words deleted under track changes, the hidden text and a hidden section, and dropped the comment, the page header and footer, and the body of the footnote.

3hidden places copied into the text file
4left out of it
A Writer document exported to text with LibreOffice 26.8, our test, 28 September 2026

In that test, 3 hidden places made it into the text file and 4 did not: the text copy shows more than the page does and less than the file holds. Read it for what it reveals, then clean the .odt itself. Writer’s Tools, Redact is a different tool again: it turns the document into a drawing that is usually finished as a PDF, which is covered in redacting a PDF properly.

RTF files: open one in a text editor and read it

Rich Text Format is older than .docx and still turns up, in exports from older systems and in documents saved for someone without Word.

Unlike an .odt or a .docx it is not a zip: it is a text file with its formatting written as codes, so Notepad or TextEdit opens it directly. That makes it the one document format you can check without special tools, once you know what to look for. Opening the same file in Word shows you the formatted page, which is the view that hides the very things you are looking for.

RTF has codes for comments, their authors, hidden text and deleted changes, so saving a Word document as RTF is no guarantee that they have gone. We wrote a short RTF with the same invented tenancy review. These are four of its lines, exactly as a text editor shows them:

{\rtf1\ansi\ansicpg1252\deff0{\fonttbl{\f0 Calibri;}}{\info{\author Ptolemy Carrow}{\company Carrow Lettings}}
\pard Tenant: Seraphina Quarrie{\*\atnid WD}{\*\atnauthor Wilhelmina Dunstable}\chatn{\*\annotation{\*\atnref 1}\pard\plain Rent arrears since June, speak to Cosmo Pengarth first}\par
\pard The inspection found no damage.{\v  Guarantor: Leofric Ashdown}\par
\pard {\deleted Previous tenant: Ottilie Brannagh. }Contact the letting agent with any questions.\par

Reading the codes

Each code tells you something, and each is a word you can search for:

Search the file for each of those words. If any turns up, go back to the source document in Word, accept the changes, delete the comments and run Inspect Document as described in our Word guide, then save as RTF again.

A last read in Notepad or TextEdit

The ICO’s guidance includes a warning worth repeating: you should still check documents yourself, even if you have used a software tool to help you. For the formats in this guide the check is quick, because all of them either are text or can be read as text.

Open the file you are about to upload, not the original, in Notepad on Windows or TextEdit on a Mac. Read the first line of a CSV and count the columns. This read is the cheap check, because an uploaded file can stay in ChatGPT’s Library after its chat is gone, which our guide to ChatGPT’s delete and archive buttons walks through.

Scroll to the end, since exports sometimes add summary rows with a name in them. Search for @, for 07 and for the surname of anyone you know is in the data. If you use Claude, Anthropic’s help centre on uploading files lists CSV, TXT, ODT and RTF among the formats it reads, so any of the four can be attached exactly as it is.

A notes cell that runs over several lines

One thing looks odd the first time you read a CSV as text. When a notes cell contains line breaks, the export wraps it in quotation marks and the text runs over several lines, so a single client can take up five lines of the file. A name written on the third of those lines does not look like part of a row at all. Read those passages with care, or better, drop the notes column before you save.

If a file has passed through several hands, check it against the firm’s list of approved AI tools as well: cleaning a file is no substitute for sending it to the right place. Practices that live on client exports, from bookkeeping to payroll, have a page for accountants, and lettings teams one for estate agents.

What Nonimo puts on the clipboard from a CSV export

Click once on the file in File Explorer or the Finder, then press the Nonimo key. The clipboard then holds its text with the personal details swapped for labels, ready for the chat, and when you paste the reply back the names return. You come away with text, not a second file.

Nonimo reads CSV, TSV, .txt and .md saved as UTF-8, with or without a byte order mark, or as UTF-16 with one. From a LibreOffice .odt or an RTF it takes the body, the tables, and the header and footer.

An invented export, saved as CSV UTF-8. Below, the file and then what Nonimo put on the clipboard, unedited:

The file
Client,NI number,Mobile,Email,Service
Morwenna Treloar,PZ 12 34 56 A,,m.treloar@example.com,Self assessment
Babajide Okonkwo,,07700 900412,,Monthly payroll

The clipboard from Nonimo
Client,NI number,Mobile,Email,Service
[PERSON_1],[REFERENCE_1],,[EMAIL_1],Self assessment
[PERSON_2],,[PHONE_1],,Monthly payroll

The services stay for the AI to work on. HMRC has not used the PZ prefix since August 2002, so that number belongs to nobody. Give the clipboard a quick read before you paste it. For what Nonimo keeps on your machine and what it sends, see how Nonimo handles your data.

Sources

Nonimo is the software that does this on your own computer: it masks client names and IDs before your text reaches ChatGPT. No account needed, and the app does it without your files leaving your machine.

Common questions

How do I upload a CSV to ChatGPT without sharing personal data?

Export only the fields the job needs, then open a copy and delete the identifying columns or swap names for codes, keeping the key in another folder. Empty or delete the notes column, save as CSV UTF-8, and read the file in Notepad or TextEdit before you attach it. The UK GDPR's minimisation principle asks for no more personal data than the purpose needs, and an analysis of fees rarely needs a mobile number.

Can a CSV file hide data the way a spreadsheet can?

Hardly. A CSV is plain text: rows, separators and nothing else, so there are no hidden sheets, comments or cell formats to worry about. That is why the ICO suggests converting complex files to txt or csv to reveal what they display. The catch is the other way round: every column in the export travels, including wide ones you never scrolled across, so read the header row before you upload.

Should I save as CSV UTF-8 or CSV (Comma delimited) in Excel?

Choose CSV UTF-8 (Comma delimited). Microsoft's own advice for CSV imports is UTF-8, because it works for all languages and alphabets, while the older option writes an older Windows character set in which a £ sign or an accented surname can arrive garbled.

Does saving a document as plain text get rid of tracked changes?

Not necessarily. In a test with LibreOffice 26.8, a Writer document saved as text kept the words deleted under track changes, the hidden text and a hidden section, while it dropped the comment and the page header. Accept or reject every change and delete the comments before you export. For the ICO, a plain text copy is something you read to find hidden content, not a cleaning step.

What does Remove personal information on saving do in LibreOffice?

It replaces the names of authors on comments and tracked changes with generic ones such as Author1, resets their times, and strips user data from the file properties when you save. It does not delete a single comment or change, so a note saying who a client is stays in the file. Switch it on under Tools, Options, LibreOffice, Security, Options, and still delete the comments themselves.

Can Claude read .odt and RTF files as well as CSV?

Claude can. Anthropic's help centre lists CSV, TXT, ODT and RTF among the document types it works with, next to PDF and DOCX, and a chat accepts up to 20 files at a time. For what OpenAI and Microsoft keep after an upload, see our UK guides to ChatGPT and Copilot. Whichever tool you use, clean the file before it goes, because the whole file is what arrives.

What should I do with the notes column in a CRM export?

Leave it out unless the task is about the notes. Free text is where staff write the name of a partner, a doctor or a solicitor, a phone number, a reason for a late payment, and none of it sits in a predictable column you can delete. If the AI really needs the notes, move them into a text file, read every line and replace each name by hand.

Can I open an RTF file to check what is in it?

Yes. An RTF is text with its formatting written as codes, so Notepad or TextEdit will show you the raw file. Search it for annotation, which marks a comment and sits next to the name of whoever wrote it, for deleted, which marks text removed under track changes, and for the author entry near the top. If any of those turn up, clean the document in Word and save it again.