[nonimo]
EN
Download

Remove personal data from a CSV before AI reads it

· Written and maintained by Nonimo

To mask data in a CSV file that is going to an AI tool, work on a copy. Delete the columns your question can do without, and replace any names that have to stay with codes whose key lives in a separate file.

Then write it out with Excel’s UTF-8 CSV option, open that copy in a plain text editor and search for two or three surnames you know were in the original. A CSV has nowhere to hide anything, so the editor shows exactly what the chatbot will get.

That is the easy format. The ones that travel with it are harder: plain text notes, documents from LibreOffice Writer and the odd RTF file that a Word document became along the way. OpenAI’s page on working with files in ChatGPT lists CSV and TXT among the formats you can upload, and the upload takes the file exactly as it is.

The regulator’s expectation is specific. In its guidance on commercially available AI products, the OAIC says it “expects organisations to consider what information is necessary and assess whether there are ways to minimise the amount of personal information that is input into the AI product”. Deleting columns is the most direct way to do that, and what the Act expects when you pick an AI tool covers the rest of that page.

How to mask data in a CSV file at a glance

A CSV is a grid written out as text, one row to a line, the cells split by commas. There are no formats, no comments and no second sheet. That makes it easier to clean than a workbook, and less forgiving: a column you forgot about is not tucked away somewhere, it is right there in the file, in every row.

The same page covers three other formats that tend to end up in the same folder. Excel workbooks and PowerPoint decks have their own guides, redacting names in Excel and redacting text in PowerPoint, and a spreadsheet kept in LibreOffice Calc is best saved as .xlsx and handled the Excel way. Word documents are covered in the Word redaction guide.

FormatWhat can travel without being obviousDo this first
CSV or TSV (.csv, .tsv)every column, every row, hidden ones includeddelete columns and rows on a copy
Plain text (.txt, .md)nothing hidden; whatever an export put thereread the whole file before sending
LibreOffice Writer (.odt)comments, hidden text, tracked changes, headers, footnotes, a page pictureaccept changes, delete comments
Rich Text (.rtf)comments and their authors, revisions, headers and footersclean the source, then save as RTF

Tested on invented files with LibreOffice 26.8 on 29 September 2026; the RTF row with a file LibreOffice wrote, not Word.

What a CSV export carries

A CSV rarely starts life as a CSV. It comes out of practice management software, a CRM, a booking system or a payroll run, or somebody saves a worksheet in that format because an upload form asked for it. Each route decides what goes in, and none of them asks you.

One sheet, and every column in it

Microsoft’s page on saving a workbook in text formats repeats one line for every CSV and text option: “Saves only the active sheet.” Any other sheets in the workbook stay behind, which is useful. The page also warns that “all formatting will be removed”, so colour coding and bold headings that told a human reader which column was sensitive disappear.

What does not disappear is data. Every column on the active sheet goes into the file, including the ones you scrolled past, and a column’s width, colour or position on screen has no effect on whether it is written.

Hidden rows and columns come along

We tested the obvious shortcut. In LibreOffice Calc we built a small job list for three invented clients, hid the Mobile column, hid the row for Maree Quintrell and added a second sheet of email addresses. Then we exported the sheet as CSV.

The file held all three clients, the Mobile column and the hidden row, with Maree Quintrell’s mobile, 0491 570 156, sitting in it. The second sheet was not there. Hiding is a way of arranging the screen, and a CSV export does not look at the screen.

Whether Excel treats hidden rows the same way we have not tested, so check your own copy in a text editor; the guide to redacting names in Excel covers hidden rows and sheets inside the workbook itself.

How to mask data in a CSV file: hidden rows and columns still go into the export On the left, the Calc sheet as it appears on screen: columns Client, Job, Hours and Status, with two clients, Bronwyn Tallis and Tobias Rendle, and a second tab called Contacts. On the right, the CSV that LibreOffice wrote from it: five columns including the hidden Mobile column, and three rows including the hidden row for Maree Quintrell with her mobile number. The Contacts sheet is not in the CSV. The sheet on screen The CSV LibreOffice wrote Client · Job · Hours · Status Bronwyn Tallis · Tax return Tobias Rendle · Bookkeeping Tabs: Jobs, Contacts Mobile column: hidden Maree Quintrell's row: hidden Client, Job, Mobile, Hours, Status Bronwyn Tallis, Tax return Maree Quintrell, BAS lodgement, mobile 0491 570 156 Tobias Rendle, Bookkeeping Contacts sheet: not written Hidden row and column: both written
Invented test sheet, exported by LibreOffice 26.8 with comma separators and UTF-8, 29 September 2026

Fields nobody looks at

An export from business software is built from what the system stores, and the list view on screen is a selection from that. Before anything else, open the file and count the columns. A client export can carry a date of birth, a residential address or a bank account number that nobody on the team has looked at in months.

Payroll exports deserve their own caution. A column of tax file numbers is TFN information, and the Privacy (Tax File Number) Rule 2015 limits what can be done with it well beyond the general principles; the Word guide sets out that rule. If the question is about hours or totals, the TFN column has no business in the upload at all.

The column to read most carefully is the free text one, whatever it is called: Notes, Comments, Description. It is a small document tucked inside the table, and it is where staff write things like “spoke to her daughter, call back Tuesday” or a partner’s name and mobile. Deleting a Name column does nothing about a name typed into a sentence two columns along, so either drop the notes column or read every cell of it.

Encoding and commas: pick the right CSV

Excel offers more than one CSV in its Save as type list, and they differ in the character set they use to write letters. That sounds academic until a client called Renée or a supplier with a macron in its name turns up in the file, or a symbol such as a pound sign sits in a notes column.

The encoding decides whether a name arrives intact, not whether it should be there at all; the list of Australian identifiers to take out answers that second question.

Microsoft’s instructions for importing contacts into Outlook tell you to save with CSV UTF-8 (Comma delimited), and explain why: UTF-8 “works for all languages and alphabets”. A second Microsoft page, on opening those files in Excel, adds that a UTF-8 CSV opens normally “if it was saved with BOM”, the short marker at the start of the file that announces the encoding.

Save as type in ExcelEncodingWorth knowing
CSV UTF-8 (Comma delimited)UTF-8 with a byte order markthe one to use for anything leaving your desk
CSV (Comma delimited)an older single byte character setan accented letter becomes one byte other tools may misread
Unicode TextUTF-16, tab separated, saved as .txtkeeps every character; some programs do not expect it
Text (Tab delimited)older character set, tab separatedsame caution as the plain CSV

Option names from Microsoft Support (Australia). The encodings are the usual Windows behaviour; Microsoft’s page does not state them.

The difference is small and specific. The word Café takes 5 bytes in UTF-8, written in hexadecimal as 43 61 66 C3 A9, because the é needs two. In the older Windows character set it takes 4: 43 61 66 E9. A program that guesses wrong shows a stray symbol in the middle of a surname, and a tool that expects UTF-8 may refuse the file outright.

How to mask data in a CSV file: Café in UTF-8 and in the older Windows set The word Café written in UTF-8 takes five bytes, 43 61 66 C3 A9, because the accented e needs two. Written in the older Windows character set it takes four bytes, 43 61 66 E9. A program that reads the second file as UTF-8 cannot make sense of the last byte. CSV UTF-8 (Comma delimited) CSV (Comma delimited) C a f é 43 61 66 C3 A9 5 bytes, what UTF-8 readers expect C a f é 43 61 66 E9 4 bytes, garbled if read as UTF-8
Byte values of one word, as Python's standard codecs write them, 29 September 2026

Commas inside a value are the other trap. RFC 4180, the note that describes the CSV format, says a field containing a comma “should be enclosed in double-quotes”, so Quintrell, Maree travels inside quotation marks. Software that honours the quotes keeps it as one cell; software that does not splits one person into two. A list stored surname first is worth a look in a text editor before it goes anywhere.

Masking the columns, one step at a time

Masking in a CSV is plain editing, and the sequence counts for more than the tool. Do it in Excel, in LibreOffice Calc or in a text editor, but always on a copy, and save the original somewhere the upload cannot reach. Some columns go altogether, some stay as they are, and a few are replaced by something less revealing.

Drop what the question does not need

Start from the question, not the file. If you want the AI to total hours by job type, it needs the job type and the hours. The client’s name, mobile, email and address are not part of that sum, so delete those columns outright rather than blanking the cells, which leaves the header announcing what used to be there.

Rows follow the same logic. Filter the list down to the period or the work type the question is about, delete the rows outside it, and only then save. A shorter file is also a quicker one to check by eye at the end.

Codes instead of names, key kept elsewhere

Sometimes the answer has to point back to a person: which client has the overdue job, which tenant raised the repair. Replace the name with a code such as C01, C02, and keep the list of which code belongs to whom in a separate file that never goes near the upload. When the answer comes back, you look the codes up yourself.

Codes like these are pseudonyms, not anonymity. You still hold the key, so the rows are still personal information in your hands, and our guide to de-identified vs pseudonymised data explains which of the two words the Privacy Act uses and what it takes to earn it.

The file you upload

C01, Tax return, 4 hours, Open. C02, BAS lodgement, 2.5 hours, Waiting on client.

The file that stays on your desk

C01 is Maree Quintrell. C02 is Mirela Quennell. Never attached, never pasted.

Two files from one invented job list: the codes travel, the key does not

A hashed ID is still an ID

Some export tools offer to hash or encrypt an ID column instead of deleting it. That keeps rows linkable, which is the point, and it also keeps the column tied to a person. The OAIC’s investigation into the 2016 release of Medicare and PBS claims data records that university researchers reversed the encryption of provider numbers about a month after publication.

If the task does not need to link rows to the same person, delete the ID column. If it does, a fresh code per row set, with the key kept elsewhere, is safer than a hash anyone with the same method can recompute. The wider list of what counts as an identifier is in our guide to personal identifiers to remove before AI.

.txt and .md files: nothing hides, and that cuts both ways

Nothing in a .txt or .md sits underneath anything else. There are no comments, no tracked changes, no properties panel and no formatting that can make a word invisible. What Notepad or TextEdit shows you is every character in the file, so checking it takes one read, and a clean up you can see is a clean up you can believe.

The flip side is that nothing is filed away either. A note pasted in from an email keeps the signature block with the mobile number in it, and a Markdown file of meeting minutes keeps the attendee list at the top. Read the whole thing, top to bottom, before it goes. Progress notes kept as plain text are a case in point, and what to take out of an NDIS case note goes through them line by line.

A text export from Writer is not a clean copy

It is tempting to turn a document into plain text to strip it back. Apple’s TextEdit guide says that Format, Make Plain Text removes “all text styles and formatting”. Styles, yes. We tried the same idea in LibreOffice, saving a test file note from Writer as a .txt, to see what else went.

The comment, the header, the footer and the text of the footnote were gone. Two things survived: a sentence deleted with tracked changes still on, which had not been accepted, and a client’s name formatted as hidden text. The deleted sentence even ran straight into the next one, with no line break to show where it ended.

What a Writer file note kept and lost when saved as plain text Two columns. Dropped from the .txt: the comment and its author, the header naming the client, the footer naming the preparer, and the text of the footnote. Carried into the .txt: the sentence deleted with tracked changes, naming the referrer and a mobile number, joined to the next sentence, and a client's name formatted as hidden text. Left out of the .txt Carried into the .txt The comment and its author Header: File note, client's name Footer: Prepared by, preparer's name The footnote's text (its number stays) A sentence deleted with tracked changes on, never accepted: the referrer's name and mobile It runs into the next sentence A name formatted as hidden text, shown as ordinary text
Invented Writer file note saved as Text (encoded), UTF-8, by LibreOffice 26.8, 29 September 2026

So the order is the same as for any document: resolve each tracked change and clear out hidden text in Writer first, then export. The plain text copy is then a fair picture of what is left, and a quick one to read through.

LibreOffice Writer files: the parts no page shows

An .odt file is a zip archive with a few XML files and a picture inside. You can rename a copy to .zip and open it, which is the fastest way to see what a Writer document really carries. We did that with the same invented file note, saved by LibreOffice 26.8.

Part inside the .odtWhat our test note held there
content.xmlthe text, the comment and its author, the deleted referral sentence with a mobile, a name formatted as hidden, the footnote
styles.xmlthe header with the client’s name and the footer with the preparer’s
meta.xmlthe author and the document title, which named the client
Thumbnails/thumbnail.pnga picture of page one, struck through sentence, footnote and footer included

Invented file note, saved by LibreOffice 26.8 on 29 September 2026. The next sections use the same note.

The metadata in meta.xml is the kind of detail the metadata guide for Word and PDF deals with for Microsoft’s formats. Here the more pressing problem is everything in content.xml, because that is the text any program reads first.

Comments and their authors

A Writer comment sits in the document next to the paragraph it annotates, signed with the writer’s name and the date. In our test the comment said who had handled the client’s BAS the year before, a remark of just the sort that belongs inside the office and nowhere else.

LibreOffice’s help page for the Comments menu lists Delete All Comments, which “Deletes all comments of the document”. Use that on the copy rather than deleting comments one at a time, and search the document afterwards for the commenters’ names.

Tracked changes and hidden text

A deletion made with tracked changes on is not a deletion yet. It stays in the file, marked as removed, until someone accepts it. LibreOffice’s help on accepting or rejecting changes points to Edit, Track Changes, Manage, where the List tab has Accept All, which “Accepts all of the changes” in one go. Word behaves the same way under different menus, and the Word steps for accepting changes and deleting comments cover Windows, Mac and the web.

Hidden text is quieter still. The Font Effects help page describes the Hidden setting simply: it “Hides the selected characters”, and they only show when Formatting Marks is on in the View menu. Turn that on, search the document, and remove anything that does not belong, rather than trusting that a hidden name will stay hidden.

Headers, footers, footnotes and the preview picture

Headers and footers live in styles.xml, apart from the body, and are easy to forget when you edit on screen. Footnotes are stored with the body, though on paper they print below it. Check both, since a line such as Prepared by, followed by a name, is the sort of thing a template adds to every page.

The picture is the surprise. LibreOffice’s File, Properties, General tab has a Save preview image with this document option, which “Saves a thumbnail preview in PNG format inside the document”. In our file that picture showed page one as it looked on screen, including the struck through referral line. Untick it on the copy.

The privacy option that renames authors

One LibreOffice option looks tailor made for this job: Remove personal information on saving, which you will find at Tools, Options, LibreOffice, Security, Options. Its help page says it removes “user data from file properties, comments and tracked changes” and replaces author names “by generic values as ‘Author1’, ‘Author2’”.

We saved the test note with that option on. The properties were cleared and the commenter became Author1, which is what the help page promises. The comment’s words, the deleted sentence with its mobile number, the footnote, the header and the footer were all still there, and so was the preview picture.

What the option changed

Author names in comments and changes became Author1. Dates were reset. The document properties were emptied.

What it left in the file

The text of the comment, the deleted sentence and its mobile, the footnote, the header, the footer and the page picture.

Our invented Writer file note saved with Remove personal information on saving turned on, LibreOffice 26.8, 29 September 2026

It is a sensible default for anything you share, and worth leaving on. It is not a clean up. The same Security Options dialogue can also warn you when you save or send a document that “contains recorded changes, versions, or comments”, which is a better prompt at the moment it matters.

RTF: a document saved as Rich Text keeps its layers

An Australian office still meets .rtf files in everyday places: letters generated by a document management system, reports from an ageing accounting package, templates nobody has updated in years. People sometimes save a document as RTF on the assumption that the simpler format sheds the extras. It does not have to.

We saved the same file note as RTF from Writer and searched it as text. All 5 layers were there: the comment with its author’s name, the deleted referral sentence with the mobile number, the footnote, the header and the footer. Between them they named 5 invented people, and only 1, the client, appears in the main paragraphs once the changes are accepted. RTF is plain text underneath, so any program sees those layers as readily as the body.

5layers of the test note kept in the RTF
5invented people named somewhere in the file
1of them in the main paragraphs once changes are accepted
Invented Writer file note saved as RTF by LibreOffice 26.8, 29 September 2026. Word's RTF not tested

We have not run the same test through Word, so treat an RTF exported from Word with the same suspicion. The fix is upstream: clean the source document first, with the steps in our Word guide or the Writer steps above, then save it as RTF. If a system hands you an RTF you did not write, open it in Writer, settle every tracked change, clear out the comments and save a fresh copy.

Open the finished file the way a machine would

Whatever the format, the last check is the same: look at the file as a program would, not as the application you edited it in presents it. For a CSV, a .txt or a .md file, that simply means opening it in a text editor. For an .odt, rename a copy to .zip and look at content.xml and styles.xml; for an RTF, open the file in a text editor and search.

Then search for names you know were in the original. Pick two or three surnames, a mobile number and part of an email address, and search the cleaned copy for each. A hit means something was missed. A clean result, on a file you have also read end to end, is about as thorough as a small office can get without dedicated tools.

CheckCSV, .txt, .md.odt.rtf
Open as textany text editorrename a copy to .zip, open content.xml and styles.xmlany text editor
Search forsurnames, mobiles, email fragmentsthe same, plus author namesthe same, plus author names
Also confirmcolumn count, row countno preview picture, no commentsno comment groups left

Where an upload has already happened, note what went, where it went and when, and work from the copy that was actually sent. Whether that upload is a notifiable data breach turns on its own test, and our guide on client data in ChatGPT and data breaches sets it out. For what the provider keeps, see the guide to what ChatGPT saves.

Exports are copies, and APP 11 follows every copy

Each export is a new copy of personal information, and the Privacy Act does not treat copies as less important than the system they came from. The OAIC’s APP guidelines say that once information is no longer needed and has to go, an organisation must deal with “all copies it holds of that personal information, including copies that have been archived or are held as back-ups”.

Australia has a clear example of a copy outliving its purpose. On 5 September 2016 an employee of a contractor to the Australian Red Cross Blood Service saved a backup of a database file from a testing environment to a publicly accessible web server. It held registration details for about 550,000 prospective blood donors.

550,000
prospective donors in a backup file saved to a public web server. OAIC, DonateBlood.com.au data breach report, 7 August 2017

An individual scanning for vulnerabilities found and copied the file on 25 October 2016. The OAIC’s report, published on 7 August 2017, found the Blood Service had breached APP 11.1, for the lack of contractual and other measures over its contractor, and APP 11.2, for keeping data on the website longer than it needed to.

Your CSV of 40 clients is not a donor database, but the pattern scales down. The working copy you cleaned for one question, the original export beside it and the key file all need a home and an end date. Delete the ones you no longer need once the answer is in. Accounting practices have a page of their own on client files and AI.

Nonimo with a job list exported as CSV

Nonimo works from the Mac’s Finder or from File Explorer in Windows: pick the CSV there and use the Nonimo key. The clipboard receives the text of the file, each name, mobile number and email address swapped for a label, and that text is what the chatbot gets. When you paste its reply, the labels turn back into the real names on your machine. Drop the file on the Nonimo window instead and you can also download a covered copy of the CSV.

It reads CSV and TSV written in UTF-8, BOM or no BOM, or in UTF-16 that starts with a BOM, and plain text and Markdown files the same way; from Excel, save with CSV UTF-8 (Comma delimited). From a Writer .odt it takes the body, tables, headers and footers. Give the notes column a quick read before you paste.

An invented job list, three rows, exported with commas:

BEFORE (the CSV as saved, UTF-8 with commas)
Client,Mobile,Email,Job,Hours,Status
Maree Quintrell,0491 570 156,,Tax return,4,Open
Mirela Quennell,,mirela.quennell@example.com,BAS lodgement,2.5,Waiting on client
Tobias Rendle,,,Bookkeeping,1.5,Closed

AFTER (Nonimo, what goes on the clipboard)
Client,Mobile,Email,Job,Hours,Status
[PERSON_1],[PHONE_1],,Tax return,4,Open
[PERSON_2],,[EMAIL_1],BAS lodgement,2.5,Waiting on client
[PERSON_3],,,Bookkeeping,1.5,Closed

Jobs, hours and status are left alone, so the model can still count and sort. Only your computer knows which label is which client, so what the chatbot sees is pseudonymised, not anonymous. Our security page lists what Nonimo keeps locally.

Sources

Nonimo is the software that does this on your own computer: it masks client names and IDs before your text reaches ChatGPT. No account needed, and the app does it without your files leaving your machine.

Common questions

How do you mask data in a CSV file?

Work on a copy. Take out every column the question does not rely on, replace the names that must stay with codes and store the code list in a separate file, then save with Excel's UTF-8 CSV option and open the result in a text editor to search for a few surnames. Nothing can hide in a CSV, so what the editor shows is what the AI receives. Nonimo can then swap what is left for labels on the clipboard.

Does saving a spreadsheet as CSV remove hidden rows?

Not in our test. We hid a row and a column in a LibreOffice Calc sheet and exported it to CSV: both were written into the file, with the client's name and mobile. What a CSV drops is the other sheets, since Microsoft says each text format saves only the active sheet. Unhide everything and delete what should not go before you export.

Which CSV option should I pick in Excel?

CSV UTF-8 (Comma delimited), the one Microsoft tells you to use on its page about importing contacts into Outlook, because UTF-8 copes with every language and alphabet. The plain CSV (Comma delimited) option usually saves in a legacy Windows encoding, so a name with an accent can arrive garbled in another program. Nonimo reads UTF-8, or UTF-16 with a byte order mark, so saving with CSV UTF-8 gets every name through intact.

Is replacing names with a hash enough for a CSV export?

Usually not. A hash or an encrypted ID that stays the same for the same person is still a pseudonym, and the rest of the row can point back to them. In the MBS and PBS release the OAIC investigated, researchers found a way to reverse the encryption of provider numbers about a month after publication. Delete the column if the task does not need it, and keep any lookup key out of the file.

Do comments survive when a Writer file is saved as .odt?

Yes. In our test file the comment and its author's name sat inside the document next to the paragraph it was attached to. Turning on LibreOffice's Remove personal information on saving replaced the author with Author1, but the words of the comment stayed. To take them out, use Delete All Comments from the Comments menu, then save the copy you will upload.

Does saving as plain text clean a Writer document?

Only partly. When LibreOffice saved our test note as .txt, the comment, header, footer and footnote text dropped out, but a deletion still pending in tracked changes and a name formatted as hidden came along. Settle every tracked change and remove hidden text first; then the plain text copy is a sound thing to check and share.

Does a Rich Text file keep comments and revisions?

It can. An RTF that LibreOffice wrote from our test note still held the comment with its author, the deleted sentence, the footnote, the header and the footer. We have not run the same test through Word. Treat an RTF like the document it came from: clear every revision and every comment in the source, save again, then search the file in a text editor.

What comes out of Nonimo when you use it on a CSV?

Text to paste, and a covered copy of the CSV. Drop the file on the Nonimo window and the rows come back with names, mobile numbers and email addresses replaced by labels, ready for the chatbot, and you can download a covered copy of the CSV itself. When you paste the reply back, your real details return on your machine. It reads CSV and TSV saved as UTF-8, or UTF-16 with a byte order mark.