[nonimo]
EN
Download

Remove personal data from a CSV before AI gets it

· Written and maintained by Nonimo

To mask data in a CSV file before an AI reads it, cut the export down to the columns your question needs, keep it in UTF-8 encoding, swap every name and identifier for a code while the code list stays in a separate file, and look through the result in a plain text editor such as Notepad before uploading. A CSV hides very little, which makes it the easiest of these formats to clean.

Plain text files are just as honest. Documents written in LibreOffice Writer, and RTF files, are not: in our test, both kept comments, a deleted line and hidden passages that the page never showed. This guide covers those four formats. Workbooks saved as .xlsx have a guide of their own, and so do slide decks and Word documents.

How to mask data in a CSV file, step by step

Work on a copy of the export, never on the file your CRM, practice management or payroll system keeps writing to. Knowing how to mask data in a CSV file is mostly knowing what to leave out of it, so the order matters: each step makes the next one smaller.

  1. Pick fields at export time. CRMs and payroll tools usually ask which fields to include. If the question is how plan choices vary by age band, the file needs the plan and the band, not the member’s name, email or SSN. A column that never leaves the system needs no masking.
  2. Save the copy as CSV UTF-8. In Excel that is File > Save As, choosing CSV UTF-8 (Comma delimited) as the type. In LibreOffice Calc, choose Text CSV, tick Edit filter settings and set the character set to Unicode (UTF-8).
  3. Delete columns, not values. A blanked cell still announces that a value used to be there, and a column header like Diagnosis tells the reader what kind of file this is even when every row is empty.
  4. Replace names and IDs with codes. Give each person a code such as P001, use the same code on every row where that person appears, and store the code list away from the export, since it is the one thing that turns P001 back into a name.
  5. Coarsen what points at one person. A full birth date becomes a year, such as 1981, a ZIP code keeps its first three digits, and a job title as precise as night charge nurse becomes nursing.
  6. Read the free text columns. Notes, comments and descriptions carry names inside sentences, and no column rule reaches them.
  7. Open the copy in a text editor. What you see there is what the AI receives: every row, every column, no formatting to hide behind.
How to mask data in a CSV file: the masked copy and the key, in two files On the left, the masked CSV that goes to the AI, with the columns Code, Plan and Birth year and two rows, P001 and P002. On the right, a separate key file that stays behind, mapping P001 and P002 to the two invented names. A dashed line between them marks that the key never travels with the copy. The copy that goes to the AI The key that stays behind Code, Plan, Birth year P001, PPO, 1981 P002, HSA, 1981 Code, Name P001, Corliss Vandegrift P002, Dolores Vantreight Invented rows. The key sits in another folder, with tighter access.
Steps 4 and 5 in one picture: codes and birth years in the copy, names only in the key

The sections below explain why each step is there. If you are not sure which fields count as personal in the first place, the US checklist of personal information is the place to start.

What a CSV export carries, and what it cannot

A CSV is a text file with one record per line and a separator between fields, usually a comma. It has no sheets, no cell comments, no formulas and no formatting, so the tricks that hide data inside a workbook have nowhere to live. Microsoft’s help for saving a workbook as text says each CSV type saves only the active sheet, and that all formatting is removed.

That same simplicity cuts the other way. Everything the export wrote is readable by any program that opens the file, including the columns and rows the person who ran the export never looked at. Unlike a Word file or a PDF, a CSV also has no author field or company name tucked into its properties; removing metadata from PDFs, Word files and photos is a job for those formats.

In the workbook or systemIn the CSV
Other sheetsNo: only the active sheet is saved
Cell comments and notesNo: the format has no place for them
Formatting, white text includedThe text stays, the formatting goes
A formula, in our Calc testIts result, written as text
A row hidden in LibreOffice CalcYes, as an ordinary row (our test)
Every field the export selectedYes, all of them

Columns the report never showed

A report on screen shows the columns someone chose. An export often writes the whole record: internal IDs, the owner’s email, date created, last activity, a free text field nobody reads. Before masking anything, list the header row and ask of each column whether the AI needs it. For a health or benefits file, the difference between PHI and PII decides how careful the rest of the job has to be.

Rows you filtered or hid

A filter or a hidden row changes what a spreadsheet shows, not what it holds. In a test we ran with LibreOffice Calc 26.8, a row hidden in the sheet was written into the CSV like any other row, with the person’s name, number and email. If the source is a workbook, clean it there first with the Excel steps, then export.

Choose CSV UTF-8, because the encoding decides what a search finds

A CSV is only characters, and the encoding decides which bytes stand for which character. Excel offers several CSV types. Microsoft’s guidance for moving contacts through a CSV says to save with UTF-8 encoding when you have that option, and the Excel step it gives is File > Save As, then the CSV UTF-8 (Comma delimited) entry in the type list. That is the one to choose for anything you are about to mask, for a practical reason.

José is 4A 6F 73 E9 in Windows-1252, the encoding many older Windows tools use on US systems, and 4A 6F 73 C3 A9 in UTF-8. A search in one encoding can miss a row written in the other, and software built for UTF-8 may show broken characters or refuse the file. You then believe every copy of a name is gone when one is not, and a detail you meant to take out rides along into the upload.

How to mask data in a CSV file: the same name in two encodings The name José written as bytes. In Windows-1252 it takes four bytes, 4A 6F 73 E9. In UTF-8 it takes five, 4A 6F 73 C3 A9. The first three bytes match; the accented letter does not. The name José, byte by byte Windows-1252 UTF-8 4A6F73E9 4A6F73C3A9 José
José in two encodings: four bytes against five. Written with Python's cp1252 and utf-8 codecs, September 29, 2026

In LibreOffice Calc the choice sits in the export dialog: pick Text CSV, tick Edit filter settings, and set Character set to Unicode (UTF-8). If you later reopen your masked copy in Excel, Microsoft notes that Excel reads a UTF-8 CSV cleanly when the file begins with a byte order mark (BOM); without one, bring it in through Data > Get Data > From Text/CSV so the accents survive.

Quoted fields and names written last name first

Many exports write a person as Vandegrift, Corliss, and because the comma is also the separator, the field goes inside double quotes. Two things follow when you mask by hand.

A search for Corliss Vandegrift never matches that cell, since the order is reversed, so search for the surname on its own. And if you replace the field, keep the quotes balanced: a code dropped in with one quote left behind can merge or shift the columns after it, and the SSN ends up under the wrong header.

The cleaner route is to split the name into two columns before masking, or to replace the whole quoted field, quotes included, with a plain code. If commas inside fields keep causing trouble, Excel’s Text (Tab delimited) type, listed in Microsoft’s import and export help, puts a tab between fields instead and sidesteps the quoting. Reopen the result afterwards and check that every row has as many fields as the header.

How to mask data in a CSV file with codes and a key kept apart

Replacing Dolores Vantreight with P002 on every row, and keeping the list that turns P002 back into a name in another file, has a name in California law. The state’s privacy statute defines pseudonymization, in Civil Code 1798.140(aa), as processing that leaves personal information “no longer attributable to a specific consumer” without additional information, as long as that information “is kept separately” under safeguards.

That is exactly the code and key method, and it is also why the result is not deidentified data. Subdivision (m) of the same section asks for more: reasonable measures against reidentification, a public commitment not to try, and contracts binding whoever receives the data. The line between deidentified and anonymized data is drawn in more detail elsewhere.

Two practical rules keep the codes useful. Use one code per person across every row and every file you send in the same conversation, or the AI cannot tell that rows 4 and 19 are the same patient. And never build the code from the person’s own data: the last four digits of an SSN or a birth date plus initials can be reversed by anyone holding the original.

HIPAA safe harbor, column by column

Among US rules, only HIPAA’s safe harbor names identifiers field by field, which makes it map neatly onto columns. It binds covered entities, and business associates acting for them, when they handle protected health information, but its list is a sound checklist for any export. Where it applies, it also demands that you have no actual knowledge the remaining data could identify someone.

Column in the exportSafe harbor item in 164.514(b)(2)(i)What to do in the CSV
Name(A) namesReplace with a code, key kept apart
Street, city, ZIP(B) geographic subdivisions smaller than a stateDrop street and city; keep three ZIP digits only if that area has over 20,000 people
Birth or service date(C) dates except the yearKeep the year; ages over 89 become one 90 or older group
Phone and email(D) and (F)Delete the column
SSN(G)Delete it, and never reuse its digits in a code
Member or account number(I) and (J)Replace with a code
Your own case or claim number(R) any other unique numberReplace with a code

The regulation also says, in 164.514(c), that a reidentification code must not be derived from information about the person. For a covered practice deciding whether any of this may reach an AI tool at all, our note on HIPAA and AI assistants comes first, and the healthcare page shows where Nonimo fits.

Notes and other free text columns

A notes column defeats every rule above. “Called Mrs. Vantreight, her son Glenn will bring the forms” puts two people in one cell, and neither sits under a Name header. Read these columns line by line, or delete them if the question does not need them.

Search the copy for the telltales: @ for emails, titles such as Mr. and Mrs., area codes you recognize, and the surnames you just replaced elsewhere. A surname that survives in a note undoes the code on every other row.

Plain text and Markdown files

A .txt file has no hidden layer. What your editor shows is the file, give or take its encoding, and the same goes for Markdown, where a table is just rows of text between pipes. These files arrive from call logs, transcripts, exports of case notes and emails pasted into a document. Their risk is volume and free text, not concealment.

Mask a Markdown table the way you mask a CSV, column by column. For prose, search and replace each name with its code, then search again for the surname alone and the first name alone, since a transcript seldom spells a person out the same way each time. If the text is going to ChatGPT, what OpenAI does with it once it arrives is worth reading too.

Encoding matters here as well. Microsoft retired WordPad from Windows 11, version 24H2, and its note on the change recommends Windows Notepad for plain text documents like .txt and Microsoft Word for .rtf. Whichever editor you use, save the masked copy as UTF-8 if it asks.

Call logs and transcripts

Text exported from a phone system or a meeting recorder has its own shape: a timestamp, a speaker label, then what was said. The speaker label is the easy part, because it repeats; replace the agent’s name at the start of each line with Agent or with a code.

What the caller says is harder. People read out a date of birth to pass security questions, spell an email address aloud and mention a spouse by first name, and none of that follows a label. Read those files for what was said, or cut the transcript down to the part your question is about before any masking starts.

How to redact in LibreOffice Writer when the file goes to AI

A Writer document is not one stream of text. An .odt is really a zip file holding several XML parts: content.xml holds the body, the tables, the comments and the tracked changes; styles.xml holds the page header and footer; meta.xml holds the document properties. Any program that opens the zip can read all three, whatever Writer shows on screen.

How to redact in LibreOffice Writer: the memo on screen against the .odt archive On the left, the memo as Writer shows it: a heading, a client line and a small table. On the right, the same .odt opened as a zip archive with three parts. content.xml holds the body and table, the comment and who wrote it, the tracked deletion and who made it, hidden text, a hidden section and a footnote. styles.xml holds the header and footer. meta.xml holds the author in the properties. What Writer shows What the .odt keeps Open enrollment memo Client line and a table one member, one row content.xml body, table, comment and its author, tracked deletion and who made it, hidden text, hidden section, footnote styles.xml header and footer meta.xml author in properties
Our invented Writer memo, saved by LibreOffice 26.8 and unzipped and read by script, September 29, 2026

To see it for ourselves, we wrote a one page benefits memo with twelve items planted where a reader would not look, had LibreOffice 26.8 save it as .odt, as .rtf and as plain text, and then read each file with a script instead of an office program. The .odt kept all twelve. The table below groups the 12 into seven rows, with the other two formats alongside.

Where it sat in the memo.odt.rtfSaved as Text
Body text and the tableyesyesyes
Page header, footer and a footnoteyesyesno
A comment, with who wrote ityesyesno
A tracked deletionyesyesyes
Who made the deletionyesyesno
Hidden text, and a section set to hideyesyesyes
The author in the propertiesyesyesno

Comments and tracked changes

A comment in Writer stores its text together with the author and the date. The Edit menu has Comment > Delete All Comments for the whole document, and Delete Comment By to clear one reviewer at a time. Hiding the comments margin does not delete anything. Word draws the same line between hiding and deleting, for comments and for tracked changes in a .docx.

Tracked changes are the bigger surprise. In our memo, a line naming a previous contact and a phone number had been deleted with change tracking on; the page no longer showed it, yet content.xml still held the whole line and the name of the colleague who deleted it. Open Edit > Track Changes > Manage and accept or reject the changes, all together or one by one, before the file goes anywhere.

Hidden text, hidden sections and footnotes

Writer can hide text under a condition, a paragraph or a whole section, and LibreOffice Help documents each through Insert > Field > More Fields and Insert > Section. Whether hidden paragraphs appear on screen depends on the Hidden paragraphs setting under Tools > Options > LibreOffice Writer > View. Delete what you hid rather than trusting the switch.

Footnotes, the header and the footer are visible in print layout but easy to skip when you are scanning the body for names. Page through them deliberately; in our memo, the footnote carried the name and email of the person who referred the client.

Remove personal information on saving renames, it does not delete

One LibreOffice option looks like the whole cure: Tools > Options > LibreOffice > Security, then Options, then Remove personal information on saving. Here is what LibreOffice Help says it does on save, next to what that description leaves alone:

What the setting does on saveWhat stays in the file
Comment and change authors become Author1, Author2The text of every comment
Their times reset to one standard valueThe line a tracked deletion removed
User data leaves the document propertiesHidden text, sections and footnotes

So the setting renames, and the renaming is worth having: a comment signed Author2 no longer tells the assistant which colleague doubted the claim. But a comment saying to call John back about his wife’s surgery reads the same under any author name, and the deleted line still carries the old contact’s phone number.

Use the setting as well as deleting comments and accepting changes, never instead. Those property fields are cousins of the ones in a Word file or a PDF.

Tools > Redact gives you a PDF of pixels

Writer’s redaction command, Tools > Redact, exports the document to LibreOffice Draw, where you draw boxes over what must go. LibreOffice Help says the covered content is removed and replaced by a block of pixels, and the result is exported as a PDF. It suits a finished document you are releasing, and our PDF guide covers the AI side of that route. For a draft you want an assistant to rewrite, you usually want the text itself, masked.

Save as plain text is not a clean up

Exporting a Writer document to .txt feels like a clean start, and it is half of one. In our test, File > Save As with the Text format lost the comment, the footnote, the document properties, and the header and footer text. It kept the body and the table, and it also kept the deleted line, the hidden sentence and the hidden section. Five of the 12 planted items survived.

5 of 12
planted items that survived Save as Text in our Writer memo, including a deleted line and two hidden passages. Our test, September 29, 2026

The deleted line came out glued to the sentence that replaced it, with no space between the old phone number and the new text, so it is easy to read past. Settle every tracked change, delete hidden passages, and only then export to text. Read that export before uploading, since it shows you everything that is left.

For a legal file, what redaction does for privilege is answered in a separate guide.

RTF still carries comments and revisions

RTF predates both .docx and .odt, and Word and LibreOffice both still open and save it. Microsoft’s note on retiring WordPad recommends Word for rich text documents like .rtf, so that is where most of these files now open.

RTF has room for everything that makes a document risky. The .rtf that LibreOffice saved from our memo kept all 12 planted items, with the comment stored under its own annotation code next to the author’s name, and the deleted line marked as a deletion rather than removed. Read it as plain text and the names sit right there between the formatting codes.

Three lines of our memo.rtf read as plain text, shortened, date codes left out
{\*\atnauthor Wendell Hesterberg}\chatn{\*\annotation ... Check the dependent form before Friday}
{\*\revtbl {Unknown;}{Priya Vandermolen;}}
{\deleted\revauthdel1 ... Previous contact: John Doe, 614-555-0182}{Current contact: the client herself.}

The fix is the one you would use for a .docx: open the file in Word, accept or reject all changes, delete all comments, and look over what the header and footer say. Our Word guide has the exact menus. Then save the cleaned copy, and read it once more as plain text.

Open the file as plain text before you upload it

Assistants differ in what they accept. On September 29, 2026, Claude’s upload page named ODT and RTF next to CSV, TXT, PDF and DOCX. Microsoft’s list for Copilot with a work or school license included .csv, .txt, .rtf and .md, and did not name .odt. Both lists change, and neither says which parts of a file the assistant reads, so assume it reads everything that is in the file.

The last check is the cheapest. Open the masked CSV, the .txt, or a text export of your cleaned document in Notepad or TextEdit, and read it top to bottom. Search for @, for three digits, a dash and two more, for the surnames you replaced, and for titles like Dr. and Mrs. If you find something, fix the source and export again.

Then keep that masked copy, dated, in the folder for the matter or the client. If a colleague, a client, an auditor or a city’s records officer later asks what went to the assistant, you can open the exact file instead of reconstructing it from memory. Keep the code list somewhere else, with tighter access, because whoever holds both files holds the original data again.

What happens to the file after upload depends on the provider and the plan: Claude’s training settings and Copilot’s each have a guide of their own.

Nonimo with a CSV export

Nonimo is installed on the Mac or Windows PC in front of you. Highlight the export in the Finder or File Explorer, then use the Nonimo shortcut. What lands on your clipboard is the text of that file, each name and identifier it recognizes swapped for a placeholder, so it can go straight into the chat. Once the assistant answers, paste its reply and the original names come back. The result is text, you can download a masked copy of it, and the export itself stays untouched.

It opens CSV, TSV and .txt files encoded in UTF-8, and UTF-16 ones that begin with a BOM, which is one more reason for the CSV UTF-8 step above. From .odt and .rtf files, Nonimo collects the main text, the tables, and the page header and footer. Here is a two row export we invented, saved as CSV UTF-8, followed by the clipboard text from Nonimo:

The CSV UTF-8 file
Name,Email,Mobile,SSN,Plan
Corliss Vandegrift,,937-555-0144,000-47-2215,PPO
Dolores Vantreight,dvantreight@example.com,,,HSA

Clipboard text from Nonimo
Name,Email,Mobile,SSN,Plan
[PERSON_1],,[PHONE_1],[REFERENCE_1],PPO
[PERSON_2],[EMAIL_1],,,HSA

Each placeholder is a pseudonym that Nonimo can turn back into the name, and the plan column stays readable for the AI.

Our security page lists what stays stored on your machine, and the license page lifts the monthly word limit. Brokers and accountants who handle client lists daily can see the same step on the insurance brokers page and the CPA firms page.

Sources

Nonimo is the software that does this on your own computer: it masks client names and IDs before your text reaches ChatGPT. No account needed, and the app does it without your files leaving your machine.

Common questions

How do I mask data in a CSV file before uploading it to an AI?

Here is how to mask data in a CSV file in one pass. Export only the columns you need, keep the copy in UTF-8, and swap each name and account number for a code such as P001 listed in a separate key. Cut birth dates to the year and ZIP codes to three digits, then read the copy in a text editor. Nonimo can do the swapping from the clipboard and restores the names when you paste the answer.

Does a CSV keep hidden rows or comments like a workbook does?

Comments, no; hidden rows, possibly. A CSV has no second sheet, no cell comments and no formatting, so nothing sits behind white text or a collapsed group. What it does have is every column and row the export wrote, including rows someone had hidden in the spreadsheet: LibreOffice Calc wrote a hidden row into the CSV in our test. Read it in Notepad or TextEdit and nothing is held back.

Which Excel save type keeps a CSV ready for masking?

The entry called CSV UTF-8 (Comma delimited) in the Save As list. When Microsoft explains moving contacts through a CSV, it asks for UTF-8 encoding and points to that same entry. UTF-8 keeps names such as José or Zoë intact, which matters when you search the file for every copy of a name. Nonimo opens CSV and .txt files encoded in UTF-8, and UTF-16 ones that begin with a BOM.

Do codes in place of names satisfy HIPAA?

On its own, no. The HIPAA safe harbor, at section 164.514(b)(2) of the Privacy Rule, lists eighteen kinds of identifier, from names and ZIP codes to phone numbers and any other unique number, and also requires that you have no actual knowledge the rest could identify someone. A code is allowed only if it is not derived from the person's own data. Nonimo's placeholders are reversible pseudonyms, so you get the real names back in the answer.

Does saving a Writer document as plain text clean it?

Only partly. In our test, a LibreOffice memo saved as Text lost the comment, the footnote, the document properties, and the header and footer text, but kept a deleted line, a hidden sentence and a hidden section, with the deleted line glued to the next one. Settle the tracked changes and delete hidden passages first. In an .odt, Nonimo takes in the main text and tables plus the page header and footer.

Can Claude or Copilot read ODT and RTF files?

On September 29, 2026, Claude's upload page named ODT and RTF next to CSV, TXT, PDF and DOCX. Microsoft's list for Copilot with a work or school license included .csv, .txt, .rtf and .md, but not .odt. Lists change, so check the provider's page on the day. Nonimo works on any of these before upload by putting the file's text, with placeholders, on your clipboard.

Is LibreOffice's personal information setting enough on its own?

No. When you save, it gives the authors of comments and tracked changes generic labels like Author1, sets their times to one standard value, and clears user data out of the properties. According to LibreOffice Help, that is the whole job: the comment text and the tracked deletion stay in the file. Delete comments and accept or reject changes as well. Nonimo is a separate step for the text you paste.

Can Nonimo mask a CSV export for me?

Yes. Highlight the .csv in the Finder or File Explorer and use the Nonimo shortcut; your clipboard receives the rows with each name and identifier it recognizes turned into a placeholder. That text goes into ChatGPT, Claude or Copilot, and pasting the assistant's answer restores the real names. You can also download a masked copy as text, and the export on disk stays untouched.