[nonimo]
EN
Download

Clear a CSV file before it goes to AI

· Written and maintained by Nonimo

ChatGPT accepts a CSV upload, and so does Claude. The catch is that the file goes up whole. A CSV exported from a client system can hold thirty columns when the question needs three, and every one of them is readable by whatever opens the upload.

So the work happens before the CSV goes up to ChatGPT: delete the columns the task does not use, replace names with codes, save in UTF-8, and read the raw file in a text editor. The same applies to the other plain formats covered here: text files, LibreOffice Writer documents and RTF.

This guide is for the Irish office that exports rather than attaches: the bookkeeper with a bank feed, the practice manager with a list of returns, the HR lead with a report from the payroll system. Whether client data belongs in a chatbot in the first place is taken up in a separate guide on client data and ChatGPT.

Workbooks and slide decks have traps of their own, covered in the Excel guide and the PowerPoint guide. For Word files there is a separate method, and what to remove before AI runs through the Irish identifiers worth taking out of any file.

Can I upload a CSV to ChatGPT? It will read every column

A CSV is the simplest file an office produces. It is text, one record per line, with commas between the values, and almost every system that holds data can export one. That is exactly why people reach for it when they want an AI tool to find a pattern: overdue returns, repeat customers, a spike in refunds.

The simplicity cuts both ways. There is no formatting to hide anything, but there is also no way to see at a glance how much the file holds. An export button rarely asks which columns you want. It gives you all of them.

What the providers say about CSV files

OpenAI’s File Uploads FAQ, as archived on 18 September 2026, gives analysing a spreadsheet such as a CSV as one of the jobs uploads were built for, and puts the limit for CSV files at about 50MB.

The same page says uploaded files stay in your account for as long as the chat that holds them, and are deleted within 30 days of deleting that chat. How ChatGPT’s retention works in more detail is in our guide on what ChatGPT keeps.

Anthropic’s page on uploading files to Claude, dated 23 July 2026, says Claude “can work with the following document types”, and the list that follows takes in every format covered here, from CSV and TXT to ODT and RTF. Whether Claude learns from what you give it is answered separately, in our guide on Claude and training.

Neither page changes what matters here. Once the file is uploaded, the provider holds every value in it, under its terms rather than yours. Deleting the chat later is a separate job, and in ChatGPT a file saved to Library has to be deleted on its own.

What survives when a spreadsheet becomes a CSV

Guidance from the DPC on redacting records, revised in August 2021, treats plain text and CSV as the simplest formats a record can take, and suggests exporting to them because that strips out concealed tables, cross references and some metadata. It is good advice, and it explains why a CSV is cleaner than the workbook it came from.

Cleaner is not the same as clean. Here is what happens to the parts of an Excel workbook when you save it as CSV:

In the workbookIn the CSV
Other worksheetsgone: Excel writes only the active sheet
Formulasonly the value each one shows
Comments, notes, headersno place for them in the format
Every column of that sheetall there, in order
Rows hidden on that sheetcheck the file; do not assume

Sources: the DPC’s redaction guidance of August 2021; Microsoft Support’s list of the file formats Excel saves.

Of the five rows, the fourth carries the weight. A CSV strips away the hiding places and leaves the data in plain view, so everything depends on which columns were in the sheet when you saved it.

That is why the workbook needs clearing before the export, not after: unhide what is hidden, delete the columns that should not travel, then save. The order for doing that in Excel is in the Excel guide, and the list of Irish identifiers to look for, column by column, is in our guide to what to take out.

When the CSV comes from somewhere else

Most CSVs in an office never passed through Excel. They are produced by the software the office already runs: the CRM, the payroll package, an online booking form, the bank’s statement download, the system that tracks client returns.

Each of those decides its own columns, and they tend to include fields nobody on the team looks at day to day: a free text notes field, the name of the staff member who created the record, a date of birth kept for identity checks, an IP address from a web form.

A bank statement export is a good example. The line describing each payment often carries the name of the person or business on the other side, so a file uploaded to sort spending into categories can also hand over a list of names.

Open the CSV in Notepad, not Excel

Double clicking a CSV on an office computer usually opens it in Excel, which is the worst way to check it. Excel parses the file into a grid, sizes the columns, and shows you a tidy table. It does not show you the file.

Open it instead in a plain text editor, TextEdit if you are on a Mac and Notepad if you are on Windows, and you see what an AI tool receives: the raw lines, the commas, the header row with every column name spelled out. Scroll to the right end of the first line and count the headings. That number, not the width of the grid on screen, is how many columns the upload carries.

Commas inside a name, and the tab alternative

A comma separates the values, so a value that contains a comma has to be wrapped in quotation marks. Irish names written surname first, “Ó Beacháin, Diarmuid”, are the everyday case, and so are addresses and company names with a comma in them. A program that splits on commas without honouring the quotes reads one name as two fields and shifts the row.

Tab delimited text sidesteps this, because nobody types a tab inside a name. LibreOffice Calc lets you choose the field delimiter and the character set when it exports text, which is the simplest way to get a tab delimited file in UTF-8.

Why a round trip through Excel changes the file

Opening an export in Excel to tidy it and saving it again is not neutral. Microsoft’s page on leading zeros says Excel removes them automatically and turns long numbers into scientific notation, so a mobile stored as digits loses its first 0 and a long account reference becomes something like 1.23E+15.

Newer versions let you switch those conversions off, but the safer habit is simpler: make the cuts at the source, or in the workbook before export, and use a text editor only to look.

If a column has to go from a CSV you already have, delete it in LibreOffice Calc or in Excel on a copy, and compare the result with the original in Notepad. How a PPS number and an Eircode should be handled once you find them is set out in the PPSN guide.

Fadas, encodings and the CSV UTF-8 choice

Every character in a text file is stored as a number, and the encoding is the table that says which number means which letter. Microsoft’s page on text encoding gives the classic failure: the same stored value shows as one letter in one encoding and a different letter in another. Irish names meet this constantly, because the síneadh fada sits outside the basic set of letters.

In UTF-8 a capital Ó is stored as two bytes, C3 93; in the older Western European (Windows) set it is a single byte, D3. Each goes wrong in its own way when a program guesses the other.

Upload a CSV to ChatGPT with fadas intact: how the letter Ó is stored in two encodings and what a mismatch shows In UTF-8 the letter Ó is stored as two bytes, C3 93. In Western European (Windows) it is one byte, D3. A file saved in the Windows encoding and read as UTF-8 shows a replacement mark instead of the letter; UTF-8 read as the Windows encoding shows two wrong characters. Saved as Bytes for Ó Read the other way UTF-8 C3 93 Diarmuid Ó Beacháin Western European (Windows) D3 Diarmuid � Beach�in Left column: how the file was saved. Right column: what a program expecting the other encoding shows. Invented name. Decoded with Python's standard library.
How one Irish name breaks when a CSV crosses encodings; test file built for this guide, 29 September 2026

For a CSV headed to an AI tool, the practical rule is short: save in UTF-8.

Excel’s Save As list offers a CSV UTF-8 entry beside the plain CSV (Comma delimited) one, and the UTF-8 entry is the one to pick. Microsoft’s own note on opening such files says Excel has no trouble with a UTF-8 CSV that carries a byte order mark, a short marker at the start that tells programs how to read the rest.

The Save As list, entry by entry

Excel offers several text formats, and they are easy to confuse. This is what each one is for, going by Microsoft’s list of the formats Excel saves, which leaves out the UTF-8 entry itself:

Save As entrySeparatorWhat to know
CSV UTF-8 (Comma delimited)commakeeps every character; the one to use
CSV (Comma delimited)commameant, going by Microsoft’s list, for other Windows systems
Text (Tab delimited)tabsaves only the active sheet
Unicode Textnot statedMicrosoft describes it only as Unicode text

Sources: Microsoft Support, file formats that are supported in Excel, and opening CSV UTF-8 files correctly in Excel.

Fadas cause a second, quieter problem when you search a file for a name, which the Word guide covers along with the vocative and the anglicised spelling of a name.

If the names in the export already look wrong when you open it in Notepad, fix the encoding before anything else. A model handed “Beach�in” has no way to know what the letter was, and a later search for the client’s name in its answer will come up empty.

Plain text and Markdown files: nothing hidden, everything visible

With a .txt, what the screen shows and what the file holds are the same thing. The DPC’s note describes plain text as holding only text and number values, with nothing more than tab stops and line breaks for layout, and that is the whole file. A .md file, the kind note apps and developers use, is the same thing with a few symbols for headings and tables.

That makes a text file easy to check and easy to get wrong. With nothing hidden, every problem is in plain sight, and the only thing standing between it and the upload is whether someone reads it.

Markdown tables deserve a second look. A note app that exports to .md writes each table as rows of text divided by vertical bars, so a client list kept in a note becomes, on export, a small CSV in all but name, with the same columns and the same need to trim them.

Email threads pasted into a .txt

The usual text file in an office is a paste: a chain of emails copied out of Outlook so an AI tool can summarise it, or notes from a call. Those carry the details people skim past: signature blocks listing a desk phone and a personal mobile, the quoted reply three levels down with the client’s address, a PPS number mentioned once in passing.

When Word saves a document as plain text, Microsoft’s page on encodings says it uses Unicode unless you pick something else, and explains that Unicode is what avoids garbled characters between systems. Leave it on Unicode, for the same reason as with a CSV.

Read the file from the first line to the last before it goes. It is the only check a text file needs, and it cannot be skipped. What the file properties of a Word document say about you, and why pasted text leaves them behind, is covered in the metadata guide.

LibreOffice Writer and the .odt file

LibreOffice is a free office suite, and its word processor, Writer, saves in the OpenDocument Text format, .odt. An .odt looks like a single document and behaves much like a Word file: it can hold comments, recorded changes, hidden text, headers and footers, and a record of who wrote it.

Under the name, an .odt is a zip package. The OpenDocument standard, published by OASIS, sets out the parts: content.xml for the text, styles.xml for page styles such as headers and footers, and meta.xml for the document’s properties, among them the person who first created it.

Inside a test .odt

We built a short draft letter from an invented accountancy practice to an invented client, the kind of file someone might hand to Claude with a request to tidy the wording. Its visible text is five short lines: a greeting, two sentences about a tax return, a closing line and a sign off. Underneath, it carried six more things, and a short Python script using only the standard library printed all of them:

A LibreOffice .odt before it goes to an AI tool: the words on the page and what a parser found in the package The page shows a short letter with no identifiers. Read as a file, it also holds a comment with its author and a PPS number, deleted text with an IBAN, hidden text with a mobile number, a header with the client's full name, a footer naming the preparer, and properties naming two members of staff. On the page Dear Diarmuid, Your Form 11 is ready to file. Kind regards, In the file as well comment by a colleague: a PPS number deleted sentence: the refund IBAN hidden text: a mobile number header: the client's full name footer: who prepared it properties: creator and last editor Invented data. The left side names nobody but the client's first name.
Test .odt assembled by hand for this page and parsed with Python's own zip and XML modules, 29 September 2026

Two of those deserve a closer look. The comment came out with its author’s name and the time it was written. The deleted sentence came out whole, because the standard says a deletion record may hold the content that was deleted while change tracking was on. Text struck out while changes are being recorded is still in the package, and only accepting the change removes it.

Writer’s help confirms the same from the other side: with recording on, deleted text stays visible but crossed out, and resting the pointer on a change shows its author and the time. For Word, how comments and changes behave in a .docx is set out in a guide of its own, and the logic carries over.

Clearing a document in Writer

The menus below are those of LibreOffice Writer with the classic menu bar, named as LibreOffice’s English help names them:

  1. Work on a duplicate. Save As, give it another file name, and leave the original untouched.
  2. Go to Edit, Track Changes, Manage, and press Accept All or Reject All. Nothing a colleague struck out is gone until you do.
  3. Go to Edit, Comment, Delete All Comments. Marking a comment Resolved only adds the word Resolved under it; the text stays.
  4. Show hidden paragraphs through View, Field Hidden Paragraphs, and check for hidden sections, then delete what should not travel.
  5. Check the header and footer on each page style.
  6. In File, Properties, press Reset Properties and untick Apply User Data. Under the Security options, Remove personal information on saving replaces the names attached to comments and recorded edits with placeholders.

That last option is worth knowing about and easy to misread. LibreOffice says it replaces the authors’ names with “Author1”, “Author2” and so on. The text of the comments stays exactly where it was.

A LibreOffice spreadsheet, the .ods, is a different case: save it as .xlsx and follow the Excel guide.

Writer’s Redact command, and where it fits

Writer has its own redaction tool, under Tools, Redact. LibreOffice’s help explains how it works: the document is exported to a drawing in LibreOffice Draw, you mark the areas to hide, and the finished result is usually a PDF in which the marked parts are replaced by blocks of pixels. The source document is not touched.

That is the right tool for publishing a document. It is the wrong one for an AI task, which usually needs text rather than a picture of a page. For text, clear the document with the steps above and then copy the words out, or save it as plain text. The PDF route brings different pitfalls, which the PDF guide walks through.

Rich Text Format, a Word file in plain letters

Rich Text Format predates .docx by years, yet offices keep receiving it, usually from older software that exports nothing else, or from someone who picked it because almost any program opens it. It is also, unusually, a format you can read in Notepad. Microsoft’s specification describes RTF as normally written in plain ASCII characters, with backslash commands that switch formatting on and off.

That makes an RTF file easy to inspect and easy to misjudge. Look at one in Notepad and, between the commands, the words of the document appear, along with some that never reach the printed page:

An RTF file opened in a text editor, with the commands that carry the author, a comment, deleted text and hidden text Five lines from a test RTF: the info group with author and company, a comment with its author, a deleted sentence with an IBAN, and hidden text with a mobile number, each marked by its RTF command. {\info{\author Declan Cullinagh}{\*\company …}} {\*\atnauthor Eimear Gilsenan} {\*\annotation … PPS number again: 0000000W} {\deleted The refund will be paid to IBAN …} {\v Mobile for texts: 087 000 0000} Invented data. Shortened with … where a line runs long.
A test RTF written in Word's style for this guide, as a text editor shows it, 29 September 2026

Each of those commands is in Microsoft’s RTF specification, version 1.9.1: an information group with fields for author and company, a comment destination that can carry its author’s name, a revision mark for text deleted since tracking was switched on, and a control word for hidden text. Because the format has a place for each of them, a Word document saved as RTF can carry all four across. Do not count on the conversion to drop them.

Clearing an RTF before it goes

Open the RTF in Word or LibreOffice, accept or reject every change, delete every comment, remove hidden text, and check the header and footer. Then save it again and open the result in Notepad or TextEdit. In Word, accepting the changes and deleting the comments have their own commands on the Review tab, set out in the Word steps for tracked changes and comments.

Search the raw file for three commands: \annotation, \deleted and \v on its own. If none of them appears, those three layers are gone. The author sits in the \info block at the top, where you can read it and, in the editor, delete it. File properties more broadly, on Word and PDF files, are in the metadata guide.

Keeping only the columns the task needs

Every format in this guide ends in the same place: the data you leave in is the data the AI tool gets. Data protection law puts a principle on that. The DPC’s summary of the GDPR principles describes data minimisation as processing that is “adequate, relevant, and limited to what is necessary”, and adds that personal data should be processed only if the purpose could not reasonably be met by other means.

For a CSV, that reads as a question to ask of each column. A typical client export might look like this:

Column in the exportDoes the question need it?
Client namerarely; a code will do
PPS numberalmost never
Email and mobileno, unless the task is about contact
Eircodethe county, at most
Return type and statusyes: this is the analysis
Notesread it first; it is often the risk

When the tool has to tell clients apart, replace each name with a code and keep the table pairing each code with a person in a separate file. The model can still count, compare and group clients; it just cannot say who they are.

That is pseudonymisation, and the DPC’s 2019 guidance is that while the source data is kept, the coded data is pseudonymised and still personal data, as our pseudonymisation guide explains. A full Eircode points to one address, which is why the PPSN guide treats it as a door rather than a detail.

Nonimo on a tax return tracker

Nonimo does the covering for you once the export is trimmed. Pick the export in File Explorer (Finder, if you work on a Mac) and use the Nonimo keyboard shortcut. Your clipboard now holds the file’s text, each personal detail replaced by a label, ready for ChatGPT or Claude. Paste the reply back and the real names return.

We ran it on 29 September 2026, using the Nonimo app, on an invented tab delimited export saved in UTF-8:

LineIn the exportWhat ChatGPT would get
headerClient · PPS number · Mobile · Email · Refund IBAN · Return · Statusunchanged
1Orla Ní Dhuibhir · 0000000W · Form 11 · Filed[PERSON_1] · [REFERENCE_1] · Form 11 · Filed
2Diarmuid Ó Beacháin · 087 000 0000 · Form 11 · Waiting on P60[PERSON_2] · [PHONE_1] · Form 11 · Waiting on P60
3Muireann Kilgallon · muireann.kilgallon@example.com · IE29AIBK93115212345678 · CT1 · Drafted[PERSON_3] · [EMAIL_1] · [IBAN_1] · CT1 · Drafted

Invented data; blank cells dropped and the tabs shown as dots.

It opens CSV and TSV exports, and .txt and .md files, saved in UTF-8, and from an .odt or RTF it takes the body, tables, header and footer. In a text note, a label such as Client: in front of each name gets it picked up straight away. Our security page explains how Nonimo treats data.

One more read before the CSV goes to ChatGPT

The person who prepared the file knows what it is meant to say, which is exactly what makes them miss what else it says. A colleague reading the raw text for five minutes is a better check than any setting. Five questions cover it:

  1. Opened in Notepad or TextEdit, how many columns does the header list, and is each one needed?
  2. Does every name with a fada read correctly, or is there a replacement mark somewhere?
  3. Is there a free text column, and has someone read every cell in it?
  4. For an .odt or RTF, are there any comments, tracked changes or hidden passages left?
  5. Could what remains, a county, a job and a year, still point to one person?

Where exports go to an AI tool weekly, it helps to write these questions into the office’s AI policy, so the check does not depend on someone’s memory. Accountants and solicitors handle client exports all the time, and we describe how Nonimo fits their work on the page for Irish accountants and the one for Irish solicitors.

Sources

Nonimo is the software that does this on your own computer: it masks client names and IDs before your text reaches ChatGPT. No account needed, and the app does it without your files leaving your machine.

Common questions

Can I upload a CSV to ChatGPT?

Yes. OpenAI's File Uploads FAQ names the CSV as a file ChatGPT can analyse, with a size ceiling of roughly 50MB. What it analyses is the whole file: every row and every column, including ones you never look at. So before a CSV goes up, cut the columns that play no part in your question, put codes where the names were, and read the raw file once in Notepad or TextEdit. Nonimo can then cover the remaining names and numbers.

Does a CSV file keep hidden sheets and comments from Excel?

No. A CSV has no room for them: the DPC describes it as values from a spreadsheet with no formulas or cross references, and Microsoft says Excel writes only the active sheet into it. What it does keep is every column of that sheet, whether or not anyone needs it. Look at it in Notepad rather than Excel to see it the way a program reading the upload would.

Which CSV option should I choose when saving from Excel?

Pick the CSV UTF-8 entry in the Save As list rather than the plain CSV one. UTF-8 is an encoding that holds every character, so names with a síneadh fada arrive as they were typed. Microsoft notes that Excel opens a UTF-8 CSV normally when it was saved with a byte order mark. Nonimo reads CSV, TSV and text files saved in UTF-8, so the same choice suits it.

Why do fadas turn into strange symbols in a CSV?

Because the file was written in one encoding and read in another. Microsoft explains that text is stored as numbers and each encoding maps them to different characters. In the older Western European (Windows) set, Ó is a single byte; a program expecting UTF-8 cannot read that byte, and shows a replacement mark. Save the file as UTF-8 and the names survive the trip into ChatGPT, Claude or Nonimo.

Is a plain .txt file free of hidden content?

Largely, yes. A .txt holds only characters and line breaks, which is why the DPC suggests plain text when redacting. Nothing sits behind the words, so what your editor shows is what gets uploaded. The risk is the visible text itself: pasted email chains, signatures with mobile numbers, quoted replies. Read every line before you upload it. Nonimo covers labelled names, PPS numbers, phones, emails and IBANs in it.

How do I clear comments from a LibreOffice .odt file?

In Writer, open Edit, Comment and choose Delete All Comments, which LibreOffice's help says removes every comment in the document. Resolving a comment is not the same: it only marks it Resolved. Then go to Edit, Track Changes, Manage and use Accept All or Reject All, because struck text is only removed once the change is accepted. Nonimo then reads the body, tables, header and footer.

Are Word comments still inside an RTF copy?

They can be. Microsoft's RTF specification has a place for comments, with an author's name, for text deleted under revision marks and for hidden text, as well as an information block with author and company. Do not count on the conversion to drop them. Accept or reject the changes and delete the comments in Word first, then save the RTF and open it in a text editor to check.

Does Nonimo work with CSV, TSV and plain text files?

Yes. Save the file in UTF-8, select it in File Explorer, or in Finder if you work on a Mac, then use the Nonimo keyboard shortcut: the clipboard receives the file's words, in which names, PPS numbers, phone numbers, emails and IBANs have become labels, and the names come back when you paste the reply. It also reads LibreOffice .odt and RTF documents.