CSV Géant

CSV Géant › Guides

Guide · Duplicates

Remove duplicates from a large CSV file

A mailing list merged from several exports, an order history pulled twice, a log with repeated lines: duplicates inflate counts and send the same email twice. Excel's Remove Duplicates works until the file no longer fits in a worksheet. Here is how to deduplicate a CSV of millions of rows in your browser, without uploading it.

Remove duplicates from my CSV

In short

Open your file in CSV Géant, tick “Remove duplicates”, choose whether two rows are duplicates when the whole row is identical or when one column is (email, customer ID…), then export. The first occurrence of each row is kept, in its original order, and the output is a clean CSV that opens in Excel.

What counts as a duplicate

SettingTwo rows are duplicates when…Typical use
Duplicate when identical on: the whole rowevery field is the sameAn export pulled twice, repeated log lines
Duplicate when identical on: one columnthat column has the same value, whatever the other fields sayOne row per email address, per order number, per SKU
Ignore case and surrounding spaces (ticked by default)values match after trimming spaces at both ends and lower-casing: “ Ann@Example.com” = “ann@example.com”Lists typed by hand or merged from several sources

The comparison setting only affects matching: the row that is kept is written exactly as it appears in your file. The header row is never treated as data. Duplicates are removed after the filter, so you can, for example, keep only one country and deduplicate it in the same pass.

Step by step

  1. Open the tool in Chrome or Edge, on a computer, and choose your file (.csv, .tsv, .txt). It is not uploaded.
  2. Check the preview: encoding and delimiter are detected automatically; the column names must appear correctly, because you will pick one of them.
  3. Under “Deduplicate”, tick “Remove duplicates” and set “Duplicate when identical on” to “the whole row” or to a column.
  4. Keep or untick “ignore case and surrounding spaces”, depending on whether “ACME” and “acme” are the same thing for you.
  5. Optional: add filter conditions, choose the columns to keep, sort, or split the result into several files.
  6. Click “Export”. The summary tells you how many rows were read, how many duplicates were removed and how many rows were written.

How it scales to millions of rows

Keeping every distinct row in memory doesn't work at 20 million rows. CSV Géant reduces each row (or the chosen column) to a 64-bit fingerprint and stores only that, in a compact table: 16 to 32 bytes per distinct value, about 256 MB for 10 million. The file itself is streamed in 8 MB chunks.

Measured on a 2.22 GB file (22 million rows, 660,441 exact duplicates)Time
Deduplicate every row and split into 21 Excel-sized files written to a folder (2.18 GB output)102.2 s and 106.9 s
Filter “country = FR” and deduplicate, single download (8,001,769 rows, 0.82 GB)71.8 s and 72.0 s

Conditions: AMD Ryzen 5 4500U, 16 GB RAM, NVMe SSD, Windows 11, Chrome 154 driven by an automated test; each result was checked against the expected counts. Timings depend on your machine.

And in Excel, Python or the command line?

  • Excel: Data › Data Tools › Remove Duplicates works well on a file that fits in a worksheet, that is 1,048,576 rows at most. Microsoft notes that it permanently deletes the duplicate values and advises copying the data first. Beyond the worksheet limit, the rows that were not loaded are never checked (see CSV too large for Excel).
  • Python (pandas), PowerShell: they can do it if you write and run a script; the simplest versions load the whole file into memory.
  • sort -u (Linux, macOS): fast, but it reorders the file, moves the header, compares raw lines only, and can break records that contain line breaks inside quotes.

Source: Microsoft Support, “Filter for unique values or remove duplicate values” and “Excel specifications and limits” (accessed October 7, 2026).

Free or Pro

  • Free: the tool reads the whole file, so the summary gives the full number of duplicates; the export writes at most 100,000 rows per operation.
  • Pro, €29 paid once: unlimited exported rows, in the browser where you enter the licence key emailed by Gumroad. No subscription.

Limits, honestly

  • One column or the whole row: you cannot yet combine two columns (for example first name + last name) as the duplicate key.
  • Exact matching only, apart from case and surrounding spaces: “Jon Smith” and “John Smith” are not duplicates.
  • The first occurrence is kept; there is no “keep the last one” option and no export of the removed rows.
  • Fingerprints are probabilistic: two different rows with the same 64-bit fingerprint would be treated as duplicates. The chance is about 1 in 80,000 for 21 million distinct rows.
  • Memory: 16 to 32 bytes per distinct value; largest file measured, 2.22 GB.
  • Tested in Chrome only. Edge uses the same engine but is not tested; Firefox and Safari are not tested.

Frequently asked questions

Which duplicate is kept?

The first one in the file. Later rows that match it are dropped, and the remaining rows keep their original order unless you also sort.

Can I remove duplicates based on one column only?

Yes: set “Duplicate when identical on” to that column (for example the email column). The other fields of the kept row stay as they are.

Is my file uploaded to a server?

No. Deduplication runs in your browser, on your computer. The server only receives anonymous counters.

Can I remove duplicates from a CSV online for free?

Yes, up to 100,000 exported rows per operation, with no account. Unlike most online tools, the file never leaves your computer.

Clean up your CSV

No account, no install, no upload.

Open the tool

Other guides