CSV Duplicate Remover
Deduplicate complete CSV records, not physical lines. Choose matching columns and keep the first or last representative, with separate exports for retained and removed rows.
Parsing, matching and serialization happen locally in your browser. Inputs are not uploaded or saved; refreshing clears the draft.Original CSV / TSV
0 B / 5 MiB UTF-8Limits: 5 MiB UTF-8, 50,000 data records, 200 columns, 20 MiB combined exports. Quoted newlines stay within records; this is not physical-line deduplication. Drafts are memory-only.
How to remove CSV duplicates
- Paste CSV/TSV or import one UTF-8 text file. Select comma, semicolon or tab explicitly, and indicate whether the first record contains column names.
- Press Parse and choose columns. Review the first 50 data records and the record/column counts. All columns are selected initially; choose at least one matching key.
- Select Keep first or Keep last. Optionally ignore edge whitespace or case in keys. These settings change matching, not the original cell values.
- Review formula protection and BOM options, then press Remove duplicates. Check the retained/removed counts, export previews and changed-cell examples.
- Copy or download each complete export. Clear / Reset offers one-step Undo of original input, parsing settings and key selection; rerun removal for fresh output.
Worked example: duplicate customer IDs
id,name 1,Alice 1,Alicia 2,Bob
Using id alone as the key, Keep first retains Alice and Bob; Keep last retains Alicia and Bob. Matching both columns retains all three because the names differ. For keys A, B, A, Keep last produces B then A: representatives stay in their original selected-record order, not the first key encounter order.
Parsing records and headers
The bundled csv-parse parser accepts LF/CRLF, quoted delimiters, doubled quote escapes, quoted newlines and one leading UTF-8 BOM. Completely empty physical records outside quotes and a terminal separator are skipped. A quoted empty field or delimiter-separated empty values remain a data record. Every cell is a string: 001 and 1 remain distinct, and dates or large identifiers are not inferred.
Header names are edge-trimmed for display and blank or duplicate trimmed names are rejected. Original header cells are preserved in exports, except for the explicitly selected formula protection. With headers off, the interface labels columns Column 1, Column 2 and so on; no header is invented in either export. Header-only input has zero data records.
Ragged records and malformed quoting block output with useful parser-record references. A parser record is a logical CSV record, not a physical line when quoted newlines are present. Invalid UTF-8 imports leave the original draft intact. Excel and OpenDocument workbooks are not parsed; export them to UTF-8 CSV first.
Collision-safe matching and source order
Selected cells form a JSON-encoded array key, so embedded separators, quotes and values such as [ab,c] versus [a,bc] do not collide. Optional normalization trims edges first, then uses locale-independent Unicode lowercasing. There is no fuzzy matching, accent folding or implicit Unicode normalization.
Keep first chooses the lowest original data-record index per key; Keep last chooses the highest. Representatives are exported in original record order. Removed rows include every other occurrence exactly once, also in original order, without deduplicating the removed set. Parsed data records always equal retained plus removed; headers are excluded.
Every explicit rerun reparses the unchanged original input in a worker and matches that original data, never the previous cleaned result. Editing input or parsing settings invalidates the preview and exports; changing matching or export options invalidates exports and retains the original inspection.
Formula protection and complete UTF-8 exports
Default-on protection adds an apostrophe to strings or headers beginning with =, +, -, @, tab or carriage return, including leading whitespace before formula markers. All CSV cells are strings, so negative numeric-looking text may change too. Matching happens before protection. Each file shows its own changed-cell count, including its header, plus examples and the actual serialized text preview.
Unchecking protection explicitly selects raw output. CSV quoting alone does not stop spreadsheet formulas; even protected export is not a universal guarantee for every importer. Choose an appropriate spreadsheet import mode for text identifiers, dates and leading zeros.
Both downloads preserve the chosen delimiter and header presence, use CRLF record separators and standard quote escaping, and contain all accepted columns and rows. Serialization may change quoting or record-separator bytes while preserving parsed values. An empty removed set exports the header only when enabled, otherwise an empty file (or BOM only if selected).
Names use the sanitized imported filename stem, or original for pasted input: original-deduplicated.csv and original-removed.csv. Tab output uses .tsv. UTF-8 BOM is optional and defaults to the imported file's BOM presence. Export names do not contain source paths. Copy includes the same complete text as download, including the selected BOM.
Limits, cancellation and previews
Limits are 5 MiB input, 50,000 data records, 200 columns and 20 MiB combined serialized output. Limits are checked without truncating source or returning partial exports. Local workers keep parsing and matching off the main thread; Cancel terminates the job, and revision guards discard obsolete file reads and clipboard replies.
Tables show the first 50 records, first 20 columns and at most 160 characters per cell or header label. Column indexes distinguish long clipped labels. Text previews show up to 8,192 characters. These are explicitly bounded previews, not complete output. Copy and downloads are complete; clipboard denial exposes the complete requested export for manual copying.
No data is persisted in URLs, browser storage or analytics. Object URLs are released on changes, reruns and unmount. Undo is in-memory only and does not restore obsolete exports. This release does not merge files, compare two datasets, infer types, or deduplicate arbitrary text by physical line.
How was this tool?
Help us improve CSV duplicate remover.
Opens your email app; you send it there. No email app? Copy your comment and email feedback@tooltulip.com.
Don't include sensitive data.
0/300