All tools
Record-aware CSV partitioning

CSV Splitter

Split one CSV or TSV into complete logical records, not raw text lines. Files, values and settings stay in your browser; no uploads or saved drafts.

5 MiB UTF-8 input, 50,000 data records, 200 columns, 200 parts/groups and 25 MiB serialized output including the manifest. Original source remains in memory; refresh clears it.

Formula protection adds an apostrophe to formula-like strings and headers, including negative numeric-looking strings. Grouping uses the original raw values before protection. CSV quoting and this protection cannot guarantee every spreadsheet import behavior.

How to split

  1. Paste or import UTF-8 and select comma, semicolon or tab. Confirm whether the first row is a header.
  2. Choose 1–50,000 records per file, or group by a column’s stable position. Inspect CSV to see the predicted parts before generating.
  3. Review raw group values, header/BOM options and formula protection. Generate explicitly, preview a selected part and download it or all parts plus manifest.json as one ZIP.

Worked examples

Five data records split by two give 2, 2 and 1 records. Each part repeats the original header when enabled. A quoted three-line cell is one record and never straddles files. Group values A, B, A create A’s file first with the first and third rows, followed by B’s file. Empty, A/a, 001/1 and whitespace differences remain distinct raw groups.

Strings, headers and empty input

No type inference, trimming or grouping-key normalization. All cells stay strings. Duplicate and empty header labels are accepted: grouping uses column index, not label. Headerless files use position labels in the UI but never gain invented data headers. Empty physical records outside quotes are skipped; quoted-empty and delimiter-only records remain. A terminal record separator is ignored. Malformed quotes and ragged records are rejected with source-record references. Empty or header-only inputs have no data parts and no ZIP.

Export format and spreadsheet protection

Output preserves the selected delimiter, uses CRLF record separators, quotes cell delimiters/line breaks correctly, and optionally includes a UTF-8 BOM. Tab files use .tsv; others .csv. Deterministic source-part-001 or source-group-001 basenames never use raw group values or source filenames, so paths, traversal segments and reserved group names cannot become ZIP paths or collide. The JSON manifest maps every exact group value to its filename and count.

Default-on protection prefixes an apostrophe to strings/headers beginning with =, +, -, @, tab, CR or whitespace followed by a formula prefix. Negative numeric-looking strings are protected too. Partitioning happens first, so protection never changes group membership. The preview displays modifications and each part’s affected-cell count. Disable protection explicitly for raw export. Neither quoting nor protection guarantees safety in all spreadsheet software.

Limits and privacy

5 MiB input, 50,000 data records, 200 columns, 200 output parts/groups, 25 MiB serialized parts plus manifest. Part/group limits are checked before generating, output growth is bounded, and ZIP creation runs locally in a worker. Progress is announced; Cancel terminates the job, and changed options discard obsolete output. Original source is preserved on errors. Preview is at most 20 data records and 20 columns with shortened cells; exports are complete. No Excel, byte-based cutting, filters or database exports. Refresh clears memory.