CSV hygiene

Clean user-supplied CSV tables without collecting external data

CSV Cleaner & Deduplication Tool trims whitespace, drops empty rows, deduplicates selected columns, and sorts records in a CSV you paste or fetch from a public URL. It is table hygiene on data you already hold — not email MX validation, not phone E.164 formatting, and not third-party enrichment.

$1.00 / 1,000 cleaned csv rows ($0.001 per cleaned CSV row)

Open CSV Cleaner & Deduplication Tool on Apify

User-supplied table hygiene, not email MX or phone E.164

Use this page when the job is cleaning a CSV you already have. Use Bulk Email Syntax & MX Validator for RFC syntax, DNS MX, and disposable-domain checks on email strings. Use Bulk Phone Format Validator for offline E.164 phone strings. Pipe cleaned columns into those Actors after this step.

This Actor Bulk Email Syntax & MX Validator Bulk Phone Format Validator
Intent Trim, drop empty rows, dedupe, sort a CSV Syntax, MX, disposable checks on emails Offline E.164 format on phone strings
Input csvUrl or csvData emails (max 1000) numbers (max 1000)
External collection None. User-supplied tables only Public DNS MX for supplied addresses None. Offline format only
Not this job Mailbox probes, HLR, third-party enrich CSV parse/dedupe of mixed tables CSV parse/dedupe of mixed tables

Store ID: taroyamada/csv-data-cleaner. Compliance: user-supplied tables only.

Use cases

How is CSV Cleaner different from Bulk Email Syntax & MX Validator and Bulk Phone Format Validator?

This Actor cleans a user-supplied CSV table: trim whitespace, drop empty rows, deduplicate selected columns, and sort. It does not collect third-party data. Store ID taroyamada/csv-data-cleaner. Bulk Email Syntax & MX Validator checks email strings you already have for RFC syntax, DNS MX, and disposable domains; it does not parse CSV tables. Bulk Phone Format Validator formats phone strings to E.164 offline; it does not clean CSV. Pass the cleaned table onward to those Actors when the columns are emails or phones. This is table hygiene on data you provide, not enrichment.

What input is required?

The live input schema has no required array. Supply csvUrl or csvData. README names datasetId, detectTypesOnly, normalizeNulls, dedupeRows, transformations, and removeDuplicates are not live schema fields. Do not invent them.

Field Type Default Notes
csvUrl string — Public URL to fetch CSV. No local file upload
csvData string — Raw CSV alternative to URL
delimiter string , Field delimiter. Configurable; schema default is comma
trimWhitespace boolean true Trim values
removeEmpty boolean true Drop all-empty rows
dedupColumns string[] — Columns to dedupe by
sortBy string — Sort column
sortOrder string asc asc or desc
delivery string dataset dataset or webhook
webhookUrl string — Used when delivery is webhook
dryRun boolean false Run without saving results

Published Store example run input:

{
  "csvData": "name,age\nAlice,30\nBob,25\nAlice,30",
  "trimWhitespace": true,
  "removeEmpty": true,
  "dedupColumns": ["name"],
  "sortBy": "name",
  "delivery": "dataset",
  "dryRun": false
}

Published Actor input object example (schema defaults, including the sample table with a duplicate and an empty row):

{
  "csvData": "name,email,city\nAlice,alice@example.com,Tokyo\nBob,bob@example.com,New York\nAlice,alice@example.com,Tokyo\n,,\nCharlie,charlie@example.com,London",
  "delimiter": ",",
  "trimWhitespace": true,
  "removeEmpty": true,
  "sortOrder": "asc",
  "delivery": "dataset",
  "dryRun": false
}

Run CSV Cleaner & Deduplication Tool on Apify

What does a result contain?

The published README output table lists rowNumber, data, changes[], dropped, and dropReason. A separate README sample shows aggregate inputRows, outputRows, duplicatesRemoved, emptyRowsRemoved, and cleanedData[]. Those are two published shapes; treat both as illustrations, not a live coverage guarantee.

Does this Actor fetch third-party data or validate types?

No. User-supplied tables only; no third-party enrichment API. There is no type validation. Local files are not accepted as uploads; fetch via a public csvUrl or paste csvData. The schema delimiter field is configurable (default comma). A README FAQ that says TSV is not supported directly conflicts with that delimiter field; prefer the live schema. README also notes in-memory processing and approximate size guidance; those are README claims, not live schema limits.

How is CSV Cleaner & Deduplication Tool priced?

Billing is pay per event. The live Store card is $1.00 / 1,000 cleaned csv rows. The billed event is Cleaned CSV row (apify-default-dataset-item) at $0.001. There is no Actor Start on the current pricing tab. The README Cost section that still quotes actor-start $0.01 plus dataset-item $0.001 ($1.01 per 1,000) is stale versus live Store rates.

$1.00 per 1,000 cleaned csv rows ($0.001 per cleaned CSV row)

See CSV Cleaner & Deduplication Tool pricing on Apify

Limits to keep in mind

Open CSV Cleaner & Deduplication Tool on Apify

Related pages