Using Gridline

Cleaning files

Fix what a report found: trim, de-duplicate, standardise dates and numbers, hide personal data, and save the result as the next version.

A quality report tells you what is wrong with a file. Clean fixes it. You choose the steps (trim spaces, remove repeated rows, write dates one way, hide a column of personal data) and Gridline writes the result as the next version of the file. Your upload is never touched, so there is always a way back.

Clean a file

  1. Open the file and choose Clean

    It is next to Upload new version, and appears once the report is ready. Only the person who uploaded the file, or an admin, sees it.

  2. Look at the steps

    Gridline lists the steps that fit what the report found, and switches on the safe ones: it only offers to remove empty rows if there are some. Steps that change meaning (dates, hiding data, placeholders) are listed but off, until you turn them on.

  3. Read what would change

    Beside the steps, What would change counts what each step does across the whole file, and shows the first rows that change with each cell before and after. It updates as you edit.

  4. Create the cleaned version

    It is made in the background and checked by the same rules as any upload. When it is ready you land on the comparison with the file it came from.

The steps

StepWhat it does
Trim spacesRemoves spaces at the start and end of cells, and doubled spaces inside them. One column or all.
Tidy the header rowThe same for the column names.
Remove empty rowsDeletes rows with nothing in them.
Remove repeated rowsKeeps the first of each. Put it after Trim spaces and it also finds rows that differed only by spaces.
Write dates as YYYY-MM-DDReads 25/12/2025, 4 Mar 2026, March 4, 2026 and ISO dates. A date like 03/04/2026 can be 3 April or 4 March, so you say which to assume, and the preview tells you how many were like that. Anything it cannot read is left alone, never guessed.
Read text as numbersTurns 1.234,50 € or $1,234.50 or (2,000) into a number. You say whether the decimal mark is a point or a comma. Percentages, ranges and words are left alone.
Replace placeholdersSwaps N/A, null, - and the like for nothing, or for a value you choose.
Fill empty cellsPuts a value into the empty cells of one column.
Change letter caseCapitals, lower case or Title Case for a column.
Rename or remove a columnChanges the header, or leaves the column out.
Hide personal dataReplaces what a column holds: hide it completely, keep only the last four characters, or replace each value with a fingerprint (the same value always gives the same fingerprint, so you can still count and join on it, but nobody can read it).

What comes out

  • A new version. Named like the file with “(cleaned)”, numbered after the newest version, shared with the same people.
  • The same kind of file. A CSV stays a CSV and a workbook stays a workbook, with real numbers and real dates. A workbook keeps only the sheet that was cleaned; a legacy .xls comes out as .xlsx.
  • Not an upload. It does not use your file allowance, because you did not upload it. It does count towards the number of versions your plan keeps, and it appears in the audit log as File cleaned.
  • Its own report. Compare it with the version it came from to see the score rise and the problems go.

Do it every month

Tick Remember these steps for this file and they are filled in the next time. On the Basic and Premium plans you can also tick Clean every new version of this file automatically: each upload is kept exactly as it arrived, and a cleaned version follows it. A cleaned version is never cleaned again.

From code

Everything the page does is available over HTTP. Cleaning needs the files:write scope.

POST/files/{id}/clean/previewCounts what a recipe would do. Writes nothing.
POST/files/{id}/cleanStarts it. Answers 202 with a job to poll.
GET/files/{id}/clean/jobs/{jobId}The job: queued, running, succeeded (with resultFileId) or failed (with the reason).
GET/datasets/{datasetId}/settingsThe saved recipe and whether every new version is cleaned.
PUT/datasets/{datasetId}/settingsSaves them.

A recipe

{ "steps": [ { "step": "trim_whitespace" }, { "step": "drop_duplicate_rows" }, { "step": "standardise_dates", "column": "joined", "order": "dmy" } ] }