DR-001 · Accepted · 2026-10-09
Synthetic data for confidentiality
Decision: Run every page, test and study on a seeded synthetic "lab world" with the shape of the 2021 HPLC exports, and never ship, read or reproduce the industry partner's data, formats or business rules.
Context
In 2021 the team worked on confidential lab-result spreadsheets and HPLC control exports supplied by the industry partner (CSL). The revival had to make the methods explorable in a public web app without exposing anything that came from the partner: values, identifiers, process-stream names, sample-naming rules or charts drawn from the real data. The 2021 coursework and its outputs stay in a private repository, and data/ and plots/ are ignored everywhere.
Recorded on 9 October 2026 for a decision taken when the site was first rebuilt, so the 2026 upgrade (evaluation studies, decision records, optional AI) inherits it explicitly.
Decision
- Every sheet, chart and cluster comes from
web/src/lib/synthetic, built from one integer seed (default 30034, the notebooks'random_state) with a small mulberry32 generator. - Column names mimic generic chromatography exports; every value, stream name, identifier format, sequence name and instrument name is invented.
- Faults are planted on purpose (out-of-control runs, a slow drift, degraded samples, blank and text cells) and listed, so the site can say what is real signal and what is noise.
- The original 2021 code is run on the same synthetic inputs (
scripts/parity_reference.py) to prove the TypeScript ports behave like it. - The optional AI feature (DR-004) only ever sees numbers derived from this synthetic data.
Options considered
- Anonymise the real data (rename streams, rescale values, jitter dates). Rejected: re-identification risk, and the partner never agreed to any release.
- Use a public HPLC data set. Rejected: none had both messy lab-result sheets and matching control-assay exports, so the join between parsing and charting (Sample Name to Injection Name) could not be shown.
- Seeded synthetic generator with planted faults (chosen).
- Static screenshots of the 2021 outputs. Rejected: they would show partner data, and nothing would be explorable.
Why
The generator keeps the structure that makes the methods interesting (free-text identifiers, several process streams, bracketing control injections, runs that carry one stream's samples) without any of the content. Planted faults also give something the real project never had: ground truth, so a method can be scored on whether it finds what was planted.
What happened
- The parity suite runs the 2021 Python on the synthetic inputs and the ports match: all 43 assay, control and measurement charts reproduce the notebook's order, seven limit lines and printed out-of-control list.
- Over 40 synthetic worlds the parser handles 99.74% of rows (95% bootstrap CI 99.69% to 99.78%; range 99.33% to 100%). That is close to the 2021 figure of about 99.7% by construction: the generator's junk-row rates were chosen to give it. It shows the parser's rules work on the formats I invented, not that they would work on new real formats. This is the weakest number on the site and I say so wherever it appears.
- The planted faults were made findable (control shifts of 5 to 6 sigma, degraded samples 8 to 11 spreads away). Recovery rates on /evaluation are therefore upper bounds for the methods, not estimates of how often a real lab would catch a real fault.
- Nothing from the partner is in this site's code, data or content: every value shown is generated from the seed.
What I'd change
- Calibrate the generator's spreads and fault sizes against published HPLC system-suitability ranges instead of my own guesses, and add autocorrelation within runs, which real controls often show.
- Generate a second, harder world (smaller faults, more realistic typing errors) and report the parse rate and fault recovery on both, so the easy case is not the only one.