MAST30034 Applied Data Science · University of Melbourne · 2021
HPLC QC LabA student capstone on lab data quality, rebuilt to run in your browser.
In 2021 our team of five worked with an industry partner on three problems: untangling messy lab-result spreadsheets, charting HPLC control assays to spot runs that drift, and testing whether clustering can surface unusual measurements. This site re-implements those methods on synthetic data so anyone can try them.
CEX control · Rel.Area Basic
n = 151
The brief, in short
What the coursework asked
Paraphrased. The original brief belongs to the subject and the partner and is not reproduced here.
Clean the lab results
A required first task: turn spreadsheets of lab submissions into consistent, machine-readable records with batch identifiers, in the spirit of the FAIR data principles.
Modernise the charts
Rebuild the control charts analysts use to check that a measurement stays within expected limits, as reusable code rather than one-off plots.
Explore clustering
Try unsupervised methods on HPLC measurements and judge, honestly, whether the clusters are reliable and useful.
Three interactive demos
Try the methods yourself
Tame the spreadsheet
Generate a messy two-workbook lab sheet, then watch a rule cascade pull batch IDs, process steps and sample names out of free text, with every leftover row routed to a person.
2021: about 99.7% of records parsed
OpenRead the control charts
Chart four control assays the way the team did (mean, sample SD, 3σ limits), then switch on Western Electric and Nelson rules and inject your own shifts and drifts.
2021: the same runs recurred across measurements
OpenFind the odd runs
Project injections with PCA, t-SNE or UMAP, cluster with K-Means or DBSCAN, tune by silhouette, and see which injections the method puts outside every group.
2021: clusters as an outlier aid, not labels
Open2026 upgrade
How far can the methods be trusted?
Monte Carlo studies of the run rules, the limits and the clustering, every estimate with its sample size, a 95% interval and a fixed seed (study seed 30034); decision records and a model card; and an optional AI explanation of a chart that runs only with your own API key and logs every call in your browser.
More rules, more false alarms
Over 4,000 simulated in-control charts of 150 points, the 2021 3σ rule raises at least one false alarm in 33.5% (95% CI 32.0% to 34.9%); all eight run rules in 90.3% (95% CI 89.3% to 91.2%).
See the studyLimits from every point can hide a shift
Over 4,000 simulated charts, a sustained 3σ shift in the last third is caught by the 2021 limits in 2.5% (95% CI 2.0% to 3.0%), and by limits from the first 40 points in 100.0% (95% CI 99.9% to 100.0%). The cost: with no shift, those limits flag an early point in 27.2% (95% CI 25.8% to 28.6%) of charts, against 21.7% (95% CI 20.5% to 23.0%).
See the studyClusters and charts find different faults
Over 4,800 runs in 20 synthetic worlds, DBSCAN outliers and 3σ signals barely agree: Cohen's kappa −0.01 (95% CI −0.03 to 0.01). The rules find faulty runs, DBSCAN finds degraded samples, so clustering stays a review aid.
See the study
Key results from 2021
What the team found
Summarised from the team's own report and slides, without the partner's data or identifiers. The team logged about 101 hours between late August and October 2021.
- Lab results
- One ordered set of pattern cases per process stream turned about 99.7% of lab-result records into tidy batch, step and sample fields. Non-standard departments and the upstream rows the team could not interpret went to an “other” table; this revival extends that to every stream so no row is dropped silently.
- Control charts
- For two of the four assays, the same one or two runs were out of control on most peak measurements, pointing at the run rather than at a noisy column. The amino acid assay never crossed a limit but showed a run of six rising results.
- Clustering
- Silhouette scores rise with more clusters, so the team looked for where the curve settles. Recommended pairings differed by assay, and clusters were proposed for spotting unusual runs, not as labels.
- Recommendations
- Re-run out-of-control experiments where possible, make result-version fields mandatory so runs can be ordered, and link injection names to sample types so clusters can be interpreted.
How it was rebuilt
A clean-room revival
The real data stays with the partner. Every sheet, chart and cluster here comes from a seeded generator that mimics the shape of HPLC exports (retention time, areas, heights, widths, plate counts, sequence names) and plants a few faults on purpose so there is something to find.
The team's methods were ported to TypeScript and checked against the original 2021 Python, which runs on the same synthetic inputs in the test suite: the preprocessing helpers, the notebook's control-chart function (all 43 assay, control and measurement combinations, same limits, order and flags) and scikit-learn's PCA, DBSCAN, K-Means and silhouette.
| Aspect | 2021 original | 2026 revival |
|---|---|---|
| Language | Python 3.8 in Jupyter | TypeScript (strict), React 19 |
| Parsing | pandas + re | Pure functions, same rule-cascade approach |
| Charts | matplotlib | Hand-rolled SVG, keyboard accessible |
| Clustering | scikit-learn, umap-learn | TS ports of PCA, K-Means, DBSCAN, silhouette, exact t-SNE; umap-js |
| Compute | Notebook kernel | Browser, heavy work in a Web Worker |
| Checks | Eyeballed outputs | Vitest unit tests + parity tests against the original Python via uv |
| Evaluation | Silhouette scores | Monte Carlo run-length, limit and stability studies with 95% intervals |
| Data | Confidential industry data | Seeded synthetic generator, nothing confidential |
About this project
Subject, team and context
MAST30034 Applied Data Science
The University of Melbourne, Semester 2 2021. An industry capstone in which student groups worked on a partner's real data problem. Our partner was CSL; this site is not affiliated with or endorsed by CSL.
Group 07, “Team 4399”
- Yuchen (Cynthia) Cai
- Rongshun Li
- Zhiliang (Tommy) Tang
- Quzihan (Martin) Wu
- Sunchuangyu (Rin) Huang
This revival
Rebuilt in 2026 by Rin Huang as a portfolio piece. The coursework repository stays private because the original work used confidential data.