Skip to content
HPLC QC LabMAST30034 · revival

MAST30034 Applied Data Science · University of Melbourne · 2021

HPLC QC LabA student capstone on lab data quality, rebuilt to run in your browser.

In 2021 our team of five worked with an industry partner on three problems: untangling messy lab-result spreadsheets, charting HPLC control assays to spot runs that drift, and testing whether clustering can surface unusual measurements. This site re-implements those methods on synthetic data so anyone can try them.

CEX control · Rel.Area Basic

n = 151

Synthetic CEX control chart for Rel.Area Basic: 2 injections beyond the 3 sigma limits.UCLCLLCL2019-072020-032020-122021-08
Synthetic data. Red points sit beyond 3σ. The run 20200217_CEX_Vega_S3 is out of control on 8 measurements at once, the pattern the 2021 team learned to look for.

The brief, in short

What the coursework asked

Paraphrased. The original brief belongs to the subject and the partner and is not reproduced here.

  1. Clean the lab results

    A required first task: turn spreadsheets of lab submissions into consistent, machine-readable records with batch identifiers, in the spirit of the FAIR data principles.

  2. Modernise the charts

    Rebuild the control charts analysts use to check that a measurement stays within expected limits, as reusable code rather than one-off plots.

  3. Explore clustering

    Try unsupervised methods on HPLC measurements and judge, honestly, whether the clusters are reliable and useful.

Key results from 2021

What the team found

Summarised from the team's own report and slides, without the partner's data or identifiers. The team logged about 101 hours between late August and October 2021.

Lab results
One ordered set of pattern cases per process stream turned about 99.7% of lab-result records into tidy batch, step and sample fields. Non-standard departments and the upstream rows the team could not interpret went to an “other” table; this revival extends that to every stream so no row is dropped silently.
Control charts
For two of the four assays, the same one or two runs were out of control on most peak measurements, pointing at the run rather than at a noisy column. The amino acid assay never crossed a limit but showed a run of six rising results.
Clustering
Silhouette scores rise with more clusters, so the team looked for where the curve settles. Recommended pairings differed by assay, and clusters were proposed for spotting unusual runs, not as labels.
Recommendations
Re-run out-of-control experiments where possible, make result-version fields mandatory so runs can be ordered, and link injection names to sample types so clusters can be interpreted.

How it was rebuilt

A clean-room revival

The real data stays with the partner. Every sheet, chart and cluster here comes from a seeded generator that mimics the shape of HPLC exports (retention time, areas, heights, widths, plate counts, sequence names) and plants a few faults on purpose so there is something to find.

The team's methods were ported to TypeScript and checked against the original 2021 Python, which runs on the same synthetic inputs in the test suite: the preprocessing helpers, the notebook's control-chart function (all 43 assay, control and measurement combinations, same limits, order and flags) and scikit-learn's PCA, DBSCAN, K-Means and silhouette.

Method notes, parity tests and credits
Original stack compared with the revived stack
Aspect2021 original2026 revival
LanguagePython 3.8 in JupyterTypeScript (strict), React 19
Parsingpandas + rePure functions, same rule-cascade approach
ChartsmatplotlibHand-rolled SVG, keyboard accessible
Clusteringscikit-learn, umap-learnTS ports of PCA, K-Means, DBSCAN, silhouette, exact t-SNE; umap-js
ComputeNotebook kernelBrowser, heavy work in a Web Worker
ChecksEyeballed outputsVitest unit tests + parity tests against the original Python via uv
EvaluationSilhouette scoresMonte Carlo run-length, limit and stability studies with 95% intervals
DataConfidential industry dataSeeded synthetic generator, nothing confidential

About this project

Subject, team and context

MAST30034 Applied Data Science

The University of Melbourne, Semester 2 2021. An industry capstone in which student groups worked on a partner's real data problem. Our partner was CSL; this site is not affiliated with or endorsed by CSL.

Group 07, “Team 4399”

  • Yuchen (Cynthia) Cai
  • Rongshun Li
  • Zhiliang (Tommy) Tang
  • Quzihan (Martin) Wu
  • Sunchuangyu (Rin) Huang

This revival

Rebuilt in 2026 by Rin Huang as a portfolio piece. The coursework repository stays private because the original work used confidential data.

rin.contactgithub.com/rNLKJA