Skip to content
HPLC QC LabMAST30034 · revival

Build notes, data and credits

How this revival was built

The aim was to make a 2021 industry capstone explorable without touching anything confidential: port the team's methods faithfully, feed them synthetic data with the same shape, and prove the ports behave like the original code.

The 2026 upgrade measures how far those methods can be trusted, on /evaluation, and records its decisions, assumptions, model card and AI use statement on /methods.

Clean-room rules

  • No partner data, code, documents, logos or business rules are used or shipped. The original data and the charts generated from it are not part of this website, and the coursework repository stays private.
  • Process-stream names, identifier formats, sample names, sequence names and instrument names on this site are invented.
  • The assignment brief is paraphrased, not reproduced. The team's report and slides are summarised, not hosted.
  • The four control assays (amino acid excipient, cation exchange, Protein A, size exclusion) are generic HPLC methods; their columns mimic common chromatography exports with invented values.

The synthetic generator

One integer seed (default 30034, a nod to the notebooks' random_state) builds a whole “lab world” with a small mulberry32 generator:

  • A lab-results sheet of 1,200 rows split across two workbooks, four invented streams, each with a canonical format plus five to eight realistic ways of typing it wrong. Each row independently has about a 0.6% chance of missing a key field and about a 0.3% chance of being genuinely unparseable, so the parsed share varies by seed.
  • 60 sequences per assay between July 2019 and August 2021, each with bracketing control injections and the sample injections submitted on the sheet. A run mostly carries one stream's samples (85% of a stream's samples go to runs of that stream), so an out-of-control run can be traced back to one or two streams through Sample Name.
  • Planted problems: runs out of control on several measurements, single-measurement jumps, a seven-point upward drift in AAE, blank cells, “ND” and “Overloaded” results, and about 2.5% of samples degraded on two features.

Synthetic assays

Columns charted per assay. Clustering drops the same columns the team dropped in 2021.
Synthetic assay definitions
AssayControlsMeasurementsTime order
AAEAmino acid excipientControl std A, low (histidine)Control std A, high (arginine)Area, Retention Time, Amount (pmol), Amino Acid (mM), VolumeDate from Sequence Name + row index, as the team did
CEXCation exchangeControl std B (system suitability)Rel.Area Basic, Rel.Area Acidic, Rel.Area, Ret.Time, Area Acidic, Area, Area Basic, Total Area, Peak Width (50%), Resolution (SM) Pre-Main, Resolution (SM) Main-Post, VolumeInject Time
Protein AProtein A titreControl std C, lowControl std C, highTotal Area, Amount, Concentration, Retention, VolumeInject Time
SECSize exclusionControl std B (system suitability)Area, Retention Time, Rel.Area Monomer, Height, Peak Width (50%), Asymmetry (EP), Plates (EP/m), Rel Area Aggreg, Rel Area Frag, Total Area, VolumeInject Time

Parity with the 2021 code

A Python script runs the original code on the same synthetic inputs and saves its outputs; the TypeScript test suite compares against them on every CI run.
Lab results preprocessing
merge_result, preprocess_department_ID, sample_name and compound are imported from the team's LabResultsPreprocessor.py and run on the synthetic workbooks (written to .xlsx first). Row identifiers, department split, Sample Name and Compound match row for row. The synthetic stream cascade is checked against an independent pandas-semantics implementation.
Control charts
The create_control_chart function is executed straight from the FA2 notebook source with its matplotlib calls recorded. For all 43 assay, control and measurement combinations the TypeScript port reproduces the plotted order, the seven limit lines and the printed out-of-control list.
Clustering
The FA4 preparation (drop controls and columns, fill 0, StandardScaler) and scikit-learn's PCA (with the 2021 sign rule), silhouette_score, DBSCAN (core points and noise), K-Means Lloyd iterations from fixed centres, and nearest-neighbour distances all agree to floating-point tolerance. The Python side runs on exactly pinned library versions.

Known differences

  • K-Means uses its own seeded random stream, so k-means++ starting points differ from NumPy's. Restarts (10), iterations and tolerance follow scikit-learn and the Lloyd iterations match it; final labels can differ the way two random seeds would. The sweep and the clustering on screen are the same fitted models.
  • The K-Means sweep covers k = 2 to 20. The team's PCA and t-SNE notebooks swept k = 2 to 49 on a few thousand rows; at 300 synthetic samples the curve has settled well before 20.
  • t-SNE is the exact algorithm with the 2021 scikit-learn optimiser: random start, 250 iterations of early exaggeration 12 at momentum 0.5, then momentum 0.8, learning rate 200, with the momentum and gains restarting between the two phases. The team used the Barnes-Hut approximation on thousands of rows; at 300 rows the exact version is fast. scikit-learn's early-stopping checks are not ported, so every run does the full number of iterations.
  • PCA axes use the sign rule of the scikit-learn the team ran in 2021 (the largest projected score on each axis is positive), so maps are oriented like the notebook plots. scikit-learn 1.5 and later can mirror an axis; the parity test checks both.
  • UMAP runs through umap-js, the PAIR team's JavaScript port of umap-learn. Embeddings are similar in character but not identical to the Python library.
  • DBSCAN uses one min_samples list (5 to 40) for every projection; the team picked a list per assay and projection, from 5 to 75. The eps grid adapts to each projection's scale instead of the notebooks' fixed ranges, running from the median nearest-neighbour distance to 1.5 times the 95th percentile of the 20th-neighbour distance. No k-distance plot is shown.
  • The notebook's k-distance plot read column 1 of the neighbour distances for every k, so all six panels showed the nearest-neighbour curve. The port reproduces that column and tests it against scikit-learn.
  • The “settles” marker is a heuristic stand-in for the team's judgement by eye: the first point after which the next four scores stay within 0.03 of each other, near the top of the curve.
  • merge_result runs with its default remove_nan=True, so rows missing a key field are set aside at the first step. The 2021 notebook switched that off and lost those rows later in the stream filters. As in pandas, cells typed as n/a or NULL count as missing.
  • Two quirks are kept on purpose: Compound concatenates matches from both columns ("Compound-1Compound-1"), and the AAE time stamp is sorted as text, so row _10 comes before row _9 on the same date.
  • Run rules, I-MR limits, Phase I windows and injected disturbances are additions. The 2021 charts flagged only points beyond 3σ, with σ from all points.

Credits

Group 07, “Team 4399”, MAST30034 Applied Data Science, The University of Melbourne, Semester 2 2021: Yuchen (Cynthia) Cai, Rongshun Li, Zhiliang (Tommy) Tang, Quzihan (Martin) Wu and Sunchuangyu (Rin) Huang. The lab-results parser, control-chart function and clustering notebooks were written together; individual stream parsers and notebooks carry their authors' names in the original repository.

The industry partner, CSL, provided the brief, the data and feedback in 2021. This revival is independent and not a CSL product.

The 2026 web revival was built by Rin Huang. Next.js, React, Tailwind CSS, shadcn/ui, Lucide icons, umap-js, Vitest, pandas and scikit-learn are open source.

Academic integrity

The original submission is preserved unchanged in the private repository for reference. This site is a reimplementation for portfolio purposes; it is not a model answer and should not be submitted as coursework.

Further reading the team used includes the FAIR Guiding Principles (Wilkinson et al., 2016) and papers on run rules and control-chart design.

rin.contactgithub.com/rNLKJAFAIR principles