Skip to content
HPLC QC LabMAST30034 · revival

05 · 2026 upgrade · Methods

How the numbers were made, and what they cannot tell you

The data, the methods and what the 2026 upgrade added to them, how every claim on the site is checked, the assumptions and limits, the decisions behind them as decision records, a model card and a data card, and what the optional AI feature does and never does.

The 2021 results stay as the team reported them: the upgrade adds analysis around the original methods and does not change what they produced. The clean-room rules and credits are on /about.

Data provenance

  • All data are synthetic, generated in the browser from one seed (default 30034): a 1,200-row lab-results sheet and four HPLC assay exports of 60 sequences each. See the data card and DR-001.
  • No industry-partner data, code, documents or business rules are used or shipped. The 2021 coursework and its confidential outputs stay in a private repository.
  • The original 2021 Python is run on the same synthetic inputs, with exactly pinned library versions, so the TypeScript ports can be compared with it.
  • The evaluation studies use simulated standard-normal streams (run rules, estimated limits, masking) and the synthetic world (clustering, agreement). Every simulation is seeded; the study seed is 30034.

Method

Ported from 2021

  • Rule-cascade parsing of free-text lab-result fields into batch, step and sample identifiers.
  • Individuals control charts with the mean and sample SD of all points, bands at 1, 2 and 3 sigma, and a list of points beyond 3 sigma; sequences flagged on several measurements, joined back to process streams.
  • PCA, t-SNE and UMAP projections; K-Means and DBSCAN tuned by silhouette, read where the curve settles.

Added in the revival and the 2026 upgrade

  • Western Electric and Nelson run rules, I-MR limits, a Phase I window and injected disturbances.
  • Review prompts from clusterings (noise, tiny clusters, far-out points) instead of labels.
  • Monte Carlo studies of the run rules and the limits, and stability and agreement studies of the clustering, all with intervals.
  • Presets that show each rule set's false-alarm cost, decision records, model and data cards.
  • An optional bring-your-own-key AI explanation of the chart on screen, with an audit log.

Evaluation design

Every quantitative claim has a check, and every estimate on /evaluation comes with a 95% interval: Wilson for proportions, distribution-free order-statistic intervals for medians and percentiles, percentile bootstrap (2,000 resamples unless stated) otherwise. Where two methods are compared they see the same simulated data, and the comparison is a paired difference or ratio with its own interval. The whole study regenerates with pnpm study (274 s).

What is checked, how, and where
ClaimHow it is checkedWhere
The TypeScript ports behave like the 2021 PythonThe original preprocessing helpers, the FA2 notebook's chart function (all 43 charts) and scikit-learn's PCA, DBSCAN, K-Means and silhouette run on the same synthetic inputs; outputs compared to floating-point tolerancetest/parity.test.ts
The statistics helpers are rightNormal CDF and quantiles, binomial CDF and quantiles, Wilson intervals, type-7 quantiles and their order-statistic intervals, the adjusted Rand index and Cohen's kappa against SciPy, statsmodels, NumPy, scikit-learn and base Rtest/stats.test.ts
The simulator counts run lengths correctlyOne-point rule checks agree with the whole-series evaluator on 300 random series, ties included; ARLs match exact values for the 3σ rule (in control and shifted) and for 8 on one side (255)test/online-rules.test.ts
The published numbers are what the code producesA slice of every study re-run with the published settings and seeds, compared with the artefact; the settings themselves compared with the codetest/study.test.ts
Run-rule false alarms against detection4,000 common-random-number streams per scenario, 9 shift sizes and 2 drifts, paired ratios between rule sets/evaluation#run-rules
Limits from estimated parametersConditional in-control ARL for 7 Phase I sizes from 10 to 500 points, 5,000 simulated samples per size, sample SD and moving range on the same samples/evaluation#phase-one
Retrospective against Phase I limits4,000 charts of 150 points with a sustained shift in the last third, four limit methods on the same charts, paired differences/evaluation#masking
Clustering stability21 algorithm seeds for every pipeline, 50 bootstrap resamples for the PCA pipelines; ARI, prompt overlap and silhouette with bootstrap intervals/evaluation#stability
DBSCAN outliers against run-rule signals20 synthetic worlds, sequences as units, Cohen's kappa and planted-fault recovery with a cluster bootstrap over worlds/evaluation#agreement
The AI client is safe with a keyAdapters, error mapping, key storage, redaction, the audit log (memory and IndexedDB) and the grounding check, with the network mockedtest/ai.test.ts
The documents quote the published studyThe decision records and model card are checked for drift from docs/ and for the study numbers they quotetest/content.test.ts

Assumptions

  • Control results are independent and normally distributed around a stable centre when the process is in control. The run-length studies use exactly that model.
  • Changes are a sustained step shift present from the first point, or a linear drift; a shift that starts later in a chart behaves somewhat differently with run rules that need history.
  • Limits in the run-length study are known; the Phase I study covers estimated limits for the 3σ rule only.
  • Blank and text results are filled with 0 before clustering, as the 2021 notebooks did.
  • Planted faults stand in for real ones. They are deliberately large enough to be findable.

Limitations

  • Synthetic data cannot validate the parser on real formats: its 99.7% parse rate was tuned to match 2021, not earned (DR-001).
  • Real HPLC control data can be autocorrelated within a run, skewed or rounded; none of that is simulated, and all three change false-alarm rates.
  • Many charts are read at once (up to 11 measurements per control); the site reports per-chart rates and does not control the family-wise rate.
  • Clustering stability under resampled data was measured for PCA pipelines only; t-SNE and UMAP refits were too slow to bootstrap.
  • The quality of the AI explanations has not been evaluated; only invented numbers are checked automatically.
  • The 3σ rule's published in-control ARL, 361.9, sits near the lower end of its interval; with 4,000 streams the Monte Carlo error is a few percent.

What I'd change

  • Add an EWMA or CUSUM chart and compare it with the run rules on the same simulated streams.
  • Control the family-wise false-alarm rate across one control's measurements, or use a multivariate chart.
  • Calibrate the generator against published system-suitability ranges and add autocorrelation within runs.
  • Report how often each injection is flagged across seeds and resamples, instead of one list of review prompts.
  • Build a small, rubric-scored evaluation set for the AI explanation and publish its results with intervals.

Decision records

Each record states the decision first, then the context, the options, why, what happened (weak numbers included) and what I would change. Records are never edited after the fact; a new record supersedes an old one. The source files are in docs/decisions/.

  1. DR-001Synthetic data for confidentiality
  2. DR-002The run-rule set, and what limits to judge it against
  3. DR-003Clustering as a review aid, not a decision
  4. DR-004Optional bring-your-own-key "Explain this chart", grounded in the chart's numbers

DR-001 · Accepted · 2026-10-09

Synthetic data for confidentiality

Decision: Run every page, test and study on a seeded synthetic "lab world" with the shape of the 2021 HPLC exports, and never ship, read or reproduce the industry partner's data, formats or business rules.

Context

In 2021 the team worked on confidential lab-result spreadsheets and HPLC control exports supplied by the industry partner (CSL). The revival had to make the methods explorable in a public web app without exposing anything that came from the partner: values, identifiers, process-stream names, sample-naming rules or charts drawn from the real data. The 2021 coursework and its outputs stay in a private repository, and data/ and plots/ are ignored everywhere.

Recorded on 9 October 2026 for a decision taken when the site was first rebuilt, so the 2026 upgrade (evaluation studies, decision records, optional AI) inherits it explicitly.

Decision

  • Every sheet, chart and cluster comes from web/src/lib/synthetic, built from one integer seed (default 30034, the notebooks' random_state) with a small mulberry32 generator.
  • Column names mimic generic chromatography exports; every value, stream name, identifier format, sequence name and instrument name is invented.
  • Faults are planted on purpose (out-of-control runs, a slow drift, degraded samples, blank and text cells) and listed, so the site can say what is real signal and what is noise.
  • The original 2021 code is run on the same synthetic inputs (scripts/parity_reference.py) to prove the TypeScript ports behave like it.
  • The optional AI feature (DR-004) only ever sees numbers derived from this synthetic data.

Options considered

  1. Anonymise the real data (rename streams, rescale values, jitter dates). Rejected: re-identification risk, and the partner never agreed to any release.
  2. Use a public HPLC data set. Rejected: none had both messy lab-result sheets and matching control-assay exports, so the join between parsing and charting (Sample Name to Injection Name) could not be shown.
  3. Seeded synthetic generator with planted faults (chosen).
  4. Static screenshots of the 2021 outputs. Rejected: they would show partner data, and nothing would be explorable.

Why

The generator keeps the structure that makes the methods interesting (free-text identifiers, several process streams, bracketing control injections, runs that carry one stream's samples) without any of the content. Planted faults also give something the real project never had: ground truth, so a method can be scored on whether it finds what was planted.

What happened

  • The parity suite runs the 2021 Python on the synthetic inputs and the ports match: all 43 assay, control and measurement charts reproduce the notebook's order, seven limit lines and printed out-of-control list.
  • Over 40 synthetic worlds the parser handles 99.74% of rows (95% bootstrap CI 99.69% to 99.78%; range 99.33% to 100%). That is close to the 2021 figure of about 99.7% by construction: the generator's junk-row rates were chosen to give it. It shows the parser's rules work on the formats I invented, not that they would work on new real formats. This is the weakest number on the site and I say so wherever it appears.
  • The planted faults were made findable (control shifts of 5 to 6 sigma, degraded samples 8 to 11 spreads away). Recovery rates on /evaluation are therefore upper bounds for the methods, not estimates of how often a real lab would catch a real fault.
  • Nothing from the partner is in this site's code, data or content: every value shown is generated from the seed.

What I'd change

  • Calibrate the generator's spreads and fault sizes against published HPLC system-suitability ranges instead of my own guesses, and add autocorrelation within runs, which real controls often show.
  • Generate a second, harder world (smaller faults, more realistic typing errors) and report the parse rate and fault recovery on both, so the easy case is not the only one.

DR-002 · Accepted · 2026-10-09

The run-rule set, and what limits to judge it against

Decision: Keep the 2021 3-sigma rule as the default view, offer "3-sigma + 8 on one side" as the one recommended addition, show the in-control false-alarm cost next to every rule-set preset, and do not recommend the trend rule or the extra Nelson rules; for monitoring new results, set limits from a Phase I window instead of from every point on the chart.

Context

The 2021 notebook flagged only points beyond 3 sigma, with sigma the sample standard deviation of every point on the chart. The revival added Western Electric and Nelson run rules as checkboxes, which made it easy to switch on all eight. Nothing on the site said what that costs: each rule adds false alarms, and a QC chart that cries wolf trains people to ignore it. The team had also spotted a run of six rising amino acid results by eye, which is exactly the trend rule (N3), so there was a case for adding it.

Decision

  • Default view on /control-charts: the 2021 rule (3 sigma only), so the page still shows what the team did.
  • Presets: "2021 notebook", "3-sigma + 8 on one side" (recommended addition), "Western Electric" (rules 1 to 4) and "Western Electric + Nelson" (all eight). Each preset shows its simulated false-alarm probability for a 150-point in-control chart.
  • Not recommended: the trend rule on its own, and the four Nelson rules beyond Western Electric.
  • Limits: the 2021 all-points limits stay the default view, but the page links to the masking study, and the Phase I switch (limits from the first 40 points) is the recommended way to judge later results.

Options considered

  1. 3 sigma only (the 2021 notebook). Fewest false alarms, slow on small sustained shifts.
  2. 3 sigma + 8 on one side (chosen as the addition).
  3. 3 sigma + trend of 6 (what the team saw by eye).
  4. Western Electric rules 1 to 4.
  5. All eight Western Electric and Nelson rules.
  6. A CUSUM or EWMA chart instead of run rules. Better suited to small shifts, but a different chart from the one the team built.

Why

From /evaluation (4,000 simulated streams per scenario, every rule set on the same streams, 95% intervals):

Rule setIn-control ARLFalse alarm in 150 pointsARL after a 1σ shiftRatio to 3σ only
3σ only361.9 (351.5 to 372.5)33.5% (32.0% to 34.9%)45.1 (43.8 to 46.6)baseline
3σ + 8 on one side151.5 (147.0 to 156.0)62.5% (61.0% to 64.0%)14.6 (14.3 to 14.9)0.32 (0.31 to 0.33)
3σ + trend of 6198.0 (192.4 to 204.0)52.5% (51.0% to 54.0%)41.2 (39.9 to 42.4)0.91 (0.90 to 0.93)
Western Electric 1 to 491.3 (88.7 to 94.0)80.8% (79.6% to 82.0%)9.5 (9.3 to 9.7)0.21 (0.20 to 0.22)
All 8 rules66.0 (64.1 to 67.7)90.3% (89.3% to 91.2%)9.4 (9.2 to 9.6)0.21 (0.20 to 0.22)
  • "8 on one side" gives most of the detection gain for sustained shifts (about three times faster at 1 sigma) for 2.4 times the false-alarm rate.
  • The trend rule costs nearly as many false alarms and buys almost nothing, even against the drifts it is meant for: at 0.05 sigma per point its ratio to the 3-sigma rule is 0.97 (0.97 to 0.98). Noise breaks long monotone runs before they reach six points.
  • The four extra Nelson rules detect nothing faster than Western Electric 1 to 4 here (9.4 against 9.5 points at 1 sigma) and cut the in-control ARL from 91.3 to 66.0.
  • A preset that shows its cost lets a visitor make the trade-off knowingly instead of ticking every box.

What happened

  • The simulation engine reproduces exact values: the 3-sigma rule's in-control ARL interval covers the theoretical 370.4, and "8 on one side" alone gives 255 (2⁸ − 1) in the tests. The published 3-sigma estimate, 361.9, sits near the bottom of its interval; it is a Monte Carlo estimate, not a disagreement with theory.
  • The bigger problem turned out to be the limits, not the rules. With the 2021 all-points limits, a sustained 3-sigma shift in the last third of a 150-point chart is caught in only 2.5% (2.0% to 3.0%) of charts, because the shift inflates the sample SD and drags the centre line. Limits from the first 40 points catch it in 100% (99.9% to 100%). At 1.5 sigma the paired gain is 75.8 percentage points (74.5 to 77.0).
  • Phase I limits have their own cost: with 40 points the limits are noisy, and 27.2% (25.8% to 28.6%) of in-control charts flag one of the first 100 points, against 21.7% (20.5% to 23.0%) for all-points limits. With 20 Phase I points, 41.1% (39.8% to 42.5%) of analysts get a chart that false-alarms at least twice as often as the textbook chart.
  • Across the synthetic worlds, the 3-sigma rule flags 2% of clean runs and Western Electric 1 to 4 flags 18%, because each control has up to 11 measurements charted at once.
  • Weak spot: all of this assumes independent, normal results with no autocorrelation. Real HPLC controls can violate all three.

What I'd change

  • Add an EWMA chart for small sustained shifts, compared on the same simulated streams, instead of stacking more run rules on a Shewhart chart.
  • Control the family-wise false-alarm rate across the measurements of one control (for example a Bonferroni-style widening, or a multivariate T² chart) and simulate it.
  • Simulate the run rules with estimated limits too; the Phase I study currently covers the 3-sigma rule only.

DR-003 · Accepted · 2026-10-09

Clustering as a review aid, not a decision

Decision: Present clustering output only as review prompts for a person (never as a label, a pass or fail, or an input to an automatic decision), show how unstable those prompts are, and keep DBSCAN outliers and run-rule signals as separate views instead of combining them into one score.

Context

The 2021 team compared PCA, t-SNE and UMAP projections with K-Means and DBSCAN, scored by silhouette, and concluded modestly that without labels clusters can only help surface unusual runs. The revival turned that into "review prompts": DBSCAN noise, tiny clusters and points far from their K-Means centre. Two questions were open. How repeatable is a list of review prompts? And does DBSCAN point at the same runs as the control charts, so that the two could be combined into one "suspicious run" score?

Decision

  • /odd-runs keeps the wording "prompts for a person to check, not verdicts" and links to the stability results.
  • No clustering output is turned into a label, a disposition or an automatic action anywhere on the site, and the optional AI feature is not given clustering output to judge.
  • DBSCAN outliers and run-rule signals stay separate views; the site does not merge them.
  • A high silhouette is not presented as evidence that a pairing is good for finding odd runs.

Options considered

  1. Clusters as labels (for example "normal" and "abnormal" groups). Rejected: no labels exist to validate them, and the 2021 team rejected it too.
  2. A combined anomaly score from DBSCAN noise and run-rule signals. Rejected after the agreement study below.
  3. Review prompts with stability shown (chosen).
  4. Drop clustering from the site. Rejected: it was a third of the 2021 brief, and showing its limits honestly is useful.

Why

From /evaluation (the whole /odd-runs procedure re-run at the reference seed and 20 other algorithm seeds, and on 50 bootstrap resamples; 95% bootstrap intervals for the means):

  • Algorithm seed alone changes the answer for some pairings. PCA + K-Means on the amino acid assay chose k = 15 at the reference seed and 4, 10, 11, 12 or 13 at the other 20 seeds. Agreement with the reference has mean ARI 0.59 (0.51 to 0.67), but the mean hides two outcomes: ARI 0.28 at the 5 seeds that chose k = 4 and about 0.70 at the rest.
  • Resampling the injections moves the prompts even where the partition is stable. Over 50 bootstrap resamples, PCA + DBSCAN on Protein A has ARI 0.42 (0.31 to 0.55) because the chosen eps jumps between about 0.4 and 1.5; on cation exchange the partition holds (ARI 0.92, 0.89 to 0.95) but the overlap of flagged injections is 0.68 (0.65 to 0.71).
  • The most repeatable pairings are the least useful for review: UMAP + K-Means and t-SNE + K-Means reproduce their partitions across seeds (mean ARI 0.96 to 1.00) and flag no injection at any seed, because they fold the odd samples into tidy clusters with a high silhouette.
  • DBSCAN and the run rules do not agree: Cohen's kappa between the 3-sigma rule and DBSCAN, over 4,800 runs in 20 synthetic worlds, is −0.01 (−0.03 to 0.01). The rules find runs with planted control faults (69%, 67% to 71%); DBSCAN finds runs with planted degraded samples (82%, 77% to 87%). A combined score would average two detectors that look at different injections.

What happened

  • The stability and agreement tables are on /evaluation and are re-checked by the test suite against the published artefact.
  • The team's own pairing for the amino acid assay, t-SNE + DBSCAN, came out well: ARI 0.97 (0.94 to 1.00) across seeds with prompt overlap 0.92 (0.85 to 0.97). For cation exchange the same pairing keeps its partition (ARI 0.99) but flags nothing at the reference seed and two to four injections at 10 of the 20 other seeds, so its prompt overlap is 0.50 (0.30 to 0.70).
  • A first version of this study used 7 other seeds. Its percentile intervals were too narrow for so few replicates (the amino acid PCA + K-Means figure read 0.52, 0.34 to 0.64), so the published study uses 20.
  • DBSCAN also flags 8% (6% to 9%) of runs with no planted problem: a missing value filled with 0 looks as odd as a degraded sample.
  • Limitation: data resampling was only run for the PCA pipelines (t-SNE and UMAP refits cost seconds each), so their stability under new data is not measured.

What I'd change

  • Report a per-injection "flag frequency" across seeds and resamples (how often each injection is prompted) instead of one prompt list, so a reviewer sees which flags are robust.
  • Replace the "where the curve settles" heuristic with a stability-based choice of k or eps (pick the setting whose partition is most repeatable), and compare.
  • Run the bootstrap for t-SNE and UMAP in a Web Worker pool, offline, to close the gap above.

DR-004 · Accepted · 2026-10-09

Optional bring-your-own-key "Explain this chart", grounded in the chart's numbers

Decision: Offer an optional "Explain this chart" that runs only with the visitor's own Anthropic or OpenAI key, called straight from the browser with nothing but the chart's numeric summary, labelled AI-generated, checked for invented numbers, recorded in a local audit log, and left to the visitor to accept, edit or reject.

Context

Control charts are easy to draw and easy to misread: a first-time reader does not know what "8 on one side" means, or that a chart with 150 points will often show a false alarm. A short plain-English reading of the chart on screen helps. The project has no budget for AI and no server: the site is static on Vercel's free tier. Any AI use also had to be transparent about what is sent, where it goes and what a person decides, informed by the Australian Government's policy for the responsible use of AI in government, the transparency principles of the EU AI Act and the NIST AI Risk Management Framework. Being informed by them is not a claim of compliance with any of them.

Decision

  • The feature is optional. Every page works without it, and nothing on the site depends on its output.
  • Bring your own key: Anthropic (default, Claude Haiku 4.5, with Claude Sonnet 5.5 as an option) or OpenAI (model id editable). The key is kept in sessionStorage by default, in localStorage only if the visitor ticks "remember on this device", and can be forgotten with one button. It is sent only in the request header to the provider, never to this site, and never logged.
  • Input: only the numeric summary of the chart on screen (what is charted, limits and how they were set, active rules and their signals, injected disturbances, run statistics, and the published false-alarm figures for the active rule set). No raw series, nothing about the visitor.
  • Output: structured JSON validated with zod, shown under an "AI-generated" label with the model, latency and token usage. A grounding check lists numbers in the reply that are not in the summary.
  • Human in the loop: the visitor accepts, edits or rejects each explanation. Every call, including failures, is written to an audit log in the visitor's IndexedDB (input without the key, output, provider, model, latency, token usage, decision), viewable and exportable as JSON or CSV at /ai-log.
  • The system prompt forbids release or disposition language: a signal is a prompt to investigate, never proof of a cause.

Options considered

  1. No AI. Simplest, but the reading of a chart stays locked behind SPC jargon.
  2. A site-funded proxy with my key. Rejected: no budget, and a public endpoint would need abuse controls.
  3. Server-side calls with the visitor's key. Rejected: the key would pass through this site's server, which is exactly what visitors should not have to trust.
  4. Bring your own key, browser-direct (chosen).
  5. A model running in the browser (WebGPU). Rejected for now: large downloads, uneven device support and weaker explanations.

Why

Browser-direct calls keep the key between the visitor and their provider. Sending only the numeric summary makes the input small (about 3,400 characters with the system prompt, roughly a thousand tokens), auditable and free of anything personal. Structured output plus the grounding check turns "the model should only use these numbers" from a hope into something visible on screen and in the log.

What happened

  • The client, both provider adapters, error mapping, key storage, redaction, the audit log (memory and IndexedDB) and the grounding check are unit-tested with the network mocked; no test calls a real provider.
  • The Anthropic request sends anthropic-dangerous-direct-browser-access: true, structured output through output_config.format, low effort only on Sonnet (Haiku 4.5 rejects the effort setting), and the server-side refusal fallback only on Sonnet 5.5.
  • What I have not measured: the quality of the explanations. With bring-your-own-key there is no budget to run an evaluation set, so there is no accuracy number to report. The grounding check catches invented numbers but not a wrong reading built from correct numbers.
  • The audit log lives in one browser. Clearing site data deletes it, and nothing is kept server-side by design.

What I'd change

  • Build a small evaluation set (charts with known correct readings, scored by a rubric), run it once with a key, and publish the results with intervals on /evaluation, with the AI explanation compared against a template-based explanation generated from the same summary.
  • Add a check that the explanation names the right rule for each listed signal, not only the right numbers.

Model card: control-chart run rules and the clustering review aid

The "models" on this site are two statistical procedures ported from a 2021 student capstone, not trained predictive models: a Shewhart individuals control chart with optional run rules, and an unsupervised clustering pipeline that turns HPLC injections into review prompts. This card follows the usual model-card headings so their intended use and limits are stated in one place. Numbers come from the published study artefact (study seed 30034), shown live with their intervals on /evaluation.

Model details

  • Control chart. Individuals chart per control injection and measurement. Centre line and sigma from the mean and sample standard deviation of every point (the 2021 notebook's create_control_chart), or from the average moving range divided by 1.128 (I-MR), optionally from a Phase I window of the first 40 points. Signals: points beyond 3 sigma, plus optional Western Electric rules 1 to 4 and Nelson rules 3, 4, 7 and 8.
  • Clustering review aid. The FA4 preparation (drop controls and identifier columns, fill blanks and text with 0, standard-scale), a PCA, t-SNE or UMAP projection, K-Means or DBSCAN with k or eps chosen where the silhouette curve settles, then review prompts: DBSCAN noise, members of clusters of one or two, and K-Means points beyond Q3 + 3 IQR of their cluster's distances.
  • Provenance. Ports of the team's Python (pandas, scikit-learn 0.23, umap-learn), checked against the original code on the same synthetic inputs. Run rules, I-MR limits and Phase I windows were added in the 2026 revival.

Intended use

  • Teaching and portfolio: showing how control charts and clustering behave on HPLC-shaped data, and what their outputs can and cannot support.
  • Prompting a person to look at a run or an injection. Every output is a prompt for review.

Out of scope

  • Any decision about a real batch, product, run or result (release, rejection, disposition, investigation closure).
  • Real or regulated data. The site runs only on synthetic data (see the data card).
  • Using cluster membership as a label or a quality grade.

Training and reference data

There is no training in the machine-learning sense. Limits are estimated from the chart's own points (or a Phase I window) of synthetic control injections; clustering is fitted to the synthetic sample injections of one assay at a time (300 per assay at the default settings). See DR-001.

Evaluation

All figures are Monte Carlo estimates with 95% intervals (Wilson for proportions, percentile bootstrap otherwise).

QuestionResult
In-control ARL, 3-sigma rule, known limits361.9 (351.5 to 372.5); theory 370.4
False alarm in a 150-point in-control chart3σ only 33.5% (32.0% to 34.9%); all 8 rules 90.3% (89.3% to 91.2%)
ARL after a 1-sigma shift3σ only 45.1 (43.8 to 46.6); 3σ + 8 on one side 14.6 (14.3 to 14.9)
Phase I of 20 points: charts at least twice as noisy as the textbook chart41.1% (39.8% to 42.5%)
Sustained 3-sigma shift in the last third, caught by the 2021 all-points limits2.5% (2.0% to 3.0%); first-40-points limits 100% (99.9% to 100%)
Agreement of DBSCAN outliers with 3-sigma signals (Cohen's kappa, runs)−0.01 (−0.03 to 0.01)
Runs with planted control faults found by the 3-sigma rule69% (67% to 71%)
Runs with planted degraded samples found by DBSCAN82% (77% to 87%)
Seed stability, PCA + K-Means on the amino acid assay (mean ARI to reference over 20 seeds)0.59 (0.51 to 0.67); 0.28 at the 5 seeds that chose k = 4, about 0.70 at the rest

The run-length engine is checked against exact values (1 / P(|Z + δ| > 3) for the 3-sigma rule, 255 for "8 on one side" alone), and the statistics helpers against SciPy, statsmodels, scikit-learn and R.

Known failure modes

  • Masking by retrospective limits. Limits from every point, including shifted ones, inflate sigma and hide sustained shifts (DR-002).
  • False alarms from many rules and many charts. Each added rule and each extra measurement charted raises the chance that something signals on an in-control process.
  • Estimated limits. Short Phase I windows give each analyst a chart with a very different false-alarm rate.
  • Assumptions. Independent, normal results. Autocorrelation within a run, skew or heavy rounding change every rate above.
  • Clustering instability. The chosen k or eps and the list of review prompts can change with the random seed or a resample; high-silhouette pairings (UMAP or t-SNE with K-Means) can flag nothing at all (DR-003).
  • Filled zeros. A blank or text result filled with 0 can look as odd as a degraded sample.

Ethical considerations

  • The original data were confidential and are not used; the synthetic generator was written from scratch (DR-001).
  • Outputs are framed as prompts so they cannot be mistaken for a quality decision. The optional AI explanation is told never to use release or disposition language, and a person accepts, edits or rejects it (DR-004).
  • Planted-fault recovery rates are only possible because the data are synthetic and are upper bounds for the methods.

Caveats and recommendations

  • For monitoring new results, set limits from a clean Phase I window that is as long as practical (with 100 points, 17.9% of charts are still at least twice as noisy as the textbook chart), and add at most "8 on one side" to the 3-sigma rule.
  • Read clustering prompts as one draw from a noisy procedure: prefer injections that are flagged across seeds and resamples.

Data card: the synthetic lab world

Every sheet, chart and cluster on this site comes from one seeded generator in web/src/lib/synthetic. No industry-partner data, code, documents or business rules are used (DR-001).

Composition

PartDefault sizeWhat it contains
Lab-results sheet1,200 rows in two workbooksFree-text department, experiment and lot fields for four invented process streams, each with a canonical format plus five to eight realistic ways of typing it wrong
HPLC assay exports4 assays × 60 sequences, July 2019 to August 2021Amino acid excipient, cation exchange, Protein A and size exclusion; bracketing control injections plus the sample injections submitted on the sheet
Control injectionsabout 150 per control and measurementOne or two control types per assay; 5 to 12 measurement columns each, one of them the constant injection volume
Sample injections300 per assayValues grouped by stream and sub-step profiles, so clusters exist

How it is generated

  • One integer seed (default 30034) drives a mulberry32 generator; named sub-streams keep parts independent, so the same seed always gives the same world in the browser, the tests and the Python parity scripts.
  • Measurements are normal around invented means, with a per-sequence shock (sd 0.35 of the measure's sd) so runs differ a little, and derived columns (relative areas, amounts, concentrations) computed from base ones.
  • 85% of a stream's samples go to runs of that stream, so an out-of-control run can be traced back to one or two streams through Sample Name.

Planted problems (the ground truth)

  • Run faults: control injections shifted by 5 to 6 sigma in named sequences (on several measurements for cation exchange and size exclusion), and a seven-point upward drift inside the limits in the amino acid assay.
  • Degraded samples: about 2.5% of sample injections pushed 8 to 11 spreads away on two measurements.
  • Messiness: about 0.6% of sheet rows miss a key field and about 0.3% are unparseable; about 1.2% of control injections and 0.6% of sample injections have one blank measurement; Protein A has "ND" and "Overloaded" text results.

Known gaps

  • No autocorrelation within runs, no instrument-specific drift, no seasonal effects.
  • Fault sizes were chosen to be findable, so recovery rates are optimistic.
  • The parse rate (99.74% across 40 worlds, 95% CI 99.69% to 99.78%) matches the 2021 figure because the junk-row rates were tuned to it; it does not validate the parser on new real formats.

Use and licence

Synthetic and free of personal or confidential information. CSV downloads on /spreadsheet are generated in the browser from the seed. The repository has no licence file: all rights reserved by the authors.

AI use statement

One optional feature uses a large language model: “Explain this chart” on /control-charts. It is informed by the Australian Government's policy for the responsible use of AI in government, the transparency principles of the EU AI Act and the NIST AI Risk Management Framework; that is not a claim of compliance with any of them. The decision is recorded in DR-004.

What the AI does

  • Writes a short plain-English reading of the control chart on screen: what is charted, how the limits were set, what the active rules flagged, and how often that rule set false-alarms.
  • Suggests up to three checks a person could make next.

What it never does

  • Never runs without your own key, and nothing on the site depends on it.
  • Never decides anything: no release, rejection or disposition language, no change to the chart, no input to another feature.
  • Never sees the raw series, the spreadsheet, the clustering, or anything about you.
  • Never sends your key, or anything else, to this site.

What is sent, and where

  • A system prompt and a JSON summary of the chart with these fields, and nothing else:
    • page: the page name (control charts, synthetic data)
    • assay: the assay
    • control_injection: the control injection charted
    • measurement: the measurement charted
    • data_seed: the data seed
    • points_on_chart: the number of points on the chart
    • points_skipped_blank_or_text: how many cells were skipped as blank or text
    • limits: the centre line, sigma and 3-sigma limits, and which points and which sigma estimate set them
    • active_rules: the active run rules, each with its signal count
    • beyond_3_sigma: how many points are beyond 3 sigma, and up to 10 of them (position, sequence name, value, z)
    • other_rule_signals: how many other rule signals there are, and up to 10 of them (rule, position, sequence name)
    • injected_by_visitor: any shift, drift or spike you injected
    • run_statistics: four run statistics: the longest run on one side of the centre line, the longest steady rise or fall, and the largest and smallest z
    • sequences_beyond_3_sigma_on_several_measurements: up to 5 sequence names beyond 3 sigma on two or more measurements, with how many
    • reference_false_alarm_figures: the published false-alarm figures for the active rule set: in-control average run length and the chance of at least one false alarm, with 95% intervals
  • It goes from your browser straight to api.anthropic.com (Anthropic Messages API, default Claude Haiku 4.5, or Claude Sonnet 5.5) or api.openai.com (default gpt-5-mini, editable), under that provider's terms for your key.
  • Your key stays in this browser: sessionStorage by default, localStorage only if you tick “remember on this device”. “Forget Anthropic key” or “Forget OpenAI key” removes that provider’s key and “Forget all keys” removes both.

Human in the loop and audit

  • Every reply is labelled AI-generated, with the model, latency and token usage.
  • A grounding check lists any number in the reply not found in the summary. It checks numbers with decimals, numbers above 10 and anything in scientific notation, sign included (a number without a minus sign matches a negative value only beside a word such as “below” or “drop”); a number followed by % is matched as a proportion and only as a proportion; whole numbers from 0 to 10 are not checked.
  • You accept, edit or reject each explanation; rejected ones are hidden. A reply you replace by asking again before deciding is logged as superseded, not left pending.
  • Every call, failed ones included, is recorded in your browser's IndexedDB with the input (never the key), output, provider, model, latency, token usage and your decision, under the feature name explain-control-chart. View or export it as JSON or CSV at /ai-log.

The reference false-alarm figure the AI receives for the 2021 rule, for example, is 33.5% (95% CI 32.0% to 34.9%) for a 150-point chart: the same number shown on /evaluation.