Skip to content
HPLC QC LabMAST30034 · revival

04 · 2026 upgrade · Evaluation

How far can the charts and clusters be trusted?

The 2021 team drew control charts and clusters. This page asks how those methods behave: how often a run-rule set raises a false alarm and how quickly it catches a real change, what happens when the limits are estimated from a short history, whether a chart can hide the shift it should catch, and whether the clusters survive a new random seed or a resampled data set.

Every number is a Monte Carlo estimate with a 95% interval (Wilson for proportions, order statistics for medians, percentile bootstrap otherwise), from fixed seeds (study seed 30034). Everything runs on synthetic data, so these figures describe the methods, not any real process or product.

Run rules: false alarms against detection

Each rule set read the same 4,000 simulated streams of individual results per scenario (known centre and sigma, the change present from the first point), so differences between sets are paired. A run length is the number of points up to the first signal.

3-sigma only, in control

362

average run length [351, 373]; theory 370.4

All 8 rules, in control

66

average run length [64, 68]

False alarm in 150 points

33% → 90%

3σ only → all 8 rules, chart in control

1-sigma shift, average run length

45 → 15

3σ only → 3σ + 8 on one side

Adding rules buys speed with false alarms. On a chart that is in control, the 2021 rule raises at least one false alarm in 150 points 33.5% (95% CI 32.0% to 34.9%) of the time; Western Electric rules 1 to 4 raise one 80.8% (95% CI 79.6% to 82.0%) of the time and all eight rules 90.3% (95% CI 89.3% to 91.2%). Adding only “8 on one side” cuts the average run length after a 1σ shift to 14.6 (95% CI 14.3 to 14.9) points from 45.1, a ratio of 0.32 (95% CI 0.31 to 0.33).

The trade-offEach point is one rule set; lines are 95% Wilson intervals.
False-alarm probability in a 150-point in-control chart against the probability of detecting a 1 sigma shift within 10 points, for five run-rule sets.better: few false alarms, fast detection0%0%25%25%50%50%75%75%100%100%P(false alarm within 150 points), in controlP(signal within 10 points of a 1σ shift)3σ only3σ + WE43σ + N3WE 1 to 4All 8
Average run length by size of shiftLog scale. Long in control is good; short after a shift is good. Whiskers are 95% bootstrap intervals.
Average run length against shift size in sigma, log scale, for five run-rule sets, with 95% intervals.1251020501002005000σ0.5σ1σ1.5σ2σ2.5σ3σShift (σ)Average run length (points)3σ theory, in control: 370
  • 3σ only
  • 3σ + WE4
  • 3σ + N3
  • WE 1 to 4
  • All 8
In-control and out-of-control run lengths for each run-rule set, with 95% intervals
Rule setIn control: average run lengthIn control: median run lengthFalse alarm in 150 points1σ shift: average run length1σ shift: ratio to 3σ onlyDrift 0.05σ per point
2021 notebook (3σ only)WE1361.9 [351.5, 372.5]254.5 [244, 264]

33.5% [32.0%, 34.9%]

45.1 [43.8, 46.6]baseline30.4 [30.1, 30.7]
3σ + 8 on one sideWE1 WE4151.5 [147.0, 156.0]107 [103, 111]

62.5% [61.0%, 64.0%]

14.6 [14.3, 14.9]0.32 [0.31, 0.33]21.9 [21.7, 22.2]
3σ + trend of 6WE1 N3198.0 [192.4, 204.0]141 [135, 147]

52.5% [51.0%, 54.0%]

41.2 [39.9, 42.4]0.91 [0.90, 0.93]29.5 [29.2, 29.8]
Western Electric (rules 1 to 4)WE1 WE2 WE3 WE491.3 [88.7, 94.0]64 [62, 67]

80.8% [79.6%, 82.0%]

9.5 [9.3, 9.7]0.21 [0.20, 0.22]18.2 [17.9, 18.3]
Western Electric + Nelson (all 8)WE1 WE2 WE3 WE4 N3 N4 N7 N866.0 [64.1, 67.7]48 [46, 49]

90.3% [89.3%, 91.2%]

9.4 [9.2, 9.6]0.21 [0.20, 0.22]17.9 [17.6, 18.1]

Read the median as well as the mean. Run lengths are strongly skewed: with the 3σ rule the in-control average is about 362 points but the median is 254.5 (95% CI 244 to 264), and one chart in ten signals within 40 points (95% CI 36 to 44). An average run length on its own oversells how quiet a chart is. Median and percentile intervals are distribution-free, from order statistics.

Which rule adds what. “8 on one side” does most of the work against sustained shifts. The trend rule (6 rising or falling) barely helps against drifts: at 0.05σ per point its ratio to the 3σ rule is 0.97 (95% CI 0.97 to 0.98), while it shortens the in-control run length to 198. The four Nelson rules beyond Western Electric detect nothing faster here (1σ shift: 9.4 against 9.5) but cut the in-control run length from 91 to 66.

Many charts at once. The cation exchange assay has 11 measurements with spread. Charted together, even the 3σ rule will usually raise somewhere. The decision this informs is in DR-002.

Every scenario, every rule set
Average run length with 95% interval, every scenario
Scenario3σ only3σ + WE43σ + N3WE 1 to 4All 8
In control361.9[351.5, 372.5]151.5[147.0, 156.0]198.0[192.4, 204.0]91.3[88.7, 94.0]66.0[64.1, 67.7]
Shift of 0.25σ279.9[271.0, 289.1]94.9[92.1, 97.7]169.0[163.7, 174.0]57.8[56.1, 59.6]46.5[45.2, 47.8]
Shift of 0.5σ151.5[146.9, 156.1]45.0[43.7, 46.3]111.1[107.9, 114.6]27.3[26.5, 28.0]25.3[24.6, 26.0]
Shift of 0.75σ78.8[76.3, 81.1]22.9[22.4, 23.5]67.0[65.0, 69.0]14.8[14.4, 15.1]14.5[14.2, 14.9]
Shift of 1σ45.1[43.8, 46.6]14.6[14.3, 14.9]41.2[39.9, 42.4]9.5[9.3, 9.7]9.4[9.2, 9.6]
Shift of 1.5σ15.6[15.1, 16.0]8.0[7.8, 8.1]15.2[14.8, 15.6]5.3[5.2, 5.3]5.3[5.2, 5.3]
Shift of 2σ6.1[5.9, 6.3]4.8[4.7, 4.9]6.1[5.9, 6.3]3.4[3.3, 3.5]3.4[3.3, 3.4]
Shift of 2.5σ3.2[3.1, 3.3]3.0[3.0, 3.1]3.2[3.1, 3.3]2.4[2.4, 2.5]2.4[2.4, 2.5]
Shift of 3σ2.0[2.0, 2.1]2.0[2.0, 2.0]2.0[2.0, 2.1]1.8[1.8, 1.8]1.8[1.8, 1.8]
Drift of 0.05σ per point30.4[30.1, 30.7]21.9[21.7, 22.2]29.5[29.2, 29.8]18.2[17.9, 18.3]17.9[17.6, 18.1]
Drift of 0.1σ per point18.4[18.2, 18.6]14.7[14.6, 14.8]18.2[18.0, 18.3]12.2[12.1, 12.3]12.1[12.0, 12.2]

Limits from a short history

Real limits come from a Phase I sample of m in-control results, so every analyst's chart has its own in-control average run length. For the 3σ rule it is exact given the estimates, 1 / P(new result outside the estimated limits), so each of the 5,000 simulated Phase I samples per size gives one value with no further simulation.

With 20 Phase I points, a chart that false-alarms at least twice as often as the textbook chart is common: 41.1% (95% CI 39.8% to 42.5%) of simulated analysts get one. The chance of a false alarm in the next 100 results is 39.5% (95% CI 38.7% to 40.4%), against 23.7% with known limits. Even with 100 points, 17.9% of charts are twice as noisy; with 500, 1.4%.

In-control average run length, conditional on the Phase I sampleLine: median across simulated analysts (whiskers: 95% bootstrap interval). Band: 10th to 90th percentile. Both axes on a log scale.
Median conditional in-control average run length against Phase I sample size, for the sample standard deviation and the moving-range estimate of sigma, with 10th to 90th percentile bands.1020501002005001,0002,0005,00010,00020,00010203050100200500Phase I sample size mIn-control average run lengthknown limits: 370
  • Sample SD (2021)
  • Moving range / 1.128 (I-MR)
Conditional in-control behaviour by Phase I size and sigma estimate
mSigma fromMedian ARL₀At least twice as noisyFalse alarm in next 100
10Sample SD188 [176, 204]

49.6% [48.2%, 51.0%]

47.4% [46.4%, 48.4%]

10Moving range217 [195, 236]

47.8% [46.5%, 49.2%]

45.9% [44.8%, 47.1%]

20Sample SD259 [246, 273]

41.1% [39.8%, 42.5%]

39.5% [38.7%, 40.4%]

20Moving range255 [238, 281]

42.6% [41.2%, 44.0%]

40.4% [39.5%, 41.4%]

30Sample SD293 [280, 305]

35.4% [34.1%, 36.7%]

35.4% [34.7%, 36.1%]

30Moving range306 [292, 323]

37.6% [36.2%, 38.9%]

36.6% [35.8%, 37.4%]

50Sample SD319 [311, 329]

27.9% [26.7%, 29.2%]

31.4% [30.9%, 32.1%]

50Moving range337 [324, 353]

30.1% [28.9%, 31.4%]

32.3% [31.6%, 33.0%]

100Sample SD340 [334, 350]

17.9% [16.9%, 19.0%]

27.9% [27.5%, 28.4%]

100Moving range355 [344, 365]

23.3% [22.1%, 24.5%]

28.9% [28.4%, 29.4%]

200Sample SD356 [350, 362]

8.6% [7.8%, 9.4%]

25.9% [25.6%, 26.2%]

200Moving range359 [349, 367]

13.9% [13.0%, 14.9%]

26.7% [26.3%, 27.1%]

500Sample SD363 [360, 368]

1.4% [1.1%, 1.7%]

24.7% [24.5%, 24.9%]

500Moving range364 [358, 368]

3.4% [3.0%, 4.0%]

24.9% [24.6%, 25.1%]

“At least twice as noisy” means a conditional in-control average run length of 185.2 or less, half the known-limit value. The Phase I switch on /control-charts uses the first 40 points, between the 30 and 50 rows above. This is the effect studied by Quesenberry (1993) and reviewed by Jensen et al. (2006); the numbers here are this site's own simulation.

Retrospective limits can hide the shift

The 2021 function set the limits from every point on the chart, shifted ones included. Here each of 4,000 simulated charts has 150 points and the last 50 are shifted; four ways of setting the limits are applied to the same charts and judged with the 3σ rule only.

A sustained 3σ shift in the last third of a chart is caught by the 2021 limits in only 2.5% (95% CI 2.0% to 3.0%) of charts: the shift inflates the sample SD and drags the centre line with it. Limits from the first 40 points catch it in 100.0% (95% CI 99.9% to 100.0%). At 1.5σ the paired gain is 75.8 percentage points (95% CI 74.5 to 77.0).

Charts with a 3σ signal among the shifted pointsAt a shift of 0 this is a false alarm among the last 50 points. Whiskers are 95% Wilson intervals.
Share of charts with a 3 sigma signal among the shifted points, against shift size, for four ways of setting the limits.0%20%40%60%80%100%0σ0.5σ1σ1.5σ2σ2.5σ3σShift in the last third (σ)Charts with a signal on the shifted points
  • 2021: all points, sample SD
  • All points, moving range (I-MR)
  • First 40 points, sample SD
  • First 40 points, moving range
Detection and early signals by limit method, 1.5 sigma shift
Limits (1.5σ shift)Signal on the shifted pointsSignal before the shiftvs 2021, paired
2021: all points, sample SD

14.8% [13.7%, 15.9%]

6.7% [6.0%, 7.5%]

baseline
All points, moving range (I-MR)

68.1% [66.6%, 69.5%]

45.4% [43.8%, 46.9%]

+53.3 pp [51.7, 54.8]
First 40 points, sample SD

90.6% [89.7%, 91.5%]

27.2% [25.8%, 28.6%]

+75.8 pp [74.5, 77.0]
First 40 points, moving range

88.1% [87.1%, 89.1%]

30.3% [28.8%, 31.7%]

+73.3 pp [72.0, 74.6]

Moving-range limits from all points are not a full fix: the centre line still moves towards the shifted points, so 45.4% (95% CI 43.8% to 46.9%) of charts also flag points before the shift, which sends a reviewer to the wrong runs. Phase I limits flag early points at their own rate, 27.2%, whatever happens later, which is the cost of estimating from 40 points (section B). The planted faults in the synthetic data are mostly single-run spikes, which the 2021 limits do catch.

Do the clusters survive a new seed, or new data?

The whole /odd-runs procedure (scale, project, sweep the silhouette, pick k or eps where the curve settles, cluster, list review prompts) was re-run on the default synthetic world. Agreement with the reference run (seed 30034) is the adjusted Rand index (ARI, 1 = the same partition, 0 = chance) and the Jaccard overlap of the injections flagged for review.

High-silhouette pairings are the most repeatable and the least useful for review: UMAP or t-SNE with K-Means reproduce their partition across seeds but flag no injection at all. Where the procedure does flag injections, which ones it flags moves with the seed and the sample (prompt overlap well below 1), and the chosen k or eps can jump between runs.

Same data, 20 other algorithm seeds (1 to 20) compared with seed 30034. Means with 95% bootstrap intervals across the 20 seeds.
AssayPipelineReference runARI to referencePrompt overlap (Jaccard)Silhouettek or eps chosen at the other seeds
AAEPCA + K-Meansk = 154 prompts

0.59 [0.51, 0.67]

0.52 [0.39, 0.63]

0.56 [0.56, 0.57]

4 ×5, 10 ×5, 11 ×8, 12, 13
AAEPCA + DBSCANeps = 2.115 promptsNo random step: every seed gives the same answer.
AAEt-SNE + DBSCAN2021 pickeps = 2.436 prompts

0.97 [0.94, 1.00]

0.92 [0.85, 0.97]

0.72 [0.71, 0.73]

1.56 to 2.58
CEXPCA + K-Meansk = 46 prompts

0.98 [0.94, 1.00]

1.00 [1.00, 1.00]

0.64 [0.64, 0.65]

3 ×2, 4 ×18
CEXPCA + DBSCANeps = 0.3811 promptsNo random step: every seed gives the same answer.
CEXt-SNE + DBSCAN2021 pickeps = 1.730 prompts

0.99 [0.99, 0.99]

0.50 [0.30, 0.70]

0.85 [0.84, 0.86]

1.22 to 1.89
CEXUMAP + K-Means2021 pickk = 70 prompts

1.00 [1.00, 1.00]

flags nothing in any run

0.94 [0.93, 0.94]

7 ×20
Protein APCA + K-Meansk = 57 prompts

0.96 [0.94, 0.97]

0.93 [0.90, 0.95]

0.56 [0.56, 0.57]

5 ×20
Protein APCA + DBSCANeps = 0.2821 promptsNo random step: every seed gives the same answer.
Protein AUMAP + K-Means2021 pickk = 30 prompts

0.96 [0.91, 1.00]

flags nothing in any run

0.85 [0.84, 0.86]

2, 3 ×18, 4
SECPCA + K-Meansk = 53 prompts

0.91 [0.90, 0.92]

0.95 [0.88, 1.00]

0.67 [0.67, 0.67]

5 ×3, 6 ×17
SECPCA + DBSCANeps = 0.736 promptsNo random step: every seed gives the same answer.
SECt-SNE + K-Means2021 pickk = 50 prompts

0.99 [0.96, 1.00]

flags nothing in any run

0.83 [0.83, 0.83]

5 ×19, 6
SECUMAP + K-Means2021 pickk = 80 prompts

0.97 [0.94, 0.99]

flags nothing in any run

0.92 [0.91, 0.92]

6, 7 ×3, 8 ×16
Same seed, 50 bootstrap resamples of the injections (distinct rows, about 63% of them each), the whole procedure refitted. PCA pipelines only: a t-SNE or UMAP refit costs seconds each.
AssayPipelineARI to referencePrompt overlap (Jaccard)Silhouettek or eps chosen (range)
AAEPCA + K-Means

0.60 [0.56, 0.64]

0.43 [0.37, 0.48]

0.58 [0.57, 0.59]

4 to 16
AAEPCA + DBSCAN

0.95 [0.92, 0.97]

0.92 [0.88, 0.95]

0.69 [0.68, 0.69]

1.58 to 2.55
CEXPCA + K-Means

0.90 [0.88, 0.92]

0.94 [0.90, 0.97]

0.64 [0.63, 0.65]

3 to 9
CEXPCA + DBSCAN

0.92 [0.89, 0.95]

0.68 [0.65, 0.71]

0.68 [0.67, 0.68]

0.42 to 0.75
Protein APCA + K-Means

0.78 [0.75, 0.82]

0.53 [0.46, 0.60]

0.60 [0.59, 0.61]

3 to 9
Protein APCA + DBSCAN

0.42 [0.31, 0.55]

0.42 [0.36, 0.48]

0.69 [0.68, 0.71]

0.36 to 1.59
SECPCA + K-Means

0.86 [0.82, 0.89]

0.37 [0.32, 0.42]

0.67 [0.66, 0.67]

3 to 7
SECPCA + DBSCAN

0.81 [0.73, 0.88]

0.83 [0.76, 0.90]

0.60 [0.60, 0.61]

0.44 to 1.07

Do DBSCAN outliers and run-rule signals agree?

The unit is a sequence (one run). A sequence is rule-flagged when a run rule signals on one of its control injections, on any measurement, with the 2021 limits; it is DBSCAN-flagged when one of its sample injections is DBSCAN noise in the default PCA pipeline. Over 20 synthetic worlds (4,800 sequences), with planted faults as the truth. Intervals resample whole worlds, because sequences in one world share their limits and their clustering.

They hardly agree, and they should not be expected to: Cohen's kappa between the 3σ rule and DBSCAN is −0.01 (95% CI −0.03 to 0.01). The 3σ rule finds 69% of the runs with planted control faults (95% CI 67% to 71%); DBSCAN finds 82% of the runs with planted degraded samples (77% to 87%). They look at different injections.

All assays, 3σ rule against DBSCANSequences, summed over worlds
Two by two table of rule flags against DBSCAN flags
DBSCAN flagsDBSCAN does not
Rule flags32191
Rule does not7683809

Of the sequences the rule flags, DBSCAN also flags 14% (95% CI 10% to 18%).

Agreement and planted-fault recovery by assay
AssayKappa, 3σ vs DBSCANFault runs foundDegraded-sample runs foundClean runs flagged
AAE

−0.03 [−0.07, 0.02]

3σ27% [23%, 31%]WE 1–438% [31%, 45%]DBSCAN8% [3%, 13%]3σ3% [0%, 6%]WE 1–416% [9%, 23%]DBSCAN76% [60%, 90%]3σ2% [1%, 3%]WE 1–413% [11%, 15%]DBSCAN4% [2%, 6%]
CEX

0.00 [−0.04, 0.04]

3σ100% [100%, 100%]WE 1–4100% [100%, 100%]DBSCAN23% [10%, 35%]3σ5% [2%, 9%]WE 1–422% [17%, 28%]DBSCAN73% [62%, 82%]3σ3% [2%, 4%]WE 1–423% [21%, 25%]DBSCAN8% [6%, 10%]
Protein A

−0.01 [−0.04, 0.02]

3σ100% [100%, 100%]WE 1–4100% [100%, 100%]DBSCAN20% [10%, 30%]3σ4% [2%, 7%]WE 1–411% [7%, 16%]DBSCAN99% [98%, 100%]3σ1% [1%, 2%]WE 1–413% [11%, 15%]DBSCAN12% [8%, 16%]
SEC

−0.01 [−0.05, 0.04]

3σ100% [100%, 100%]WE 1–4100% [100%, 100%]DBSCAN13% [3%, 23%]3σ4% [1%, 7%]WE 1–423% [15%, 32%]DBSCAN81% [74%, 88%]3σ1% [0%, 2%]WE 1–422% [19%, 24%]DBSCAN8% [6%, 11%]
All assays

−0.01 [−0.03, 0.01]

3σ69% [67%, 71%]WE 1–474% [70%, 77%]DBSCAN14% [9%, 18%]3σ4% [2%, 6%]WE 1–418% [15%, 21%]DBSCAN82% [77%, 87%]3σ2% [1%, 2%]WE 1–418% [16%, 19%]DBSCAN8% [6%, 9%]

The amino acid assay is the exception for the rules: most of its planted fault runs carry a slow drift inside the limits, which the 3σ rule misses. Over 20 worlds it finds 24 of the 89 fault runs, 27% (95% CI 23% to 31%). Western Electric rules 1 to 4, read on the same runs, find 34, 38% (95% CI 31% to 45%), but flag 13% (95% CI 11% to 15%) of the assay's clean runs, against 2% for the 3σ rule. Over all four assays they flag 18% and 2% of clean runs. “Planted” recall is only possible because the data are synthetic; it says nothing about real faults.

How it was computed

  • Run lengths: 4,000 streams per scenario, cut at 20,000 points (no stream reached it); intervals from 2,000 bootstrap resamples; ratios by paired bootstrap on the same streams. The engine is checked against exact values: 1 / P(|Z + δ| > 3) for the 3σ rule and 2⁸ − 1 = 255 for “8 on one side” alone.
  • Estimated limits: 5,000 Phase I samples for each size, both sigma estimates computed on the same samples.
  • Retrospective limits: 4,000 charts of 150 points, the same noise for every shift size and limit method.
  • Clustering: the default world (seed 30034); 21 algorithm seeds (the reference and 20 more); 50 bootstrap resamples per PCA pipeline. Agreement: 20 worlds (seeds 30034 and 1 to 19).
  • All of it is produced by pnpm study into one JSON artefact (274 s on a laptop). The test suite re-runs a slice of every study with the same seeds and requires identical results, and checks the statistics helpers against SciPy, statsmodels, scikit-learn and R.

What these numbers cannot tell you

  • The run-length studies assume independent, normally distributed results. Real HPLC controls can be autocorrelated within a run, skewed, or rounded, and all three change false-alarm rates.
  • The clustering and agreement studies use the synthetic generator, whose faults were planted to be findable. They show how the methods behave, not how often a real lab would catch a real fault.
  • Western Electric and Nelson rules with estimated limits were not simulated; section B covers the 3σ rule only.
  • More in the methods notes, decision records and model card.