04 · 2026 upgrade · Evaluation
How far can the charts and clusters be trusted?
The 2021 team drew control charts and clusters. This page asks how those methods behave: how often a run-rule set raises a false alarm and how quickly it catches a real change, what happens when the limits are estimated from a short history, whether a chart can hide the shift it should catch, and whether the clusters survive a new random seed or a resampled data set.
Every number is a Monte Carlo estimate with a 95% interval (Wilson for proportions, order statistics for medians, percentile bootstrap otherwise), from fixed seeds (study seed 30034). Everything runs on synthetic data, so these figures describe the methods, not any real process or product.
Run rules: false alarms against detection
3-sigma only, in control
362
average run length [351, 373]; theory 370.4
All 8 rules, in control
66
average run length [64, 68]
False alarm in 150 points
33% → 90%
3σ only → all 8 rules, chart in control
1-sigma shift, average run length
45 → 15
3σ only → 3σ + 8 on one side
Adding rules buys speed with false alarms. On a chart that is in control, the 2021 rule raises at least one false alarm in 150 points 33.5% (95% CI 32.0% to 34.9%) of the time; Western Electric rules 1 to 4 raise one 80.8% (95% CI 79.6% to 82.0%) of the time and all eight rules 90.3% (95% CI 89.3% to 91.2%). Adding only “8 on one side” cuts the average run length after a 1σ shift to 14.6 (95% CI 14.3 to 14.9) points from 45.1, a ratio of 0.32 (95% CI 0.31 to 0.33).
- 3σ only
- 3σ + WE4
- 3σ + N3
- WE 1 to 4
- All 8
| Rule set | In control: average run length | In control: median run length | False alarm in 150 points | 1σ shift: average run length | 1σ shift: ratio to 3σ only | Drift 0.05σ per point |
|---|---|---|---|---|---|---|
| 2021 notebook (3σ only)WE1 | 361.9 [351.5, 372.5] | 254.5 [244, 264] | 33.5% [32.0%, 34.9%] | 45.1 [43.8, 46.6] | baseline | 30.4 [30.1, 30.7] |
| 3σ + 8 on one sideWE1 WE4 | 151.5 [147.0, 156.0] | 107 [103, 111] | 62.5% [61.0%, 64.0%] | 14.6 [14.3, 14.9] | 0.32 [0.31, 0.33] | 21.9 [21.7, 22.2] |
| 3σ + trend of 6WE1 N3 | 198.0 [192.4, 204.0] | 141 [135, 147] | 52.5% [51.0%, 54.0%] | 41.2 [39.9, 42.4] | 0.91 [0.90, 0.93] | 29.5 [29.2, 29.8] |
| Western Electric (rules 1 to 4)WE1 WE2 WE3 WE4 | 91.3 [88.7, 94.0] | 64 [62, 67] | 80.8% [79.6%, 82.0%] | 9.5 [9.3, 9.7] | 0.21 [0.20, 0.22] | 18.2 [17.9, 18.3] |
| Western Electric + Nelson (all 8)WE1 WE2 WE3 WE4 N3 N4 N7 N8 | 66.0 [64.1, 67.7] | 48 [46, 49] | 90.3% [89.3%, 91.2%] | 9.4 [9.2, 9.6] | 0.21 [0.20, 0.22] | 17.9 [17.6, 18.1] |
Read the median as well as the mean. Run lengths are strongly skewed: with the 3σ rule the in-control average is about 362 points but the median is 254.5 (95% CI 244 to 264), and one chart in ten signals within 40 points (95% CI 36 to 44). An average run length on its own oversells how quiet a chart is. Median and percentile intervals are distribution-free, from order statistics.
Which rule adds what. “8 on one side” does most of the work against sustained shifts. The trend rule (6 rising or falling) barely helps against drifts: at 0.05σ per point its ratio to the 3σ rule is 0.97 (95% CI 0.97 to 0.98), while it shortens the in-control run length to 198. The four Nelson rules beyond Western Electric detect nothing faster here (1σ shift: 9.4 against 9.5) but cut the in-control run length from 91 to 66.
Many charts at once. The cation exchange assay has 11 measurements with spread. Charted together, even the 3σ rule will usually raise somewhere. The decision this informs is in DR-002.
Every scenario, every rule set
| Scenario | 3σ only | 3σ + WE4 | 3σ + N3 | WE 1 to 4 | All 8 |
|---|---|---|---|---|---|
| In control | 361.9[351.5, 372.5] | 151.5[147.0, 156.0] | 198.0[192.4, 204.0] | 91.3[88.7, 94.0] | 66.0[64.1, 67.7] |
| Shift of 0.25σ | 279.9[271.0, 289.1] | 94.9[92.1, 97.7] | 169.0[163.7, 174.0] | 57.8[56.1, 59.6] | 46.5[45.2, 47.8] |
| Shift of 0.5σ | 151.5[146.9, 156.1] | 45.0[43.7, 46.3] | 111.1[107.9, 114.6] | 27.3[26.5, 28.0] | 25.3[24.6, 26.0] |
| Shift of 0.75σ | 78.8[76.3, 81.1] | 22.9[22.4, 23.5] | 67.0[65.0, 69.0] | 14.8[14.4, 15.1] | 14.5[14.2, 14.9] |
| Shift of 1σ | 45.1[43.8, 46.6] | 14.6[14.3, 14.9] | 41.2[39.9, 42.4] | 9.5[9.3, 9.7] | 9.4[9.2, 9.6] |
| Shift of 1.5σ | 15.6[15.1, 16.0] | 8.0[7.8, 8.1] | 15.2[14.8, 15.6] | 5.3[5.2, 5.3] | 5.3[5.2, 5.3] |
| Shift of 2σ | 6.1[5.9, 6.3] | 4.8[4.7, 4.9] | 6.1[5.9, 6.3] | 3.4[3.3, 3.5] | 3.4[3.3, 3.4] |
| Shift of 2.5σ | 3.2[3.1, 3.3] | 3.0[3.0, 3.1] | 3.2[3.1, 3.3] | 2.4[2.4, 2.5] | 2.4[2.4, 2.5] |
| Shift of 3σ | 2.0[2.0, 2.1] | 2.0[2.0, 2.0] | 2.0[2.0, 2.1] | 1.8[1.8, 1.8] | 1.8[1.8, 1.8] |
| Drift of 0.05σ per point | 30.4[30.1, 30.7] | 21.9[21.7, 22.2] | 29.5[29.2, 29.8] | 18.2[17.9, 18.3] | 17.9[17.6, 18.1] |
| Drift of 0.1σ per point | 18.4[18.2, 18.6] | 14.7[14.6, 14.8] | 18.2[18.0, 18.3] | 12.2[12.1, 12.3] | 12.1[12.0, 12.2] |
Limits from a short history
With 20 Phase I points, a chart that false-alarms at least twice as often as the textbook chart is common: 41.1% (95% CI 39.8% to 42.5%) of simulated analysts get one. The chance of a false alarm in the next 100 results is 39.5% (95% CI 38.7% to 40.4%), against 23.7% with known limits. Even with 100 points, 17.9% of charts are twice as noisy; with 500, 1.4%.
- Sample SD (2021)
- Moving range / 1.128 (I-MR)
| m | Sigma from | Median ARL₀ | At least twice as noisy | False alarm in next 100 |
|---|---|---|---|---|
| 10 | Sample SD | 188 [176, 204] | 49.6% [48.2%, 51.0%] | 47.4% [46.4%, 48.4%] |
| 10 | Moving range | 217 [195, 236] | 47.8% [46.5%, 49.2%] | 45.9% [44.8%, 47.1%] |
| 20 | Sample SD | 259 [246, 273] | 41.1% [39.8%, 42.5%] | 39.5% [38.7%, 40.4%] |
| 20 | Moving range | 255 [238, 281] | 42.6% [41.2%, 44.0%] | 40.4% [39.5%, 41.4%] |
| 30 | Sample SD | 293 [280, 305] | 35.4% [34.1%, 36.7%] | 35.4% [34.7%, 36.1%] |
| 30 | Moving range | 306 [292, 323] | 37.6% [36.2%, 38.9%] | 36.6% [35.8%, 37.4%] |
| 50 | Sample SD | 319 [311, 329] | 27.9% [26.7%, 29.2%] | 31.4% [30.9%, 32.1%] |
| 50 | Moving range | 337 [324, 353] | 30.1% [28.9%, 31.4%] | 32.3% [31.6%, 33.0%] |
| 100 | Sample SD | 340 [334, 350] | 17.9% [16.9%, 19.0%] | 27.9% [27.5%, 28.4%] |
| 100 | Moving range | 355 [344, 365] | 23.3% [22.1%, 24.5%] | 28.9% [28.4%, 29.4%] |
| 200 | Sample SD | 356 [350, 362] | 8.6% [7.8%, 9.4%] | 25.9% [25.6%, 26.2%] |
| 200 | Moving range | 359 [349, 367] | 13.9% [13.0%, 14.9%] | 26.7% [26.3%, 27.1%] |
| 500 | Sample SD | 363 [360, 368] | 1.4% [1.1%, 1.7%] | 24.7% [24.5%, 24.9%] |
| 500 | Moving range | 364 [358, 368] | 3.4% [3.0%, 4.0%] | 24.9% [24.6%, 25.1%] |
“At least twice as noisy” means a conditional in-control average run length of 185.2 or less, half the known-limit value. The Phase I switch on /control-charts uses the first 40 points, between the 30 and 50 rows above. This is the effect studied by Quesenberry (1993) and reviewed by Jensen et al. (2006); the numbers here are this site's own simulation.
Retrospective limits can hide the shift
A sustained 3σ shift in the last third of a chart is caught by the 2021 limits in only 2.5% (95% CI 2.0% to 3.0%) of charts: the shift inflates the sample SD and drags the centre line with it. Limits from the first 40 points catch it in 100.0% (95% CI 99.9% to 100.0%). At 1.5σ the paired gain is 75.8 percentage points (95% CI 74.5 to 77.0).
- 2021: all points, sample SD
- All points, moving range (I-MR)
- First 40 points, sample SD
- First 40 points, moving range
| Limits (1.5σ shift) | Signal on the shifted points | Signal before the shift | vs 2021, paired |
|---|---|---|---|
| 2021: all points, sample SD | 14.8% [13.7%, 15.9%] | 6.7% [6.0%, 7.5%] | baseline |
| All points, moving range (I-MR) | 68.1% [66.6%, 69.5%] | 45.4% [43.8%, 46.9%] | +53.3 pp [51.7, 54.8] |
| First 40 points, sample SD | 90.6% [89.7%, 91.5%] | 27.2% [25.8%, 28.6%] | +75.8 pp [74.5, 77.0] |
| First 40 points, moving range | 88.1% [87.1%, 89.1%] | 30.3% [28.8%, 31.7%] | +73.3 pp [72.0, 74.6] |
Moving-range limits from all points are not a full fix: the centre line still moves towards the shifted points, so 45.4% (95% CI 43.8% to 46.9%) of charts also flag points before the shift, which sends a reviewer to the wrong runs. Phase I limits flag early points at their own rate, 27.2%, whatever happens later, which is the cost of estimating from 40 points (section B). The planted faults in the synthetic data are mostly single-run spikes, which the 2021 limits do catch.
Do the clusters survive a new seed, or new data?
High-silhouette pairings are the most repeatable and the least useful for review: UMAP or t-SNE with K-Means reproduce their partition across seeds but flag no injection at all. Where the procedure does flag injections, which ones it flags moves with the seed and the sample (prompt overlap well below 1), and the chosen k or eps can jump between runs.
| Assay | Pipeline | Reference run | ARI to reference | Prompt overlap (Jaccard) | Silhouette | k or eps chosen at the other seeds |
|---|---|---|---|---|---|---|
| AAE | PCA + K-Means | k = 154 prompts | 0.59 [0.51, 0.67] | 0.52 [0.39, 0.63] | 0.56 [0.56, 0.57] | 4 ×5, 10 ×5, 11 ×8, 12, 13 |
| AAE | PCA + DBSCAN | eps = 2.115 prompts | No random step: every seed gives the same answer. | |||
| AAE | t-SNE + DBSCAN2021 pick | eps = 2.436 prompts | 0.97 [0.94, 1.00] | 0.92 [0.85, 0.97] | 0.72 [0.71, 0.73] | 1.56 to 2.58 |
| CEX | PCA + K-Means | k = 46 prompts | 0.98 [0.94, 1.00] | 1.00 [1.00, 1.00] | 0.64 [0.64, 0.65] | 3 ×2, 4 ×18 |
| CEX | PCA + DBSCAN | eps = 0.3811 prompts | No random step: every seed gives the same answer. | |||
| CEX | t-SNE + DBSCAN2021 pick | eps = 1.730 prompts | 0.99 [0.99, 0.99] | 0.50 [0.30, 0.70] | 0.85 [0.84, 0.86] | 1.22 to 1.89 |
| CEX | UMAP + K-Means2021 pick | k = 70 prompts | 1.00 [1.00, 1.00] | flags nothing in any run | 0.94 [0.93, 0.94] | 7 ×20 |
| Protein A | PCA + K-Means | k = 57 prompts | 0.96 [0.94, 0.97] | 0.93 [0.90, 0.95] | 0.56 [0.56, 0.57] | 5 ×20 |
| Protein A | PCA + DBSCAN | eps = 0.2821 prompts | No random step: every seed gives the same answer. | |||
| Protein A | UMAP + K-Means2021 pick | k = 30 prompts | 0.96 [0.91, 1.00] | flags nothing in any run | 0.85 [0.84, 0.86] | 2, 3 ×18, 4 |
| SEC | PCA + K-Means | k = 53 prompts | 0.91 [0.90, 0.92] | 0.95 [0.88, 1.00] | 0.67 [0.67, 0.67] | 5 ×3, 6 ×17 |
| SEC | PCA + DBSCAN | eps = 0.736 prompts | No random step: every seed gives the same answer. | |||
| SEC | t-SNE + K-Means2021 pick | k = 50 prompts | 0.99 [0.96, 1.00] | flags nothing in any run | 0.83 [0.83, 0.83] | 5 ×19, 6 |
| SEC | UMAP + K-Means2021 pick | k = 80 prompts | 0.97 [0.94, 0.99] | flags nothing in any run | 0.92 [0.91, 0.92] | 6, 7 ×3, 8 ×16 |
| Assay | Pipeline | ARI to reference | Prompt overlap (Jaccard) | Silhouette | k or eps chosen (range) |
|---|---|---|---|---|---|
| AAE | PCA + K-Means | 0.60 [0.56, 0.64] | 0.43 [0.37, 0.48] | 0.58 [0.57, 0.59] | 4 to 16 |
| AAE | PCA + DBSCAN | 0.95 [0.92, 0.97] | 0.92 [0.88, 0.95] | 0.69 [0.68, 0.69] | 1.58 to 2.55 |
| CEX | PCA + K-Means | 0.90 [0.88, 0.92] | 0.94 [0.90, 0.97] | 0.64 [0.63, 0.65] | 3 to 9 |
| CEX | PCA + DBSCAN | 0.92 [0.89, 0.95] | 0.68 [0.65, 0.71] | 0.68 [0.67, 0.68] | 0.42 to 0.75 |
| Protein A | PCA + K-Means | 0.78 [0.75, 0.82] | 0.53 [0.46, 0.60] | 0.60 [0.59, 0.61] | 3 to 9 |
| Protein A | PCA + DBSCAN | 0.42 [0.31, 0.55] | 0.42 [0.36, 0.48] | 0.69 [0.68, 0.71] | 0.36 to 1.59 |
| SEC | PCA + K-Means | 0.86 [0.82, 0.89] | 0.37 [0.32, 0.42] | 0.67 [0.66, 0.67] | 3 to 7 |
| SEC | PCA + DBSCAN | 0.81 [0.73, 0.88] | 0.83 [0.76, 0.90] | 0.60 [0.60, 0.61] | 0.44 to 1.07 |
Do DBSCAN outliers and run-rule signals agree?
They hardly agree, and they should not be expected to: Cohen's kappa between the 3σ rule and DBSCAN is −0.01 (95% CI −0.03 to 0.01). The 3σ rule finds 69% of the runs with planted control faults (95% CI 67% to 71%); DBSCAN finds 82% of the runs with planted degraded samples (77% to 87%). They look at different injections.
| DBSCAN flags | DBSCAN does not | |
|---|---|---|
| Rule flags | 32 | 191 |
| Rule does not | 768 | 3809 |
Of the sequences the rule flags, DBSCAN also flags 14% (95% CI 10% to 18%).
| Assay | Kappa, 3σ vs DBSCAN | Fault runs found | Degraded-sample runs found | Clean runs flagged |
|---|---|---|---|---|
| AAE | −0.03 [−0.07, 0.02] | 3σ27% [23%, 31%]WE 1–438% [31%, 45%]DBSCAN8% [3%, 13%] | 3σ3% [0%, 6%]WE 1–416% [9%, 23%]DBSCAN76% [60%, 90%] | 3σ2% [1%, 3%]WE 1–413% [11%, 15%]DBSCAN4% [2%, 6%] |
| CEX | 0.00 [−0.04, 0.04] | 3σ100% [100%, 100%]WE 1–4100% [100%, 100%]DBSCAN23% [10%, 35%] | 3σ5% [2%, 9%]WE 1–422% [17%, 28%]DBSCAN73% [62%, 82%] | 3σ3% [2%, 4%]WE 1–423% [21%, 25%]DBSCAN8% [6%, 10%] |
| Protein A | −0.01 [−0.04, 0.02] | 3σ100% [100%, 100%]WE 1–4100% [100%, 100%]DBSCAN20% [10%, 30%] | 3σ4% [2%, 7%]WE 1–411% [7%, 16%]DBSCAN99% [98%, 100%] | 3σ1% [1%, 2%]WE 1–413% [11%, 15%]DBSCAN12% [8%, 16%] |
| SEC | −0.01 [−0.05, 0.04] | 3σ100% [100%, 100%]WE 1–4100% [100%, 100%]DBSCAN13% [3%, 23%] | 3σ4% [1%, 7%]WE 1–423% [15%, 32%]DBSCAN81% [74%, 88%] | 3σ1% [0%, 2%]WE 1–422% [19%, 24%]DBSCAN8% [6%, 11%] |
| All assays | −0.01 [−0.03, 0.01] | 3σ69% [67%, 71%]WE 1–474% [70%, 77%]DBSCAN14% [9%, 18%] | 3σ4% [2%, 6%]WE 1–418% [15%, 21%]DBSCAN82% [77%, 87%] | 3σ2% [1%, 2%]WE 1–418% [16%, 19%]DBSCAN8% [6%, 9%] |
The amino acid assay is the exception for the rules: most of its planted fault runs carry a slow drift inside the limits, which the 3σ rule misses. Over 20 worlds it finds 24 of the 89 fault runs, 27% (95% CI 23% to 31%). Western Electric rules 1 to 4, read on the same runs, find 34, 38% (95% CI 31% to 45%), but flag 13% (95% CI 11% to 15%) of the assay's clean runs, against 2% for the 3σ rule. Over all four assays they flag 18% and 2% of clean runs. “Planted” recall is only possible because the data are synthetic; it says nothing about real faults.
How it was computed
- Run lengths: 4,000 streams per scenario, cut at 20,000 points (no stream reached it); intervals from 2,000 bootstrap resamples; ratios by paired bootstrap on the same streams. The engine is checked against exact values: 1 / P(|Z + δ| > 3) for the 3σ rule and 2⁸ − 1 = 255 for “8 on one side” alone.
- Estimated limits: 5,000 Phase I samples for each size, both sigma estimates computed on the same samples.
- Retrospective limits: 4,000 charts of 150 points, the same noise for every shift size and limit method.
- Clustering: the default world (seed 30034); 21 algorithm seeds (the reference and 20 more); 50 bootstrap resamples per PCA pipeline. Agreement: 20 worlds (seeds 30034 and 1 to 19).
- All of it is produced by
pnpm studyinto one JSON artefact (274 s on a laptop). The test suite re-runs a slice of every study with the same seeds and requires identical results, and checks the statistics helpers against SciPy, statsmodels, scikit-learn and R.
What these numbers cannot tell you
- The run-length studies assume independent, normally distributed results. Real HPLC controls can be autocorrelated within a run, skewed, or rounded, and all three change false-alarm rates.
- The clustering and agreement studies use the synthetic generator, whose faults were planted to be findable. They show how the methods behave, not how often a real lab would catch a real fault.
- Western Electric and Nelson rules with estimated limits were not simulated; section B covers the 3σ rule only.
- More in the methods notes, decision records and model card.