Plan: Calibration — calibrated control limits for processbehavior 0.2.0
Context
processbehavior computes natural process limits entirely from the current study's
data. A recurring workflow it doesn't support: "I have a known baseline mean and
within-subgroup standard deviation from a prior stable period — chart new data against
those limits instead of recomputing them." This issue adds calibration: a
user-supplied (μ, σ) attached to a Study that all charts substitute into the
existing limit math.
Longer-term app UX: let a user select a window of points, compute (μ, σ), store it on
the Study, and re-execute. That requires calibration to be a Study-level property
(persists across pickle; applies to every subsequent execution), not a per-execute
parameter.
Ships in 0.2.0, after the 0.1.0 PyPI release. Adding calibration later is
non-breaking (optional frozen-dataclass field).
Goal
Calibration(mean: float, sd: float, source: str | None = None) immutable public
value object. source is free-text provenance (e.g. "Q3 2024 baseline").
Study.calibration: Calibration | None + Study.with_calibration(cal) returning a
new Study via dataclasses.replace. Cheap — reuses the existing
_ads/_spec/_plan; no re-formulate().
- When calibration is set, response control charts (X, mR, Xbar, S) substitute
(μ, σ) into their existing limit formulas.
- Residual charts substitute σ as a calibrated noise floor (see "Residual &
recentered semantics"); μ is unused on residuals. Gated on Tom Q1–Q5.
- Chart metadata declares the limits source for downstream consumers.
- Calibration persists transparently via pickle.
- Zero impact on uncalibrated paths — byte-for-byte identical results when
calibration is None.
Architecture decision (resolves the earlier confusion)
The chart calculators cannot "read study.calibration directly" — execute()
dispatches Analysis(self._spec, request, analysis_dataset=self._ads) and Analysis
holds no Study reference. Wiring one in would be circular.
Calibration threads through ChartRequest — the existing ephemeral per-execute
carrier (formulation_spec.py:96) that already conveys n_sigma, n_mode, phased,
read by calculators as self.request.*. execute() snapshots self.calibration into
the request; calculators read self.request.calibration. This leaves Analysis
decoupled from Study and reuses the established pattern.
study.calibration (persisted, frozen Study field)
│ execute() copies it in ── execute() signature UNCHANGED
▼
ChartRequest.calibration (ephemeral)
▼ self.request.calibration
analysis.py limit-math sites → substitute μ, σ
Locked design decisions
- Study-level only. No per-execute
calibration= parameter; execute() signature
preserved byte-for-byte.
Calibration(mean, sd, source=None) — source is optional provenance.
ValidationError for all validation (never raw ValueError).
- σ on residuals: calibration.σ replaces the data-estimated R2 noise floor as the
dispersion basis for residual-chart limits; center stays ν≈0; μ unused on residuals.
- Xbar c₄ subtlety:
calculate_limits(limits_type='Xbar') computes
Wd = sd / c4(N), so the calculator passes sd = σ·c4(N) to yield μ ± 3σ/√n
exactly. Pinned by a test on an unbalanced design.
Residual & recentered semantics (findings + pending decisions)
What the code does today (confirmed):
- Residual chart center: non-recentered → ν≈0; recentered (RCR) → a VAS baseline
reconstruction (analysis_dataset.py:354-361), e.g. RCR1=Ȳ+R1,
RCR3=(Ȳ_k+Ȳ_t−Ȳ)+R3. Recentering moves the center only; limit width is unchanged.
- Residual limit width: for effect residuals R1/R3/R4/R5,
_resolve_limits_column
(analysis.py:572-588) swaps in R2 (within-cell noise floor) as the dispersion
basis — not the residual's own spread. (Caveat: applied on the Xbar/S path; the
X/mR-on-residual path currently uses the residual column's own moving range — an
existing inconsistency.)
Reframe: residual limits are already built from the noise floor, and calibration.σ
is a noise floor → σ substitutes cleanly. calibration.μ is response-location only and
has no role on a mean-removed residual (except possibly the recentered grand-mean
anchor — speculative).
Confirmed: σ-on-residuals enabled; μ ignored for residuals; center stays ν≈0.
Pending Tom — placeholders that drive the design:
| # |
Question for Tom |
Tom's answer |
Design consequence |
| 1 |
May VAS substitute a calibrated within-subgroup σ for the data-estimated R2 σ on residual charts? |
TBD |
YES → σ substitutes R2 dispersion on residual limits. NO → raise on residual+calibration. |
| 2 |
Is calibrated σ the single yardstick for all effect residuals (R1,R3,R4,R5,R6), as R2 is today? |
TBD |
YES → uniform substitution across residuals. PARTIAL → restrict to the named residuals. |
| 3 |
With an already-unbiased calibrated σ, still apply c₄ bias correction + per-cell √n scaling on Xbar/S residuals, or only √n? |
TBD |
Determines whether calculator passes σ·c4(N) or σ into calculate_limits. |
| 4 |
Should X/mR residual limits also be driven by calibrated σ (mR ≡ d₂·σ), unifying with Xbar/S? |
TBD |
YES → normalize the X/mR-vs-Xbar/S inconsistency under calibration (own decision/changelog note). NO → leave X/mR residual on own mR. |
| 5 |
On a non-recentered residual (center ν=0), confirm μ has no role — only σ? |
TBD |
Confirms "μ unused on residuals." If NO → revisit. |
| 6 |
On a recentered residual, may calibrated μ replace the grand-mean (Ȳ) anchor while effect means stay data-driven, or does that violate the decomposition? |
TBD |
YES → inject μ into RCR baseline's Ȳ term. NO → recentered baseline stays fully data-driven. |
| 7 |
Is "recenter a residual onto a calibrated μ" a workflow you endorse, or is recentering inherently data-reconstruction? |
TBD |
Endorsed → support recentered+calibration. Not → raise on recentered+calibration (MVP boundary). |
| 8 |
Preferred VAS term / reference section for calibrating limits to a prior baseline; any caution about the word "calibration"? |
TBD |
Confirms naming + docs citation in docs/guide/calibration.md. |
Non-Goals (MVP)
- No window-selection logic in the library (app/user computes
(μ, σ)).
- No per-stratum calibration; one
(μ, σ) applies to every lane.
- No per-execute
calibration= override.
- No
phased=True + calibration (per-phase data limits conflict with a global
override) → raises ValidationError.
- No capability/loss-function interaction.
- No backfill into 0.1.0.
Files to Touch
New
processbehavior/calibration.py (~45 lines) — frozen Calibration(mean, sd, source=None);
__post_init__ raises ValidationError on non-numeric mean / non-positive sd.
tests/test_calibration.py — dataclass validation; with_calibration round-trip
- clear; calibrated X (
center=μ, UPL=μ+3σ); mR (center=1.128σ); Xbar c₄ test
on an unbalanced design; S; stratified; residual σ-substitution; the
phased/(gated) recentered/(gated) residual exclusions; pickle (incl. source);
uncalibrated regression gate; plot notation.
docs/guide/calibration.md (MyST) — use case, API, four-chart math, residual σ
semantics, exclusions; cite the VAS reference per Tom Q8.
Modified
processbehavior/__init__.py — re-export Calibration beside SpecLimits
(line 58 / __all__ line 82).
processbehavior/formulation_spec.py — add calibration: Calibration | None = None
to ChartRequest (after line 133).
processbehavior/study.py — calibration field on frozen Study;
with_calibration(); pass calibration=self.calibration into the ChartRequest
built in execute() (line ~1853); __repr__ line when set; validation guards
(phased+calibration always raises; residual/recentered guards per Tom answers).
processbehavior/analysis.py — substitution at limit-math sites reading
self.request.calibration: X/mR (mR = d2·σ, center = μ), Xbar (center = μ,
pass sd = σ·c4(N) per Tom Q3), S (center = c4(N)·σ); residual paths substitute σ
as the noise floor (gated on Tom Q1–Q4). Chart-info metadata via a helper:
'limits_source': 'calibration'|'data',
'calibration': {'mean','sd','source'} | None.
processbehavior/plotting/plotter.py — _generate_title (~1437) and
_generate_subplot_title (~1465) append " (σ given)" when
metadata['limits_source']=='calibration'; add a "Calibrated: μ=…, σ=…"
annotation driven by metadata.
CHANGELOG.md — [Unreleased] > Added entry.
Behavioral Changes
| Surface |
Before |
After |
Study.calibration |
— |
new field, default None |
study.execute(...), calibration is None |
data limits |
identical (regression-gated) |
study.execute(...), calibration set, response chart |
n/a |
limits from (μ, σ); per-cell √n scaling preserved |
study.execute(...), calibration set, residual chart |
n/a |
limit width from calibrated σ (noise floor); center ν≈0; μ unused — gated on Tom |
chart metadata limits_source |
— |
always 'data' or 'calibration' |
| pickle round-trip |
no field |
preserves calibration (incl. source) |
phased=True + calibration |
n/a |
raises ValidationError |
| recentered / residual + calibration |
n/a |
per Tom answers (substitute σ, or raise) |
Invariants to Preserve
Study stays frozen; with_calibration returns a new instance.
- Uncalibrated path byte-for-byte identical.
Analysis keeps no Study reference — calibration arrives only via ChartRequest.
- Per-cell variable limits (n_kt √n scaling) still apply — calibration substitutes σ.
- Bishop VAS unweighted cell-mean center preserved on the uncalibrated path.
get_statistics() dict shape unchanged: {N, center, lpl, upl}.
Test Strategy
pytest tests/test_calibration.py -v.
pytest tests/ -m "not slow" -x — zero regression.
- Validate calibrated limits against hand-computed values on
validation/PBTESTDATABASE_T100.csv.
Risk Areas / Edge Cases
- Xbar σ vs c₄ — naive
sd=σ inflates limits by 1/c4(N); must pass σ·c4(N)
(pending Tom Q3). Test on an unbalanced design so a missing factor fails loudly.
- σ semantics — within-subgroup vs pooled; mitigate via docstring + guide.
- Empty / single-stratum — calibrated chart still emits metadata with no points.
- 0.1.0 pickle forward-compat — old Studies lack the field;
= None default absorbs.
- X/mR vs Xbar/S residual basis inconsistency — calibrated σ could normalize it
(Tom Q4); if so, document as an intentional behavior change.
- Title pollution —
" (σ given)" on every calibrated title; app can override via
result.plot(title=...).
Sequencing
- Branch
feat/calibration.
calibration.py + __init__ export + dataclass test.
Study.calibration + with_calibration() + Study tests (pickle, no-op when unset).
ChartRequest.calibration + execute() wiring + phased guard.
- Response-chart substitution (X/mR, then Xbar/S) + tests + regression baseline.
- Resolve Tom Q1–Q7, then residual/recentered substitution or guards + tests.
- Metadata fields + plot notation + tests.
docs/guide/calibration.md (cite per Tom Q8), CHANGELOG, version bump.
- App follow-up PR (storage + window-selection UI) — separate scope.
Plan: Calibration — calibrated control limits for processbehavior 0.2.0
Context
processbehavior computes natural process limits entirely from the current study's
data. A recurring workflow it doesn't support: "I have a known baseline mean and
within-subgroup standard deviation from a prior stable period — chart new data against
those limits instead of recomputing them." This issue adds calibration: a
user-supplied
(μ, σ)attached to aStudythat all charts substitute into theexisting limit math.
Longer-term app UX: let a user select a window of points, compute
(μ, σ), store it onthe Study, and re-execute. That requires calibration to be a Study-level property
(persists across pickle; applies to every subsequent execution), not a per-execute
parameter.
Ships in 0.2.0, after the 0.1.0 PyPI release. Adding calibration later is
non-breaking (optional frozen-dataclass field).
Goal
Calibration(mean: float, sd: float, source: str | None = None)immutable publicvalue object.
sourceis free-text provenance (e.g."Q3 2024 baseline").Study.calibration: Calibration | None+Study.with_calibration(cal)returning anew Study via
dataclasses.replace. Cheap — reuses the existing_ads/_spec/_plan; no re-formulate().(μ, σ)into their existing limit formulas.recentered semantics"); μ is unused on residuals. Gated on Tom Q1–Q5.
calibration is None.Architecture decision (resolves the earlier confusion)
The chart calculators cannot "read
study.calibrationdirectly" —execute()dispatches
Analysis(self._spec, request, analysis_dataset=self._ads)andAnalysisholds no
Studyreference. Wiring one in would be circular.Calibration threads through
ChartRequest— the existing ephemeral per-executecarrier (
formulation_spec.py:96) that already conveysn_sigma,n_mode,phased,read by calculators as
self.request.*.execute()snapshotsself.calibrationintothe request; calculators read
self.request.calibration. This leavesAnalysisdecoupled from
Studyand reuses the established pattern.Locked design decisions
calibration=parameter;execute()signaturepreserved byte-for-byte.
Calibration(mean, sd, source=None)—sourceis optional provenance.ValidationErrorfor all validation (never rawValueError).dispersion basis for residual-chart limits; center stays ν≈0; μ unused on residuals.
calculate_limits(limits_type='Xbar')computesWd = sd / c4(N), so the calculator passessd = σ·c4(N)to yieldμ ± 3σ/√nexactly. Pinned by a test on an unbalanced design.
Residual & recentered semantics (findings + pending decisions)
What the code does today (confirmed):
reconstruction (
analysis_dataset.py:354-361), e.g.RCR1=Ȳ+R1,RCR3=(Ȳ_k+Ȳ_t−Ȳ)+R3. Recentering moves the center only; limit width is unchanged._resolve_limits_column(
analysis.py:572-588) swaps in R2 (within-cell noise floor) as the dispersionbasis — not the residual's own spread. (Caveat: applied on the Xbar/S path; the
X/mR-on-residual path currently uses the residual column's own moving range — an
existing inconsistency.)
Reframe: residual limits are already built from the noise floor, and calibration.σ
is a noise floor → σ substitutes cleanly. calibration.μ is response-location only and
has no role on a mean-removed residual (except possibly the recentered grand-mean
anchor — speculative).
Confirmed: σ-on-residuals enabled; μ ignored for residuals; center stays ν≈0.
Pending Tom — placeholders that drive the design:
σ·c4(N)orσintocalculate_limits.docs/guide/calibration.md.Non-Goals (MVP)
(μ, σ)).(μ, σ)applies to every lane.calibration=override.phased=True+ calibration (per-phase data limits conflict with a globaloverride) → raises
ValidationError.Files to Touch
New
processbehavior/calibration.py(~45 lines) — frozenCalibration(mean, sd, source=None);__post_init__raisesValidationErroron non-numeric mean / non-positive sd.tests/test_calibration.py— dataclass validation;with_calibrationround-tripcenter=μ,UPL=μ+3σ); mR (center=1.128σ); Xbar c₄ teston an unbalanced design; S; stratified; residual σ-substitution; the
phased/(gated)recentered/(gated) residual exclusions; pickle (incl.source);uncalibrated regression gate; plot notation.
docs/guide/calibration.md(MyST) — use case, API, four-chart math, residual σsemantics, exclusions; cite the VAS reference per Tom Q8.
Modified
processbehavior/__init__.py— re-exportCalibrationbesideSpecLimits(line 58 /
__all__line 82).processbehavior/formulation_spec.py— addcalibration: Calibration | None = Noneto
ChartRequest(after line 133).processbehavior/study.py—calibrationfield on frozenStudy;with_calibration(); passcalibration=self.calibrationinto theChartRequestbuilt in
execute()(line ~1853);__repr__line when set; validation guards(
phased+calibration always raises; residual/recentered guards per Tom answers).processbehavior/analysis.py— substitution at limit-math sites readingself.request.calibration: X/mR (mR = d2·σ,center = μ), Xbar (center = μ,pass
sd = σ·c4(N)per Tom Q3), S (center = c4(N)·σ); residual paths substitute σas the noise floor (gated on Tom Q1–Q4). Chart-info metadata via a helper:
'limits_source': 'calibration'|'data','calibration': {'mean','sd','source'} | None.processbehavior/plotting/plotter.py—_generate_title(~1437) and_generate_subplot_title(~1465) append" (σ given)"whenmetadata['limits_source']=='calibration'; add a"Calibrated: μ=…, σ=…"annotation driven by metadata.
CHANGELOG.md—[Unreleased] > Addedentry.Behavioral Changes
Study.calibrationNonestudy.execute(...),calibration is Nonestudy.execute(...), calibration set, response chart(μ, σ); per-cell √n scaling preservedstudy.execute(...), calibration set, residual chartlimits_source'data'or'calibration'source)phased=True+ calibrationValidationErrorInvariants to Preserve
Studystays frozen;with_calibrationreturns a new instance.Analysiskeeps noStudyreference — calibration arrives only viaChartRequest.get_statistics()dict shape unchanged:{N, center, lpl, upl}.Test Strategy
pytest tests/test_calibration.py -v.pytest tests/ -m "not slow" -x— zero regression.validation/PBTESTDATABASE_T100.csv.Risk Areas / Edge Cases
sd=σinflates limits by1/c4(N); must passσ·c4(N)(pending Tom Q3). Test on an unbalanced design so a missing factor fails loudly.
= Nonedefault absorbs.(Tom Q4); if so, document as an intentional behavior change.
" (σ given)"on every calibrated title; app can override viaresult.plot(title=...).Sequencing
feat/calibration.calibration.py+__init__export + dataclass test.Study.calibration+with_calibration()+ Study tests (pickle, no-op when unset).ChartRequest.calibration+execute()wiring +phasedguard.docs/guide/calibration.md(cite per Tom Q8), CHANGELOG, version bump.