Skip to content

Enable Calibration: calibrated control limits from a known (μ, σ) (0.2.0) #79

Description

@cnicholas

Plan: Calibration — calibrated control limits for processbehavior 0.2.0

Context

processbehavior computes natural process limits entirely from the current study's
data. A recurring workflow it doesn't support: "I have a known baseline mean and
within-subgroup standard deviation from a prior stable period — chart new data against
those limits instead of recomputing them." This issue adds calibration: a
user-supplied (μ, σ) attached to a Study that all charts substitute into the
existing limit math.

Longer-term app UX: let a user select a window of points, compute (μ, σ), store it on
the Study, and re-execute. That requires calibration to be a Study-level property
(persists across pickle; applies to every subsequent execution), not a per-execute
parameter.

Ships in 0.2.0, after the 0.1.0 PyPI release. Adding calibration later is
non-breaking (optional frozen-dataclass field).

Goal

  1. Calibration(mean: float, sd: float, source: str | None = None) immutable public
    value object. source is free-text provenance (e.g. "Q3 2024 baseline").
  2. Study.calibration: Calibration | None + Study.with_calibration(cal) returning a
    new Study via dataclasses.replace. Cheap — reuses the existing
    _ads/_spec/_plan; no re-formulate().
  3. When calibration is set, response control charts (X, mR, Xbar, S) substitute
    (μ, σ) into their existing limit formulas.
  4. Residual charts substitute σ as a calibrated noise floor (see "Residual &
    recentered semantics"); μ is unused on residuals. Gated on Tom Q1–Q5.
  5. Chart metadata declares the limits source for downstream consumers.
  6. Calibration persists transparently via pickle.
  7. Zero impact on uncalibrated paths — byte-for-byte identical results when
    calibration is None.

Architecture decision (resolves the earlier confusion)

The chart calculators cannot "read study.calibration directly" — execute()
dispatches Analysis(self._spec, request, analysis_dataset=self._ads) and Analysis
holds no Study reference. Wiring one in would be circular.

Calibration threads through ChartRequest — the existing ephemeral per-execute
carrier (formulation_spec.py:96) that already conveys n_sigma, n_mode, phased,
read by calculators as self.request.*. execute() snapshots self.calibration into
the request; calculators read self.request.calibration. This leaves Analysis
decoupled from Study and reuses the established pattern.

study.calibration  (persisted, frozen Study field)
        │  execute() copies it in  ──  execute() signature UNCHANGED
        ▼
ChartRequest.calibration  (ephemeral)
        ▼  self.request.calibration
analysis.py limit-math sites  →  substitute μ, σ

Locked design decisions

  • Study-level only. No per-execute calibration= parameter; execute() signature
    preserved byte-for-byte.
  • Calibration(mean, sd, source=None) — source is optional provenance.
  • ValidationError for all validation (never raw ValueError).
  • σ on residuals: calibration.σ replaces the data-estimated R2 noise floor as the
    dispersion basis for residual-chart limits; center stays ν≈0; μ unused on residuals.
  • Xbar c₄ subtlety: calculate_limits(limits_type='Xbar') computes
    Wd = sd / c4(N), so the calculator passes sd = σ·c4(N) to yield μ ± 3σ/√n
    exactly. Pinned by a test on an unbalanced design.

Residual & recentered semantics (findings + pending decisions)

What the code does today (confirmed):

  • Residual chart center: non-recentered → ν≈0; recentered (RCR) → a VAS baseline
    reconstruction (analysis_dataset.py:354-361), e.g. RCR1=Ȳ+R1,
    RCR3=(Ȳ_k+Ȳ_t−Ȳ)+R3. Recentering moves the center only; limit width is unchanged.
  • Residual limit width: for effect residuals R1/R3/R4/R5, _resolve_limits_column
    (analysis.py:572-588) swaps in R2 (within-cell noise floor) as the dispersion
    basis — not the residual's own spread. (Caveat: applied on the Xbar/S path; the
    X/mR-on-residual path currently uses the residual column's own moving range — an
    existing inconsistency.)

Reframe: residual limits are already built from the noise floor, and calibration.σ
is a noise floor → σ substitutes cleanly. calibration.μ is response-location only and
has no role on a mean-removed residual (except possibly the recentered grand-mean
anchor — speculative).

Confirmed: σ-on-residuals enabled; μ ignored for residuals; center stays ν≈0.

Pending Tom — placeholders that drive the design:

# Question for Tom Tom's answer Design consequence
1 May VAS substitute a calibrated within-subgroup σ for the data-estimated R2 σ on residual charts? TBD YES → σ substitutes R2 dispersion on residual limits. NO → raise on residual+calibration.
2 Is calibrated σ the single yardstick for all effect residuals (R1,R3,R4,R5,R6), as R2 is today? TBD YES → uniform substitution across residuals. PARTIAL → restrict to the named residuals.
3 With an already-unbiased calibrated σ, still apply c₄ bias correction + per-cell √n scaling on Xbar/S residuals, or only √n? TBD Determines whether calculator passes σ·c4(N) or σ into calculate_limits.
4 Should X/mR residual limits also be driven by calibrated σ (mR ≡ d₂·σ), unifying with Xbar/S? TBD YES → normalize the X/mR-vs-Xbar/S inconsistency under calibration (own decision/changelog note). NO → leave X/mR residual on own mR.
5 On a non-recentered residual (center ν=0), confirm μ has no role — only σ? TBD Confirms "μ unused on residuals." If NO → revisit.
6 On a recentered residual, may calibrated μ replace the grand-mean (Ȳ) anchor while effect means stay data-driven, or does that violate the decomposition? TBD YES → inject μ into RCR baseline's Ȳ term. NO → recentered baseline stays fully data-driven.
7 Is "recenter a residual onto a calibrated μ" a workflow you endorse, or is recentering inherently data-reconstruction? TBD Endorsed → support recentered+calibration. Not → raise on recentered+calibration (MVP boundary).
8 Preferred VAS term / reference section for calibrating limits to a prior baseline; any caution about the word "calibration"? TBD Confirms naming + docs citation in docs/guide/calibration.md.

Non-Goals (MVP)

  • No window-selection logic in the library (app/user computes (μ, σ)).
  • No per-stratum calibration; one (μ, σ) applies to every lane.
  • No per-execute calibration= override.
  • No phased=True + calibration (per-phase data limits conflict with a global
    override) → raises ValidationError.
  • No capability/loss-function interaction.
  • No backfill into 0.1.0.

Files to Touch

New

  1. processbehavior/calibration.py (~45 lines) — frozen Calibration(mean, sd, source=None);
    __post_init__ raises ValidationError on non-numeric mean / non-positive sd.
  2. tests/test_calibration.py — dataclass validation; with_calibration round-trip
    • clear; calibrated X (center=μ, UPL=μ+3σ); mR (center=1.128σ); Xbar c₄ test
      on an unbalanced design
      ; S; stratified; residual σ-substitution; the
      phased/(gated) recentered/(gated) residual exclusions; pickle (incl. source);
      uncalibrated regression gate; plot notation.
  3. docs/guide/calibration.md (MyST) — use case, API, four-chart math, residual σ
    semantics, exclusions; cite the VAS reference per Tom Q8.

Modified

  1. processbehavior/__init__.py — re-export Calibration beside SpecLimits
    (line 58 / __all__ line 82).
  2. processbehavior/formulation_spec.py — add calibration: Calibration | None = None
    to ChartRequest (after line 133).
  3. processbehavior/study.py — calibration field on frozen Study;
    with_calibration(); pass calibration=self.calibration into the ChartRequest
    built in execute() (line ~1853); __repr__ line when set; validation guards
    (phased+calibration always raises; residual/recentered guards per Tom answers).
  4. processbehavior/analysis.py — substitution at limit-math sites reading
    self.request.calibration: X/mR (mR = d2·σ, center = μ), Xbar (center = μ,
    pass sd = σ·c4(N) per Tom Q3), S (center = c4(N)·σ); residual paths substitute σ
    as the noise floor (gated on Tom Q1–Q4). Chart-info metadata via a helper:
    'limits_source': 'calibration'|'data',
    'calibration': {'mean','sd','source'} | None.
  5. processbehavior/plotting/plotter.py — _generate_title (~1437) and
    _generate_subplot_title (~1465) append " (σ given)" when
    metadata['limits_source']=='calibration'; add a "Calibrated: μ=…, σ=…"
    annotation driven by metadata.
  6. CHANGELOG.md — [Unreleased] > Added entry.

Behavioral Changes

Surface Before After
Study.calibration — new field, default None
study.execute(...), calibration is None data limits identical (regression-gated)
study.execute(...), calibration set, response chart n/a limits from (μ, σ); per-cell √n scaling preserved
study.execute(...), calibration set, residual chart n/a limit width from calibrated σ (noise floor); center ν≈0; μ unused — gated on Tom
chart metadata limits_source — always 'data' or 'calibration'
pickle round-trip no field preserves calibration (incl. source)
phased=True + calibration n/a raises ValidationError
recentered / residual + calibration n/a per Tom answers (substitute σ, or raise)

Invariants to Preserve

  • Study stays frozen; with_calibration returns a new instance.
  • Uncalibrated path byte-for-byte identical.
  • Analysis keeps no Study reference — calibration arrives only via ChartRequest.
  • Per-cell variable limits (n_kt √n scaling) still apply — calibration substitutes σ.
  • Bishop VAS unweighted cell-mean center preserved on the uncalibrated path.
  • get_statistics() dict shape unchanged: {N, center, lpl, upl}.

Test Strategy

  1. pytest tests/test_calibration.py -v.
  2. pytest tests/ -m "not slow" -x — zero regression.
  3. Validate calibrated limits against hand-computed values on
    validation/PBTESTDATABASE_T100.csv.

Risk Areas / Edge Cases

  1. Xbar σ vs c₄ — naive sd=σ inflates limits by 1/c4(N); must pass σ·c4(N)
    (pending Tom Q3). Test on an unbalanced design so a missing factor fails loudly.
  2. σ semantics — within-subgroup vs pooled; mitigate via docstring + guide.
  3. Empty / single-stratum — calibrated chart still emits metadata with no points.
  4. 0.1.0 pickle forward-compat — old Studies lack the field; = None default absorbs.
  5. X/mR vs Xbar/S residual basis inconsistency — calibrated σ could normalize it
    (Tom Q4); if so, document as an intentional behavior change.
  6. Title pollution — " (σ given)" on every calibrated title; app can override via
    result.plot(title=...).

Sequencing

  1. Branch feat/calibration.
  2. calibration.py + __init__ export + dataclass test.
  3. Study.calibration + with_calibration() + Study tests (pickle, no-op when unset).
  4. ChartRequest.calibration + execute() wiring + phased guard.
  5. Response-chart substitution (X/mR, then Xbar/S) + tests + regression baseline.
  6. Resolve Tom Q1–Q7, then residual/recentered substitution or guards + tests.
  7. Metadata fields + plot notation + tests.
  8. docs/guide/calibration.md (cite per Tom Q8), CHANGELOG, version bump.
  9. App follow-up PR (storage + window-selection UI) — separate scope.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions