Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
8 changes: 8 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -6,6 +6,14 @@ Versions follow [SemVer](https://semver.org).

## [Unreleased]

- Add optional per-benchmark `regression` allowances with a free tolerance,
exclusive hard cap, and optional measured climbed-gain requirement, in
relative or absolute units. Suite rows and reports identify the applied rule.
Upgrading: no action for existing contracts; their gate behavior and measurement
signatures are unchanged. Legacy suite rows without `rule` default to
`legacy-floor`; the new row field is additive. Contracts using the new block
require this kernel version and must remove it before rolling back.

### Upgrading

Operator actions (everything else needs no action; details in each entry):
Expand Down
45 changes: 45 additions & 0 deletions docs/contract.md
Original file line number Diff line number Diff line change
Expand Up @@ -21,6 +21,7 @@ The knobs that shape a climb, all optional:
| Knob | What it decides |
| --- | --- |
| `seed_env`, `min_delta` / `min_delta_rel` | Paired seeding for resampled evals, and the significance floor a delta must clear — calibrate it from seed variance, the gate enforces it |
| `regression.free_rel`, `max_rel`, `requires_gain_rel` (or absolute `free`, `max`, `requires_gain`) | Per-benchmark sibling regression allowance, independent of its win floor |
| `eval_minutes`, `gpus` | Evals that need their own job (and GPUs) are dispatched to the cluster rather than run in the author's job |
| `baseline: paired \| cached` | Re-measure the base tree beside every candidate, or measure it once per base and run only candidates |
| `depth_k`, `sleep_k` | How many experiments an author may launch and how many times it may sleep for results |
Expand Down Expand Up @@ -128,3 +129,47 @@ from either the fleet author or a rebound author. Supported judge backends are
`kind:backend:model[endpoint=judge-profile]` with their own judge credential.
Omitted models retain today's author-model inheritance. This existing setting
needs no additional contract pin; self-report verification bypasses the panel.

### Conditional sibling regressions

When shared code changes, a sibling can opt into a separate regression policy:

```yaml
- name: rollout-mem
command: ./bench --memory --json
metric: rollout_mem_bytes_f16
direction: min
min_delta_rel: 0.05
regression:
free_rel: 0.005
max_rel: 0.5
requires_gain_rel: 0.05
```

A win on this benchmark still requires at least 5% saved. As a sibling,
regressions up to and including 0.5% are free; larger regressions below 50%
require at least 5% measured gain on the climbed benchmark. A regression of
50% or more is refused regardless of gain. Reports and suite verdict rows
record the applied rule, such as `free-allowance`, `gain-unlocked`,
`insufficient-gain`, or `hard-cap`.

Use one unit family per block. Relative values are fractions of the absolute
baseline; `free_rel` defaults to this sibling's `min_delta_rel`, or zero.
The absolute twins `free`, `max`, and `requires_gain` use sibling metric units
for losses and climbed metric units for gains; `free` defaults to `min_delta`,
or zero. An empty block selects relative units when `min_delta_rel` is set,
otherwise absolute units. The default free allowance must also be at most max.

All values must be finite and nonnegative. `requires_gain_rel` needs `max_rel`
(and `requires_gain` needs `max`). Without a cap, the free allowance is the
entire tolerance. With a cap but no gain requirement, every regression below
the cap is allowed. The cap is exclusive and takes precedence when free equals
max. Non-regressing siblings pass, including when max is zero. Non-finite
measurements fail closed. A zero sibling baseline scales relative allowances
to zero; a zero climbed baseline cannot unlock a relative gain requirement.
Negative baselines use their absolute magnitude as the scale.

Omitting `regression` preserves the existing gate exactly: a regression is
refused only when it exceeds the larger of the sibling's absolute and scaled
relative significance floors. This block changes only sibling checks, not
whether this benchmark qualifies as the climbed benchmark's improvement.
55 changes: 55 additions & 0 deletions src/outerloop/contract.py
Original file line number Diff line number Diff line change
Expand Up @@ -84,6 +84,42 @@ class _StrictModel(BaseModel):
model_config = ConfigDict(extra="forbid")


class Regression(_StrictModel):
"""Sibling regression allowance, in one unit family per block.

free_rel: fraction of the sibling's absolute baseline always allowed;
defaults to its min_delta_rel (or zero). max_rel: optional exclusive
hard cap. requires_gain_rel: climbed benchmark's minimum relative gain
to unlock the interval above free_rel; requires max_rel. The absolute
twins free/max/requires_gain use sibling/climbed metric units respectively.
Without a cap, free is the entire allowance. All values must be finite
and nonnegative; fractions may exceed one for unbounded metrics.
"""

free_rel: float | None = Field(default=None, ge=0, allow_inf_nan=False)
max_rel: float | None = Field(default=None, ge=0, allow_inf_nan=False)
requires_gain_rel: float | None = Field(default=None, ge=0, allow_inf_nan=False)
free: float | None = Field(default=None, ge=0, allow_inf_nan=False)
max: float | None = Field(default=None, ge=0, allow_inf_nan=False)
requires_gain: float | None = Field(default=None, ge=0, allow_inf_nan=False)

@model_validator(mode="after")
def _bounds(self) -> Regression:
relative = any(v is not None for v in (self.free_rel, self.max_rel, self.requires_gain_rel))
absolute = any(v is not None for v in (self.free, self.max, self.requires_gain))
if relative and absolute:
raise ValueError("regression must use either relative or absolute units, not both")
for free, cap, gain in (
(self.free_rel, self.max_rel, self.requires_gain_rel),
(self.free, self.max, self.requires_gain),
):
if gain is not None and cap is None:
raise ValueError("regression requires_gain requires max in the same units")
if free is not None and cap is not None and free > cap:
raise ValueError("regression free must be <= max")
return self


class Benchmark(_StrictModel):
# Slug shape only: the name reaches branch names, ledger keys, and log
# labels — contract text must not shape refs or paths beyond a slug.
Expand Down Expand Up @@ -168,6 +204,8 @@ def measurement_signature(self) -> tuple:
so a future field joins the signature by default and the base-sync
skip fails toward re-measuring."""
data = self.model_dump()
if self.regression is None:
data.pop("regression")
# Preserve existing gate ledger signatures across the schema addition.
if self.verification == "gate":
data.pop("verification")
Expand Down Expand Up @@ -209,6 +247,23 @@ def _gpu_benchmarks_dispatch(self) -> Benchmark:
)
return self

@model_validator(mode="after")
def _regression_defaults(self) -> Benchmark:
r = self.regression
if r is not None and r.free is None and r.free_rel is None:
absolute = r.max is not None or r.requires_gain is not None
relative = r.max_rel is not None or r.requires_gain_rel is not None
if relative or (not absolute and self.min_delta_rel is not None):
defaults = {"free_rel": self.min_delta_rel or 0.0}
else:
defaults = {"free": self.min_delta or 0.0}
self.regression = Regression.model_validate(r.model_dump() | defaults)
return self

# Optional sibling-only policy; never changes this benchmark's win floor.
# Omission preserves the legacy gate and measurement signature.
regression: Regression | None = None

# Cross-seed noise floor. A comparison against the RECORDED best was
# measured under a different seed, so a delta inside the floor is noise,
# not progress; same-seed paired comparisons are exempt by construction.
Expand Down
120 changes: 99 additions & 21 deletions src/outerloop/orchestrator.py
Original file line number Diff line number Diff line change
Expand Up @@ -36,6 +36,7 @@
from outerloop.contract import (
Benchmark,
Contract,
Regression,
_fold,
load_contract,
normalize_path,
Expand Down Expand Up @@ -477,6 +478,7 @@ class SuiteMeasurement:
candidate: float
regressed: bool
display_digits: int | None = None
rule: str = "legacy-floor"


@dataclass(frozen=True)
Expand Down Expand Up @@ -544,7 +546,9 @@ def report(self, config: RunConfig, redact_secrets: tuple[str, ...] = ()) -> str
lines.append(f"Candidate{label}: {self.candidate}")
for row in self.suite:
verdict = "REGRESSED" if row.regressed else "ok"
lines.append(f"Suite {row.name}: {row.baseline} -> {row.candidate} ({verdict})")
lines.append(
f"Suite {row.name}: {row.baseline} -> {row.candidate} ({verdict}; {row.rule})"
)
if self.panel_rounds:
if self.panel_blocking_open:
state = "blocking findings OPEN at the cap"
Expand Down Expand Up @@ -736,19 +740,80 @@ def suite_regressed(
direction: str,
min_delta: float | None = None,
min_delta_rel: float | None = None,
*,
regression: Regression | None = None,
climbed_gain_rel: float | Fraction | None = None,
climbed_gain: float | Fraction | None = None,
) -> bool:
"""Did a sibling benchmark move the WRONG way beyond its own floor?

Both sides are same-seed paired, so with no floor declared any wrong-way
move counts (paired noise is ~0 by construction); a declared floor gives
a stochastic eval its honest tolerance. Non-finite values fail closed —
an unmeasurable sibling must never read as "no regression"."""
"""Whether a sibling violates its allowance (legacy floor when omitted)."""
return suite_regression_verdict(
baseline,
candidate,
direction,
min_delta,
min_delta_rel,
regression=regression,
climbed_gain_rel=climbed_gain_rel,
climbed_gain=climbed_gain,
)[0]


def suite_regression_verdict(
baseline: float,
candidate: float,
direction: str,
min_delta: float | None = None,
min_delta_rel: float | None = None,
*,
regression: Regression | None = None,
climbed_gain_rel: float | Fraction | None = None,
climbed_gain: float | Fraction | None = None,
) -> tuple[bool, str]:
"""Decision and rule for reports. New policy boundaries use exact decimal
arithmetic, like reaches_floor. A zero baseline cannot unlock relative
gain; relative sibling thresholds scale to zero. Hard caps win ties."""
if not (math.isfinite(baseline) and math.isfinite(candidate)):
return True
drop = baseline - candidate if direction == "max" else candidate - baseline
if drop <= 0:
return False
return drop > benchmark_floor(baseline, min_delta, min_delta_rel)
return True, "non-finite"
if regression is None:
# Keep the original floating-point comparison for existing contracts.
drop = baseline - candidate if direction == "max" else candidate - baseline
refused = drop > 0 and drop > benchmark_floor(baseline, min_delta, min_delta_rel)
return refused, "legacy-floor"
r = regression
relative = any(v is not None for v in (r.free_rel, r.max_rel, r.requires_gain_rel))
if not relative and all(v is None for v in (r.free, r.max, r.requires_gain)):
relative = min_delta_rel is not None
free = r.free_rel if relative else r.free
if free is None:
free = (min_delta_rel if relative else min_delta) or 0.0
cap = r.max_rel if relative else r.max
required = r.requires_gain_rel if relative else r.requires_gain
gain = climbed_gain_rel if relative else climbed_gain
if not all(
isinstance(v, Fraction) or math.isfinite(v)
for v in (free, cap, required, gain)
if v is not None
):
return True, "non-finite"
p, c = Fraction(repr(baseline)), Fraction(repr(candidate))
loss = p - c if direction == "max" else c - p
if loss <= 0:
return False, "no-regression"
scale = abs(p) if relative else Fraction(1)
if cap is not None and loss >= Fraction(repr(cap)) * scale:
return True, "hard-cap"
if loss <= Fraction(repr(free)) * scale:
return False, "free-allowance"
if cap is None:
return True, "free-exceeded"
if required is None:
return False, "unconditional-allowance"
if gain is None:
return True, "insufficient-gain"
exact_gain = gain if isinstance(gain, Fraction) else Fraction(repr(gain))
if exact_gain >= Fraction(repr(required)):
return False, "gain-unlocked"
return True, "insufficient-gain"


def improved(baseline: float, candidate: float, direction: str, min_rel: float) -> bool:
Expand Down Expand Up @@ -793,8 +858,8 @@ def make_task(
f"`{bench.command}` on a private seed to verify any improvement "
"claim, and the PR's CI runs the repository tests"
+ (
"; changes touching shared paths are suite-gated, so no sibling "
"benchmark may regress beyond its floor"
"; changes touching shared paths are suite-gated, so every sibling "
"benchmark must satisfy its regression policy (its floor by default)"
if suite_gated
else ""
)
Expand Down Expand Up @@ -1049,18 +1114,31 @@ def measure_and_decide(
run_seed=seed,
)

# Exact decimal gain avoids rounding an inclusive unlock boundary down.
main_base, main_cand = Fraction(repr(baseline)), Fraction(repr(candidate))
gain = main_cand - main_base if bench.direction == "max" else main_base - main_cand
gain_rel = gain / abs(main_base) if main_base else None
suite_rows: list[SuiteMeasurement] = []
for b in siblings:
sib_base = vals[f"sib-{b.name}-base"]
sib_cand = vals[f"sib-{b.name}-cand"]
refused, rule = suite_regression_verdict(
sib_base,
sib_cand,
b.direction,
b.min_delta,
b.min_delta_rel,
regression=b.regression,
climbed_gain_rel=gain_rel,
climbed_gain=gain,
)
suite_rows.append(
SuiteMeasurement(
name=b.name,
baseline=sib_base,
candidate=sib_cand,
regressed=suite_regressed(
sib_base, sib_cand, b.direction, b.min_delta, b.min_delta_rel
),
regressed=refused,
rule=rule,
display_digits=b.display_digits,
)
)
Expand Down Expand Up @@ -2526,13 +2604,13 @@ def pr_body(
suite_lines = [
"",
"Shared code was touched, so every sibling benchmark was re-measured "
"on both sides (paired seed): none regressed beyond its floor.",
"on both sides (paired seed): all passed their sibling regression policies.",
"",
"| suite benchmark | baseline | candidate |",
"| --- | --- | --- |",
"| suite benchmark | baseline | candidate | rule |",
"| --- | --- | --- | --- |",
] + [
f"| {row.name} | {fmt_metric(row.baseline, row.display_digits)} "
f"| {fmt_metric(row.candidate, row.display_digits)} |"
f"| {fmt_metric(row.candidate, row.display_digits)} | {row.rule} |"
for row in result.suite
]
if result.panel_blocking_open:
Expand Down
1 change: 1 addition & 0 deletions tests/fixtures/suite_measurement_legacy.json
Original file line number Diff line number Diff line change
@@ -0,0 +1 @@
{"name": "memory", "baseline": 100.0, "candidate": 100.0, "regressed": false, "display_digits": null}
Loading
Loading