Problem
The M4 qualification gate was one-sided: a task passed after at least 4 re-proposals in 6 comparator runs. That fixed M1/M2's empty instrument but selected the opposite failure mode.
Four of M4's eight tasks were 7/7 in both measured arms and contributed no discordance. Four of the five candidates that had scored 6/6 in qualification became those saturated tasks. The strongest qualification result was therefore evidence of no headroom, not a better instrument.
Registered rule for the next measurement
Run six qualification seeds on the registered primary comparator arm and qualify a task only at 4 or 5 re-proposals out of 6.
- The floor of 4 is retained from M4: M1/M2 had seven of ten control tasks at zero and an aggregate base rate near 20%, so a lower floor returns to an empty instrument.
- The ceiling of 5 is the first point below the observed failure boundary. It requires at least one measured non-event; 6/6 is excluded because M4 showed that boundary saturating in both arms.
- If fewer than the preregistered minimum number of tasks qualify, the experiment stops. The bounds are not relaxed after seeing qualification.
- The treatment arm is neither run nor inspected during qualification.
Applied to M4's qualification data, only the 5/6 and two 4/6 candidates would pass, so the six-task proceed gate would stop the run. This is a design diagnostic, not permission to drop the four saturated tasks from the completed M4 analysis.
Acceptance criteria
- Qualification logic enforces the inclusive 4–5/6 window on the primary comparator.
- The proceed gate fails closed when too few tasks qualify.
- Qualification output includes Wilson or Clopper-Pearson intervals for every task proportion.
- Tests cover empty (0/6), lower boundary (4/6), upper boundary (5/6), and saturated (6/6) tasks.
Problem
The M4 qualification gate was one-sided: a task passed after at least 4 re-proposals in 6 comparator runs. That fixed M1/M2's empty instrument but selected the opposite failure mode.
Four of M4's eight tasks were 7/7 in both measured arms and contributed no discordance. Four of the five candidates that had scored 6/6 in qualification became those saturated tasks. The strongest qualification result was therefore evidence of no headroom, not a better instrument.
Registered rule for the next measurement
Run six qualification seeds on the registered primary comparator arm and qualify a task only at 4 or 5 re-proposals out of 6.
Applied to M4's qualification data, only the 5/6 and two 4/6 candidates would pass, so the six-task proceed gate would stop the run. This is a design diagnostic, not permission to drop the four saturated tasks from the completed M4 analysis.
Acceptance criteria