Repository navigation
A PR is one idea, not one knob - #430
Conversation
The panel instruction, the verifier rubric, the author brief and the architecture doc now treat one idea, not one change, as the unit: an idea may bring the few changes it needs when the report attributes each, independent ideas bundled together stay blocking, and a new "landscape" category lets the panel block single-point tuning with no sweep or ablation behind it. Authors are told to sweep in one array launch and that a promising idea without a win is a success to report. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
There was a problem hiding this comment.
Round 1 — reviewed head c5aff9f4 — reviewer summarizer:hermes/gpt-5.6-terra over coverage+credentials+deployment+general+lifecycle+prose.
terra
Advisory findings from outerloop — the code owner decides. Reply to disagree; the outerloop:no-review label opts this PR out.
Verdict: nothing blocking — 2 advisory notes.
Advisory (non-blocking):
- Suggestion: Research-line brief still requires one clean contribution. [coverage+prose] A PR to main must be one idea, and may include the related changes it needs when the report gives each change’s own effect. (
src/outerloop/brief.py:280; high confidence) - Generated PRs still tell authors to use one hypothesis. [credentials] The changed policy allows one idea with several attributed changes, but generated PR bodies still say one hypothesis per PR, which can discourage the combinations this change permits. (
src/outerloop/orchestrator.py:2268; high confidence)
Verdict: two non-blocking policy-alignment changes remain. Rejected: none; the coverage and prose reports are merged because they identify the same stale one-contribution rule in the research-line brief, with prose supplying the required replacement wording.
…othesis Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
|
Both advisories batched in 86ca856: the research-line brief now says a PR to main is one idea with the few changes it needs, each change's effect in the report; generated PR bodies say one idea per PR. Pinned brief test updated. |
|
Compatibility statement (added for the 0.3.0rc1 release audit, per RELEASING.md).
|
Recent gpt-speedrun PRs show two patterns the old rule produced.
Owner decisions: a combination is fine when the joint effect is genuine and each single effect is documented. A single-point tuning PR may be blocking when the report gives no picture of the landscape (a sweep, neighbouring values, or an ablation); this is a judgment, not an absolute rule. Fundamental ideas are encouraged even when they don't yet beat the best.
What changes (wording and one additive category)
aggregationnow accepts a genuine combination with per-change effects. It still rejects an even blend of small tweaks with no stated mechanism.landscapecategory covers single-point hyperparameter or capacity changes with no sweep, neighbours or ablation behind them. It also asks a size-only win to argue why size is the right lever.Compatibility
Prompts and rubric text only, plus an additive finding category; a reader that does not know
landscapetreats it asother. No persisted state changes. It applies from the next panel read.🤖 Generated with Claude Code