Skip to content

Measurement provenance in PR bodies and judge briefs - #440

Merged
renmengye merged 2 commits into
mainfrom
fix/measurement-provenance
Sep 29, 2026
Merged

renmengye merged 2 commits into
mainfrom
fix/measurement-provenance

Conversation

@renmengye

@renmengye renmengye commented Sep 29, 2026 •

Copy link
Copy Markdown
Member

In a judge evaluation, every configuration we tested misread one benchmark PR the same way: a reporting floor existed in older trees and was removed on main, and judges compared the author's older-tree experiment numbers with the orchestrator's current-tree measurement as if comparable. The PR body gave them no way to tell which tree each number came from.

This PR gives judges the evidence rather than a rule:

  • Measured names the base and candidate commits the orchestrator measured, and that both used the same eval command read from the base tree.
  • Experiments shows the commit each launch ran on (the sealed tree passed to the launcher); older ledger rows show unknown.
  • The verify and review briefs (bare and agentic, which wrap them) carry one sentence: numbers measured on different commits are not directly comparable; check the commits in the checkout history before attributing a gap.

No change to how measurement works.

Compatibility (RELEASING.md)

The only persisted schema change is an optional commit string on submitted
rows in launches.jsonl, taken from the sealed tree passed to the launcher.
History and experiment rendering tolerate all prior rows without that field,
including ended runs; these display unknown rather than guessing a commit.
Existing submission identities remain authoritative on retries.

On the first tick after upgrade, in-flight runs continue normally and existing
PR bodies remain unchanged. Newly created PR bodies include measurement and
launch provenance. No backfill or migration is required. Rolling back is safe:
older ledger readers ignore the additive field. Run records are unchanged.

tests/fixtures/launches-before-provenance.jsonl captures the prior submitted
and ended row format; its test covers repeated reads, an idempotent re-park,
and retry after an interrupted append alongside new records.

Tests

PR body with provenance in Measured; experiments tables with and without recorded commits (legacy fixture); both briefs carry the caution. Mutation check: dropping the SHAs from Measured fails the provenance test. Gate: 2397 passed, 2 skipped; ruff check, ruff format --check, mypy clean.

Built by codex from my brief; codex gate plus my cross-review (all brief paths covered, ledger change additive, compatibility note moved from a new docs page into this PR body per RELEASING.md).

🤖 Generated with Claude Code

…ges told to check commits before comparing numbers

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

@github-actions github-actions Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Round 1 — reviewed head 273031b3 — reviewer summarizer:hermes/gpt-5.6-terra over coverage+credentials+deployment+general+lifecycle+prose.

terra
Advisory findings from outerloop — the code owner decides. Reply to disagree; the outerloop:no-review label opts this PR out.

Verdict: 1 blocking, 0 advisory.

1 finding attached to the lines below.

One blocking finding: the stale-base checkpoint can record an unsubmitted sibling launch against candidate_sha, producing false PR provenance. Rejected findings: none; the remaining lenses reported no findings.

Comment thread src/outerloop/attempt.py Outdated
…eirs)

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@renmengye renmengye added the autoresearch:review Request a fresh advisory review of this PR's current state label Sep 29, 2026

@github-actions github-actions Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Round 1 — reviewed head 928245ac — reviewer hermes/gpt-5.6-terra.

terra
Advisory findings from outerloop — the code owner decides. Reply to disagree; the outerloop:no-review label opts this PR out.

Verdict: no defects found.

@renmengye
renmengye merged commit 3ff1a1d into main Sep 29, 2026
5 checks passed
@renmengye
renmengye deleted the fix/measurement-provenance branch September 29, 2026 00:33
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

autoresearch:review Request a fresh advisory review of this PR's current state

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant