feat(bench): correction rounds and multi-reference paired lift for the agent benchmark - #494
Merged
Merged
Conversation
Teakowa
added this pull request to stack #495
October 3, 2026 17:33
e54-bot
force-pushed
the
feat/466-bench-lift
branch
from
October 4, 2026 03:18
1a80e55 to
a56ad3c
Compare
e54-bot
force-pushed
the
feat/466-bench-lift
branch
from
October 4, 2026 03:54
a56ad3c to
31c9854
Compare
e54-bot
force-pushed
the
feat/466-bench-lift
branch
from
October 4, 2026 04:09
31c9854 to
160405d
Compare
e54-bot
force-pushed
the
feat/466-bench-lift
branch
from
October 4, 2026 04:51
160405d to
938394a
Compare
e54-bot
force-pushed
the
feat/466-bench-lift
branch
from
October 4, 2026 05:09
938394a to
3c5abb4
Compare
e54-bot
force-pushed
the
feat/466-bench-lift
branch
from
October 4, 2026 05:17
3c5abb4 to
57c7421
Compare
e54-bot
force-pushed
the
feat/466-bench-lift
branch
from
October 4, 2026 05:23
57c7421 to
71302e9
Compare
e54-bot
force-pushed
the
feat/466-bench-lift
branch
from
October 4, 2026 06:02
71302e9 to
8398766
Compare
Teakowa
force-pushed
the
feat/466-bench-lift
branch
from
October 4, 2026 06:06
8398766 to
fd358a0
Compare
Teakowa
approved these changes
Oct 4, 2026
e54-bot
force-pushed
the
feat/466-bench-lift
branch
from
October 4, 2026 06:15
fd358a0 to
f0a6140
Compare
…erence Refs #466. Two metrics additions for the effectiveness benchmark: - correctionRounds: failed-validation -> workspace-edit rounds per trial. The condition tool's validating op (wright check/lint/analyze/compile as CLI calls or serve ops on any transport, overpy compile under the opy cell) reporting exit 1 followed by an edit counts once; consecutive failures merge, a pass resets, and usage errors, refusals, crashes, and protocol errors are neutral. Reported as a corr column and per-cell mean. - Repeatable --reference on report and evaluate: emits one paired section per named reference label, so lift can be measured against docs, skill, or any other control cell, not only none/none/off. serve_result_payload normalizes op payloads across stdio, jsonrpc, and MCP (content[] text blocks; isError refusals carry no exit and are neutral). Verified end to end with scripted agents through the real harness: a failed check then repair yields correctionRounds 1 under wright and 0 under none, and both requested paired sections render.
Refs #466. The generated progressive-disclosure skill was named workshop-wiki while the condition vocabulary, cells, and docs all call it workshop-skill, so the generated artifact never installed under the name the harness and result records expect. SKILL_NAME, the output-directory requirement, BUILD.json identity, help text, examples, docs, and test fixtures now agree on workshop-skill.
Teakowa
force-pushed
the
feat/466-bench-lift
branch
from
October 4, 2026 07:18
f0a6140 to
cfb4808
Compare
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Refs #466. Stacked on #492 (
feat/474-mcp-level).Three additions to the agent benchmark from the issue's punch list, plus the transport normalization they need.
Correction rounds (
correctionRounds)Counts
failed validation -> workspace editrounds per trial: the condition tool's validating op reportingexit1 followed by an edit (snapshot of a watched file changing).check/lint/analyze/compile— CLI calls underbin,tools/callserve ops undermcp/stdio/jsonrpc.overpy compilecounts under theopycell, so the comparison stays fair across conditions.corrcolumn in the outcome table andcorrectionRoundsper cell insummary.json.serve_result_payloadnormalizes op payloads across transports:{"result": envelope}for stdio/jsonrpc,result.content[]text blocks for MCP;isErrorrefusal payloads ({code, message}, noexit) are detected and skipped.Repeatable
--referenceonreportandevaluate--reference LABELis now repeatable: the report emits onePaired against \LABEL`section per named reference, so lift can be measured againstdocs, a skill-only control, or any other named cell — not onlynone/none/off`. Default is unchanged (baseline only).workshop-wiki→workshop-skillrenameThe generated wiki skill was named
workshop-wikiwhile the condition vocabulary (SKILLS, cells, docs) calls itworkshop-skill— the generated artifact never installed under the name the harness records.SKILL_NAME, the output-dir requirement,BUILD.jsonidentity, help, examples, docs, and fixtures now agree.(Already in place, per the punch list: clustered-bootstrap intervals via
rate_runs/cluster_interval, and model/effort/skill-hash/tool-list recording viaagent.id+agentInfo+environment.skills.)Verification
WRIGHT_BIN(correction-round coverage for CLI calls, stdio serve ops, MCP payload unwraps, overpy compile, refusal/error/unanswered neutrality, multi-reference pairing, and the renamed skill fixtures).checkthen repair yieldscorrectionRounds: 1underwright/none/offand0undernone/none/off; a two-referencereportrenders both paired sections.git diff --checkclean.Independent review
One review pass done on the metrics commit: found that non-0/1 exits cleared pending state like a pass (fixed — only exits 0/1 mark a verdict now), plus nits applied (docstring lag caveat,
payload.get("exit"), wording). Deferred pre-existing finding: the jsonrpc transport mapscompilerequests tojsonrpc:compilerather than thecompileop (predates this diff, also affectstoolUse/friction/E02/E04) — worth a follow-up issue.Limits
correctionRoundsobserves only watched files (the scenariowatchlist); edits elsewhere are invisible.