Skip to content

feat(bench): correction rounds and multi-reference paired lift for the agent benchmark - #494

Merged
Teakowa merged 2 commits into
mainfrom
feat/466-bench-lift
Oct 4, 2026
Merged

Teakowa merged 2 commits into
mainfrom
feat/466-bench-lift

Conversation

@e54-bot

@e54-bot e54-bot commented Oct 3, 2026 •

Copy link
Copy Markdown
Collaborator

Refs #466. Stacked on #492 (feat/474-mcp-level).

Three additions to the agent benchmark from the issue's punch list, plus the transport normalization they need.

Correction rounds (correctionRounds)

Counts failed validation -> workspace edit rounds per trial: the condition tool's validating op reporting exit 1 followed by an edit (snapshot of a watched file changing).

  • Wright: check/lint/analyze/compile — CLI calls under bin, tools/call serve ops under mcp/stdio/jsonrpc. overpy compile counts under the opy cell, so the comparison stays fair across conditions.
  • Consecutive failures before one edit count once; a pass resets the sequence.
  • Usage errors (exit 2), crashes (exit 3/4), serve refusals, protocol errors, and unanswered requests are all neutral — they neither count as a correction nor clear a pending failure.
  • Reported as a corr column in the outcome table and correctionRounds per cell in summary.json.

serve_result_payload normalizes op payloads across transports: {"result": envelope} for stdio/jsonrpc, result.content[] text blocks for MCP; isError refusal payloads ({code, message}, no exit) are detected and skipped.

Repeatable --reference on report and evaluate

--reference LABEL is now repeatable: the report emits one Paired against \LABEL`section per named reference, so lift can be measured againstdocs, a skill-only control, or any other named cell — not only none/none/off`. Default is unchanged (baseline only).

workshop-wiki → workshop-skill rename

The generated wiki skill was named workshop-wiki while the condition vocabulary (SKILLS, cells, docs) calls it workshop-skill — the generated artifact never installed under the name the harness records. SKILL_NAME, the output-dir requirement, BUILD.json identity, help, examples, docs, and fixtures now agree.

(Already in place, per the punch list: clustered-bootstrap intervals via rate_runs/cluster_interval, and model/effort/skill-hash/tool-list recording via agent.id + agentInfo + environment.skills.)

Verification

  • 111 unittest tests pass with WRIGHT_BIN (correction-round coverage for CLI calls, stdio serve ops, MCP payload unwraps, overpy compile, refusal/error/unanswered neutrality, multi-reference pairing, and the renamed skill fixtures).
  • Scripted agents through the real harness: a deliberately failed check then repair yields correctionRounds: 1 under wright/none/off and 0 under none/none/off; a two-reference report renders both paired sections.
  • git diff --check clean.

Independent review

One review pass done on the metrics commit: found that non-0/1 exits cleared pending state like a pass (fixed — only exits 0/1 mark a verdict now), plus nits applied (docstring lag caveat, payload.get("exit"), wording). Deferred pre-existing finding: the jsonrpc transport maps compile requests to jsonrpc:compile rather than the compile op (predates this diff, also affects toolUse/friction/E02/E04) — worth a follow-up issue.

Limits

  • Edit markers use the snapshot poller's detection time (~0.2s granularity); sub-interval edit/validation orderings can merge rounds.
  • correctionRounds observes only watched files (the scenario watch list); edits elsewhere are invisible.

@Teakowa
Teakowa added this pull request to stack #495 October 3, 2026 17:33
@e54-bot
e54-bot force-pushed the feat/466-bench-lift branch from 1a80e55 to a56ad3c Compare October 4, 2026 03:18
@e54-bot
e54-bot force-pushed the feat/466-bench-lift branch from a56ad3c to 31c9854 Compare October 4, 2026 03:54
@e54-bot
e54-bot force-pushed the feat/466-bench-lift branch from 31c9854 to 160405d Compare October 4, 2026 04:09
@e54-bot
e54-bot force-pushed the feat/466-bench-lift branch from 160405d to 938394a Compare October 4, 2026 04:51
@e54-bot
e54-bot force-pushed the feat/466-bench-lift branch from 938394a to 3c5abb4 Compare October 4, 2026 05:09
@e54-bot
e54-bot force-pushed the feat/466-bench-lift branch from 3c5abb4 to 57c7421 Compare October 4, 2026 05:17
@e54-bot
e54-bot force-pushed the feat/466-bench-lift branch from 57c7421 to 71302e9 Compare October 4, 2026 05:23
@e54-bot
e54-bot force-pushed the feat/466-bench-lift branch from 71302e9 to 8398766 Compare October 4, 2026 06:02
@Teakowa
Teakowa force-pushed the feat/466-bench-lift branch from 8398766 to fd358a0 Compare October 4, 2026 06:06
@e54-bot
e54-bot force-pushed the feat/466-bench-lift branch from fd358a0 to f0a6140 Compare October 4, 2026 06:15
Base automatically changed from feat/474-mcp-level to main October 4, 2026 07:18
…erence

Refs #466. Two metrics additions for the effectiveness benchmark:

- correctionRounds: failed-validation -> workspace-edit rounds per trial.
  The condition tool's validating op (wright check/lint/analyze/compile as
  CLI calls or serve ops on any transport, overpy compile under the opy
  cell) reporting exit 1 followed by an edit counts once; consecutive
  failures merge, a pass resets, and usage errors, refusals, crashes, and
  protocol errors are neutral. Reported as a corr column and per-cell mean.
- Repeatable --reference on report and evaluate: emits one paired section
  per named reference label, so lift can be measured against docs, skill,
  or any other control cell, not only none/none/off.

serve_result_payload normalizes op payloads across stdio, jsonrpc, and MCP
(content[] text blocks; isError refusals carry no exit and are neutral).

Verified end to end with scripted agents through the real harness: a
failed check then repair yields correctionRounds 1 under wright and 0
under none, and both requested paired sections render.
Refs #466. The generated progressive-disclosure skill was named
workshop-wiki while the condition vocabulary, cells, and docs all call it
workshop-skill, so the generated artifact never installed under the name
the harness and result records expect. SKILL_NAME, the output-directory
requirement, BUILD.json identity, help text, examples, docs, and test
fixtures now agree on workshop-skill.
@Teakowa
Teakowa force-pushed the feat/466-bench-lift branch from f0a6140 to cfb4808 Compare October 4, 2026 07:18
@Teakowa
Teakowa merged commit 92179a2 into main Oct 4, 2026
3 checks passed
@Teakowa
Teakowa deleted the feat/466-bench-lift branch October 4, 2026 07:18
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

Development

Successfully merging this pull request may close these issues.

2 participants