Skip to content

Simplify operational-value grading to one-shot evaluators - #60682

Merged
mnkiefer merged 8 commits into
mainfrom
new-value-skill
Sep 14, 2026
Merged

mnkiefer merged 8 commits into
mainfrom
new-value-skill

Conversation

@mnkiefer

@mnkiefer mnkiefer commented Sep 13, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

  • Replace historical replay, maturation, baseline, cache, registry, and report machinery with a deterministic one-shot Bash evaluator contract.
  • Freeze inline or file-backed evaluator bytes and their SHA-256 digest into each compiled workflow run, with strict metric-array validation and size limits.
  • Supply current-run validated safe-output requests as request.outputs while preserving the distinction between requested and applied mutations.
  • Refocus the operational-value designer on ultimate goals and the strongest observable, attributable, independently verifiable outcome; remove domain-specific metric templates.
  • Add a Daily File Diet evaluator that independently checks the largest non-test Go file at the run SHA against the workflow's 800-line threshold and validates the required issue/noop decision.

Why

Operational value should measure the strongest effect a run can actually prove at grading time. Workflow-specific semantics now live in frozen evaluators instead of generic historical infrastructure, avoiding hindsight tuning and claims about future or unapplied outcomes.

Validation

  • Operational-value designer tests and deterministic evaluator fixtures pass.
  • Bash syntax and prospective contract-change verification pass.
  • Historical Daily File Diet run Test aliases #263 evaluates to file-diet-decision-conformance=1 and largest-file-under-threshold=0 under the new evaluator.
  • All 299 workflows recompile successfully with lock files in sync.
  • make agent-report-progress passes.

…on, improve metric validation, and add comprehensive metric patterns

- Updated the operational-value-designer skill description and metadata for clarity and versioning.
- Refined the design procedure to emphasize the importance of semantic translation and evaluator specificity.
- Introduced a new references file containing operational value metric patterns to guide metric design.
- Enhanced the verification script to enforce stricter output validation and ensure compliance with the new evaluator structure.
- Improved test coverage for the evaluator, including checks for invalid outputs and oversized responses.
Copilot AI balanced review requested due to automatic review settings September 13, 2026 19:49
@mnkiefer mnkiefer self-assigned this Sep 13, 2026
@mnkiefer
mnkiefer marked this pull request as draft September 13, 2026 19:49

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟡 Changes recommended

The new evaluator contract is incompatible with current runtime consumers, and stdout validation has correctness gaps.

Get a fresh assessment by requesting another Copilot review.

Review tier: Balanced
Findings: 1 High severity · 2 Medium severity

Open findings (3)
What changed in this PR

Refactors the operational-value designer around workflow-specific, deterministic per-run metrics.

Changes:

  • Replaces the legacy evaluator contract with an ordered metric-array design.
  • Adds domain-specific metric patterns.
  • Strengthens evaluator validation and negative tests.
File Description
SKILL.md Defines the revised design process and evaluator contract.
references/​metric-patterns.md Adds metric-design examples.
scripts/​verify-operational-value-evaluator.sh Validates the new output format.
tests/​test.sh Adds invalid and oversized-output cases.

Comment thread .github/skills/operational-value-designer/SKILL.md
Comment thread .github/skills/operational-value-designer/tests/test.sh Outdated
Comment thread .github/skills/operational-value-designer/SKILL.md
@github-actions

This comment has been minimized.

@github-actions

Copy link
Copy Markdown
Contributor

Nice work on the operational-value-designer refactor! 🎯

This PR looks well-structured and focused on improving the skill's documentation, testing, and verification infrastructure. As a core team member, you've followed the appropriate process, and the changes are clearly scoped to a single skill with strong test coverage.

Key strengths:

  • ✅ All changes contained within .github/skills/operational-value-designer/
  • ✅ New reference file (metric-patterns.md) provides helpful guidance
  • ✅ Test coverage improvements with invalid output and size checks
  • ✅ Clearer skill metadata and refined design procedures
  • ✅ Clear PR description with actionable bullet points

Noted: The PR is currently marked as draft with "Do not merge" label, which is appropriate for WIP status. Ready for team review when you mark it ready.

Generated by ✅ Contribution Check · copilot · auto · 41.2 AIC · ⌖ 6.22 AIC · ⊞ 9.4K · ◷

@mnkiefer mnkiefer changed the title Refactor operational-value-designer skill Simplify operational-value grading to one-shot evaluators Sep 14, 2026
@mnkiefer
mnkiefer marked this pull request as ready for review September 14, 2026 10:50
@github-actions

github-actions Bot commented Sep 14, 2026 •

Copy link
Copy Markdown
Contributor

🧠 Matt Pocock Skills Reviewer has completed the skills-based review. ✅

Warning

Firewall blocked 1 domain

The following domain was blocked by the firewall during workflow execution:

  • github.com

To allow these domains, add them to the network.allowed list in your workflow frontmatter:

network:
  allowed:
    - defaults
    - "github.com"

See Network Configuration for more information.

🧠 Reviewed using Matt Pocock's skills by Matt Pocock Skills Reviewer

@github-actions

github-actions Bot commented Sep 14, 2026 •

Copy link
Copy Markdown
Contributor

✅ Test Quality Sentinel completed test quality analysis.

Test Quality Sentinel skipped because pre-fetch PR data was unavailable: unable to fetch test file diff

🧪 Test quality analysis by Test Quality Sentinel

@github-actions

github-actions Bot commented Sep 14, 2026 •

Copy link
Copy Markdown
Contributor

⚠️ Security scanning failed for Design Decision Gate 🏗️. Review the logs for details.

🏗️ ADR gate enforced by Design Decision Gate 🏗️

@github-actions

github-actions Bot commented Sep 14, 2026 •

Copy link
Copy Markdown
Contributor

✅ Ponytail Reviewer completed successfully!

Warning

Threat Detection Engine Failure — The analysis engine could not complete. This is a tooling failure, not a security finding.

What happened

The threat detection engine failed to produce results.

Review the workflow run logs for details.

Warning

Firewall blocked 1 domain

The following domain was blocked by the firewall during workflow execution:

  • ab.chatgpt.com

To allow these domains, add them to the network.allowed list in your workflow frontmatter:

network:
  allowed:
    - defaults
    - "ab.chatgpt.com"

See Network Configuration for more information.

Generated by Ponytail Reviewer for #60682

@github-actions

github-actions Bot commented Sep 14, 2026 •

Copy link
Copy Markdown
Contributor

⚠️ PR Code Quality Reviewer failed during code quality review.

Warning

Threat Detection Engine Failure — The analysis engine could not complete. This is a tooling failure, not a security finding.

What happened

The threat detection engine failed to produce results.

Review the workflow run logs for details.

🔎 Code quality review by PR Code Quality Reviewer

@github-actions

Copy link
Copy Markdown
Contributor
🏗️ ADR Required - draft added

An ADR-backed design decision is required for this PR because the prefetch summary shows 550 added lines in default business-logic directories, which is above the 100-line enforcement threshold.

Evidence used

  • /tmp/gh-aw/agent/adr-prefetch-summary.json: requires_adr_by_default_volume=true, default_business_additions=550
  • /tmp/gh-aw/agent/pr.json: PR title and summary describe a repo-wide architectural change from historical operational-value grading to one-shot evaluators
  • /tmp/gh-aw/agent/pr-files.json and /tmp/gh-aw/agent/pr.diff: changes rewrite the operational-value designer contract, verifier scripts, tests, and Daily File Diet evaluator
  • docs/adr/: no ADR linked in the PR body and no ADR on branch for this decision before this review

Action taken

  • Added draft ADR: docs/adr/60682-simplify-operational-value-grading-to-one-shot-evaluators.md

Next action for the author

  • Review and refine the draft ADR so it accurately captures the final decision and trade-offs for this PR.

🏗️ ADR gate enforced by Design Decision Gate 🏗️ · pi · gpt54 · 25.5 AIC · ⊞ 9.9K · ◷
Comment /review to run again

@github-actions github-actions Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Skills-Based Review 🧠

Applied /codebase-design and /diagnosing-bugs. The architectural simplification (removing ~2,900 lines of historical replay/maturation/baseline/cache machinery in favor of a frozen one-shot evaluator contract) is well executed and consistently threaded through the compiler, CLI runner, JS grader shim, docs, and spec. Requesting changes on two regressions that reintroduce bugs a prior review already flagged on this branch.

📋 Key Themes & Highlights

Key Themes

  • Regressed fixes: two inline comments call out spots where earlier Copilot review feedback (unquoted $skill_dir, NUL-byte-unsafe command substitution in verify-operational-value-evaluator.sh) was not carried into the final diff.
  • Contract consistency: the new {schemaVersion, run, event, outputs, config} request shape and [{id, value}] metric array are validated symmetrically in Go (pkg/cli/graders_run.go, pkg/workflow/graders_operational_value.go), JS (operational_value_grader.cjs), and the verification/test scripts — good deep-module boundary with a single, simple interface replacing several prior modes.
  • Test coverage: graders_operational_value_test.go, graders_run_test.go, and operational_value_grader.test.cjs were all updated to match the new contract with solid edge-case coverage (empty arrays, duplicate/empty ids, out-of-range values, timeouts).

Positive Highlights

  • ✅ AGENTS.md rule 10 codifies the new evaluator-change contract (co-locate evaluator + workflow changes, verify-operational-value-contract-change.sh), keeping governance in sync with the code change.
  • ✅ Docs (graders-specification.md, cli.md) were updated in the same commit to drop dead concepts (evidence cutoff, opportunity keys, historical regrading) rather than leaving stale references.
  • ✅ Daily File Diet evaluator rewrite is much simpler and self-documenting (comment block states metric semantics up front).

Note: the pre-fetched diff was capped at 3000 lines, so some Go files (pkg/cli/graders_run.go, pkg/workflow/graders_config.go, pkg/workflow/graders_operational_value.go, associated tests) were reviewed directly from the working tree rather than the diff; no additional issues were found there.

Warning

Firewall blocked 1 domain

The following domain was blocked by the firewall during workflow execution:

  • github.com

To allow these domains, add them to the network.allowed list in your workflow frontmatter:

network:
  allowed:
    - defaults
    - "github.com"

See Network Configuration for more information.

🧠 Reviewed using Matt Pocock's skills by Matt Pocock Skills Reviewer · copilot · sonnet50 · 112.1 AIC · ⌖ 15.1 AIC · ⊞ 10.4K
Comment /matt to run again

Comment thread .github/skills/operational-value-designer/tests/test.sh Outdated

@github-actions github-actions Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Applied Impeccable distill (large refactor/cleanup removing historical replay machinery) and audit (technical/correctness quality) modes given this is a refactor_cleanup change.

The core simplification is well-executed: the new one-shot evaluator contract, fixture-driven verification, and updated spec/docs are consistent with each other, and the daily-file-diet evaluator + fixtures pass local verification. Two issues found tied to the refactor itself:

  1. Dead env var (pkg/workflow/compiler_yaml_graders.go:68) — GH_AW_RUN_CREATED_AT is still emitted into every compiled workflow even though the JS side (trace_graders.cjs, operational_value_grader.cjs) no longer reads or forwards createdAt anywhere. Leftover wiring from the removed maturation/evidence-cutoff machinery.
  2. Regression: unquoted variable (.github/skills/operational-value-designer/tests/test.sh:23) — this PR changed a previously-quoted "$skill_dir" invocation to an unquoted $skill_dir expansion, inconsistent with every other call in the same script and unsafe for paths with spaces/glob characters.

Requesting changes to clean up both before merge; neither is a security risk but both are avoidable defects introduced directly by this diff.

Warning

Firewall blocked 1 domain

The following domain was blocked by the firewall during workflow execution:

  • github.com

To allow these domains, add them to the network.allowed list in your workflow frontmatter:

network:
  allowed:
    - defaults
    - "github.com"

See Network Configuration for more information.

🧵 Reviewed using Impeccable skills by Impeccable Skills Reviewer · copilot · sonnet50 · 234.5 AIC · ⌖ 13.9 AIC · ⊞ 8.4K

Comments that could not be inline-anchored

pkg/workflow/compiler_yaml_graders.go:68

This env var is now dead: trace_graders.cjs (this PR) removed process.env.GH_AW_RUN_CREATED_AT / the operationalValueRunMetadata plumbing (main no longer builds or forwards runMetadata), and operational_value_grader.cjs's buildRunSubject(env) no longer accepts or reads metadata.createdAt. Keeping this GH_TOKEN/GH_AW_RUN_CREATED_AT env block still wires needs.activation.outputs.run_created_at into every compiled workflow (e.g. .github/workflows/daily-file-diet.lock.yml) f…

.github/skills/operational-value-designer/tests/test.sh:23

This line was changed by this PR from the quoted form path=$("$skill_dir/scripts/...") to an unquoted $skill_dir expansion: evaluator_path=$($skill_dir/scripts/operational-value-evaluator-path.sh daily-file-diet). Every other invocation in this script quotes "$skill_dir". If the repository checkout path contains whitespace or glob metacharacters, this line will word-split/glob and fail or run the wrong command. Please quote it: `evaluator_path=$("$skill_dir/scripts/operational-value-eva…

@mnkiefer
mnkiefer enabled auto-merge (squash) September 14, 2026 11:52
@mnkiefer
mnkiefer merged commit 70581f4 into main Sep 14, 2026
45 checks passed
@mnkiefer
mnkiefer deleted the new-value-skill branch September 14, 2026 11:58
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants