feat: restore tier-3 eval coverage across eight plugins#193
Conversation
Recreates the eval/test-coverage attrition from the salvage sweep (items #13, #15-#23 plus the flagged review-toolkit and docs-hygiene items), genericized and validated against evals.schema.json: - songwriting 0.4.0: 13 behavioral evals mapped onto the multi-skill split - ai-briefing 0.4.0: 6 engine evals + 3 synthetic fixtures via the audience-defaults seam - event-storming 0.4.0: offline board-export eval + fixture for --discover-bcs (disjoint from the live-Miro eval) - codebase-audit 0.3.0: scope-boundary routing eval - discovery 0.5.0: research floor-scaling + broad-topic-minimums evals - source-control 0.2.0: readiness security-gate (+fixture), mixed-actor, and three worktree evals - docs-hygiene 0.4.0: self-contained compress/declutter fixtures (empirically verified against detect.sh), rename-references eval-case clauses - review-toolkit 0.6.0: fanout evals 6 -> 20 (dedup/severity, fix-pass safety fence, run-everything reconciliation/ordering) Per-plugin version bumps + CHANGELOG entries; all validations pass. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_011xHkkNc7CR98L8Xz9Mu7ZA
|
Claude finished @kyle-sexton's task in 0s —— View job I'll analyze this and get back to you. |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 617d750f2a
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
…contradiction, offline-export branch - source-control: grader ground-truth removed from the security fixture (already covered by eval 6 expectations); eval 7 no longer gates the required human classification reply — only reaction/fix pause - event-storming: --discover-bcs protocol documents the supplied-export input path (no Miro) alongside the live board-URL read; board-write step marked live-path-only Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_011xHkkNc7CR98L8Xz9Mu7ZA
|
Claude finished @kyle-sexton's task in 0s —— View job I'll analyze this and get back to you. |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 8b945f4a9f
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
- review-toolkit: evals 7-20 files arrays held scenario prose, not fixture paths; prose folded into prompts, arrays emptied - discovery: comparison eval targets real tools (RabbitMQ/Kafka/NATS) so primary-source expectations are satisfiable - source-control: evals.json rewritten with LF + final newline (editorconfig hygiene failure from the previous rewrite) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_011xHkkNc7CR98L8Xz9Mu7ZA
|
Claude finished @kyle-sexton's task in 0s —— View job I'll analyze this and get back to you. |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 0f5f761a69
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
- discovery: production-readiness eval names FusionCache so primary-source expectations are satisfiable - source-control: cleanup --dry-run now specifies git worktree prune --dry-run (report-only) in Step 1; eval 6 wording aligned. A dry-run that mutates worktree metadata violated its own contract Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_011xHkkNc7CR98L8Xz9Mu7ZA
|
Claude finished @kyle-sexton's task in 0s —— View job I'll analyze this and get back to you. |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 0a333dd9fe
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
Fixture checks are all terminal, so Gate 1 is satisfied; only Gates 2-3 block. Expectation narrowed accordingly. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_011xHkkNc7CR98L8Xz9Mu7ZA
|
Claude finished @kyle-sexton's task in 1s —— View job I'll analyze this and get back to you. |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: fdd0ed998b
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
Eval 6 required classification while asserting Gate 2 open. NOT-ready now rests on Gates 3 and 5 only. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_011xHkkNc7CR98L8Xz9Mu7ZA
|
Claude finished @kyle-sexton's task in 0s —— View job I'll analyze this and get back to you. |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 405d82ee97
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
Fixture preamble narrated the expected verdict to the skill under test. Fixture is raw gh output only; the cue lives in expected_output now. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_011xHkkNc7CR98L8Xz9Mu7ZA
|
Claude finished @kyle-sexton's task in 0s —— View job I'll analyze this and get back to you. |
Summary
Restores the eval/test-coverage attrition identified by the salvage sweep (tier-3: items #13, #15–#23 plus the two flagged cosmetic-severity items), recreated genericized and validated against
plugins/skill-quality/reference/evals.schema.json. Per-plugin version bumps + CHANGELOG entries.--discover-bcs— disjoint from the live-Miro-required eval, reconciliation noted inline/claude-config-audit:settings-auditdetect.sh; rename-references "add an eval case" clausesVerification
claude plugin validate;scripts/validate-plugin-contracts.mjspasses (1311 files)🤖 Generated with Claude Code
https://claude.ai/code/session_011xHkkNc7CR98L8Xz9Mu7ZA