Skip to content

feat: restore tier-3 eval coverage across eight plugins#193

Merged
kyle-sexton merged 7 commits into
mainfrom
restore/tier3-evals
Jul 15, 2026
Merged

feat: restore tier-3 eval coverage across eight plugins#193
kyle-sexton merged 7 commits into
mainfrom
restore/tier3-evals

Conversation

@kyle-sexton

Copy link
Copy Markdown
Contributor

Summary

Restores the eval/test-coverage attrition identified by the salvage sweep (tier-3: items #13, #15#23 plus the two flagged cosmetic-severity items), recreated genericized and validated against plugins/skill-quality/reference/evals.schema.json. Per-plugin version bumps + CHANGELOG entries.

Plugin Version Restored coverage
songwriting 0.4.0 All 13 medley behavioral evals mapped onto the multi-skill split (workflow 3, diagnosis 3, rhyme 2, song-form 2, co-write 2, object-writing 1); zero drops
ai-briefing 0.4.0 6 engine evals + 3 synthetic fixtures via the audience-defaults seam; the legacy Grok-preload case re-derived as an unreachable-RSS visible-degrade scenario (the flag never shipped; CI contract bans the token)
event-storming 0.4.0 Offline board-export eval (id 8) + 74-line fixture for --discover-bcs — disjoint from the live-Miro-required eval, reconciliation noted inline
codebase-audit 0.3.0 Scope-boundary eval: decline settings/MCP/hooks claim-extraction, route to /claude-config-audit:settings-audit
discovery 0.5.0 Research floor-scaling ("floors are not targets") + broad-topic doubled-minimums evals, vendor-neutral
source-control 0.2.0 Readiness security-gate eval + genericized fixture, mixed-actor (bot-fix-now vs human-pause) eval, three worktree evals (dry-run report-only, invalid-name rejection, batched-gh status + graceful degrade)
docs-hygiene 0.4.0 Two self-contained fixture-backed cases (compress classification-table, declutter opt-out/section-exemption) — fixtures empirically verified against detect.sh; rename-references "add an eval case" clauses
review-toolkit 0.6.0 code-review-fanout evals 6 → 20: dedup/severity-derivation, fix-pass safety fence (correctness never routed to /simplify), run-everything null-reconciliation + priority-ordering; 2 medley cases skipped (fixed-roster counts deliberately generalized away), their surviving assertions folded in

Verification

  • All 8 plugins pass claude plugin validate; scripts/validate-plugin-contracts.mjs passes (1311 files)
  • Every evals.json jq-parses and validates against the schema (ajv/check-jsonschema)
  • markdownlint-cli2 0 errors on new fixtures + CHANGELOGs; coupling grep (medley/consumer refs) clean
  • Rebased onto main post-feat: restore tier-2 salvage items across nine plugins #192; per-plugin validate re-run green after rebase

🤖 Generated with Claude Code

https://claude.ai/code/session_011xHkkNc7CR98L8Xz9Mu7ZA

Recreates the eval/test-coverage attrition from the salvage sweep (items
#13, #15-#23 plus the flagged review-toolkit and docs-hygiene items),
genericized and validated against evals.schema.json:

- songwriting 0.4.0: 13 behavioral evals mapped onto the multi-skill split
- ai-briefing 0.4.0: 6 engine evals + 3 synthetic fixtures via the
  audience-defaults seam
- event-storming 0.4.0: offline board-export eval + fixture for
  --discover-bcs (disjoint from the live-Miro eval)
- codebase-audit 0.3.0: scope-boundary routing eval
- discovery 0.5.0: research floor-scaling + broad-topic-minimums evals
- source-control 0.2.0: readiness security-gate (+fixture), mixed-actor,
  and three worktree evals
- docs-hygiene 0.4.0: self-contained compress/declutter fixtures
  (empirically verified against detect.sh), rename-references eval-case
  clauses
- review-toolkit 0.6.0: fanout evals 6 -> 20 (dedup/severity, fix-pass
  safety fence, run-everything reconciliation/ordering)

Per-plugin version bumps + CHANGELOG entries; all validations pass.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011xHkkNc7CR98L8Xz9Mu7ZA
@claude

claude Bot commented Jul 15, 2026

Copy link
Copy Markdown

Claude finished @kyle-sexton's task in 0s —— View job


I'll analyze this and get back to you.

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 617d750f2a

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread plugins/source-control/skills/pull-request/evals/evals.json Outdated
Comment thread plugins/event-storming/skills/simulation/evals/evals.json
…contradiction, offline-export branch

- source-control: grader ground-truth removed from the security fixture
  (already covered by eval 6 expectations); eval 7 no longer gates the
  required human classification reply — only reaction/fix pause
- event-storming: --discover-bcs protocol documents the supplied-export
  input path (no Miro) alongside the live board-URL read; board-write
  step marked live-path-only

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011xHkkNc7CR98L8Xz9Mu7ZA
@claude

claude Bot commented Jul 15, 2026

Copy link
Copy Markdown

Claude finished @kyle-sexton's task in 0s —— View job


I'll analyze this and get back to you.

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 8b945f4a9f

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread plugins/review-toolkit/skills/code-review-fanout/evals/evals.json Outdated
Comment thread plugins/discovery/skills/research/evals/evals.json Outdated
- review-toolkit: evals 7-20 files arrays held scenario prose, not
  fixture paths; prose folded into prompts, arrays emptied
- discovery: comparison eval targets real tools (RabbitMQ/Kafka/NATS)
  so primary-source expectations are satisfiable
- source-control: evals.json rewritten with LF + final newline
  (editorconfig hygiene failure from the previous rewrite)

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011xHkkNc7CR98L8Xz9Mu7ZA
@claude

claude Bot commented Jul 15, 2026

Copy link
Copy Markdown

Claude finished @kyle-sexton's task in 0s —— View job


I'll analyze this and get back to you.

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 0f5f761a69

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread plugins/discovery/skills/research/evals/evals.json Outdated
Comment thread plugins/source-control/skills/worktree/evals/evals.json Outdated
- discovery: production-readiness eval names FusionCache so primary-source
  expectations are satisfiable
- source-control: cleanup --dry-run now specifies git worktree prune
  --dry-run (report-only) in Step 1; eval 6 wording aligned. A dry-run
  that mutates worktree metadata violated its own contract

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011xHkkNc7CR98L8Xz9Mu7ZA
@claude

claude Bot commented Jul 15, 2026

Copy link
Copy Markdown

Claude finished @kyle-sexton's task in 0s —— View job


I'll analyze this and get back to you.

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 0a333dd9fe

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread plugins/source-control/skills/pull-request/evals/evals.json Outdated
Fixture checks are all terminal, so Gate 1 is satisfied; only Gates 2-3
block. Expectation narrowed accordingly.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011xHkkNc7CR98L8Xz9Mu7ZA
@claude

claude Bot commented Jul 15, 2026

Copy link
Copy Markdown

Claude finished @kyle-sexton's task in 1s —— View job


I'll analyze this and get back to you.

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: fdd0ed998b

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread plugins/source-control/skills/pull-request/evals/evals.json Outdated
Eval 6 required classification while asserting Gate 2 open. NOT-ready now
rests on Gates 3 and 5 only.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011xHkkNc7CR98L8Xz9Mu7ZA
@claude

claude Bot commented Jul 15, 2026

Copy link
Copy Markdown

Claude finished @kyle-sexton's task in 0s —— View job


I'll analyze this and get back to you.

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 405d82ee97

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Fixture preamble narrated the expected verdict to the skill under test.
Fixture is raw gh output only; the cue lives in expected_output now.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011xHkkNc7CR98L8Xz9Mu7ZA
@claude

claude Bot commented Jul 15, 2026

Copy link
Copy Markdown

Claude finished @kyle-sexton's task in 0s —— View job


I'll analyze this and get back to you.

@kyle-sexton
kyle-sexton merged commit dfb774d into main Jul 15, 2026
13 checks passed
@kyle-sexton
kyle-sexton deleted the restore/tier3-evals branch July 15, 2026 09:27
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant