| PR Sous Chef |
6 |
100.0% |
0.0% |
"Does the agent output show a specific reason why the selected PR needs a nudge toward maintainer investigation?" (0.0% YES) |
| Issue Monster |
3 |
100.0% |
0.0% |
"Does the agent output show that at most one issue was assigned to Copilot per run?" (0.0% YES) |
| Avenger |
2 |
100.0% |
0.0% |
"Did the agent assess the current CI state and determine if intervention was needed? [2 UNKNOWN]" (0.0% YES) |
| [aw] Failure Investigator (6h) |
1 |
100.0% |
100.0% |
"Did the agent investigate agentic workflow failures from the last 6 hours and produce findings?" (100.0% YES) |
| AI Moderator |
1 |
100.0% |
0.0% |
"Did the agent apply at least one label (spam, ai-generated, link-spam, or ai-inspected) or call noop? [1 UNKNOWN]" (0.0% YES) |
| Auto-Triage Issues |
1 |
100.0% |
0.0% |
"Did the agent label a non-report issue describing unexpected behavior, expected versus actual results, or a reproducible failure with bug and the relevant component label (or needs-triage if no component is clear)?" (0.0% YES) |
| CLI Version Checker |
1 |
100.0% |
100.0% |
"Did the agent check for new versions of agentic CLI tools (Claude Code, GitHub Copilot CLI, Codex, MCP servers, etc.)?" (100.0% YES) |
| Code Scanning Fixer |
1 |
100.0% |
0.0% |
"Did the agent analyze code scanning alerts and identify at least one fixable alert, or correctly skip when no fixable alerts were found?" (0.0% YES) |
| Contribution Check |
1 |
100.0% |
0.0% |
"Did the agent dispatch PRs to the contribution-checker subagent and produce evaluation results for each PR? [1 UNKNOWN]" (0.0% YES) |
| Copilot Session Insights |
1 |
100.0% |
100.0% |
"Was a report produced with usage patterns, success rates, and performance metrics?" (100.0% YES) |
| Daily AgentRx Trace Optimizer |
1 |
100.0% |
0.0% |
"Does the agent output show that the objective for experiment sub_agent_strategy was successfully completed? [1 UNKNOWN]" (0.0% YES) |
| Daily Assign Issue To User |
1 |
100.0% |
0.0% |
"Did the agent assign one or more unassigned issues to a user, or correctly call noop when no unassigned issues were found?" (0.0% YES) |
| Daily Cli Tools Tester |
1 |
100.0% |
0.0% |
"Did the agent run exploratory tests on the audit, logs, and compile CLI tools?" (0.0% YES) |
| Daily Container Image Security Scan |
1 |
100.0% |
100.0% |
"Did the agent analyze container images for vulnerabilities, updates, and rejected licenses?" (100.0% YES) |
| Daily GitHub Docs SEO Optimizer |
1 |
100.0% |
0.0% |
"Did the agent report an actionable documentation recommendation or explain why no update was needed? [1 UNKNOWN]" (0.0% YES) |
| Daily Go Test Parallelizer |
1 |
100.0% |
100.0% |
"Did the agent create a pull request for a safe test change, or use noop when no safe change was available?" (100.0% YES) |
| Daily Safe Outputs Conformance Checker |
1 |
100.0% |
100.0% |
"Did the agent run a conformance check against the Safe Outputs specification implementation?" (100.0% YES) |
| Daily VulnHunter Scan |
1 |
100.0% |
100.0% |
"Was a security issue created for verified exploitable findings, or was noop used when VulnHunter found nothing actionable?" (100.0% YES) |
| Daily Workflow Updater |
1 |
100.0% |
0.0% |
"Did the agent check GitHub Actions versions for available updates?" (0.0% YES) |
| Design Decision Gate 🏗️ |
1 |
100.0% |
0.0% |
"Does the agent output confirm that it checked for existing ADRs before deciding on an action?" (0.0% YES) |
| ESLint Refiner |
1 |
100.0% |
100.0% |
"Did the agent analyze ESLint diagnostics trends to identify rule refinement opportunities?" (100.0% YES) |
| Issue Arborist |
1 |
100.0% |
0.0% |
"Did the agent analyze recent issues and identify related issue relationships? [1 UNKNOWN]" (0.0% YES) |
| Multi-Device Docs Tester |
1 |
100.0% |
100.0% |
"Did the agent test the documentation site across the requested device form factors?" (100.0% YES) |
| PR Code Quality Reviewer |
1 |
100.0% |
0.0% |
"Does the agent output show that the review findings are limited to changes in the pull request diff rather than unrelated code? [1 UNKNOWN]" (0.0% YES) |
| PR Triage Agent |
1 |
100.0% |
0.0% |
"Does the agent output include a triage report summarizing the PRs processed? [1 UNKNOWN]" (0.0% YES) |
| Sub-Issue Closer |
1 |
100.0% |
100.0% |
"Did the agent check parent issues for the completion status of all their sub-issues?" (100.0% YES) |
| Test Quality Sentinel |
1 |
100.0% |
0.0% |
"Does the agent output show that the objective for experiment model_size was successfully completed? [1 UNKNOWN]" (0.0% YES) |
| Tidy |
1 |
100.0% |
100.0% |
"Was a pull request created with formatting and tidying fixes, or was noop used when no changes were needed?" (100.0% YES) |
Executive Summary
Analyzed 36 evals-producing runs across 28 workflows over the last 7 days. The evals pipeline itself is stable with a 100.0% job success rate, but the aggregate YES rate is only 50.0%, so the feature is DEGRADED. 23 answers were emitted as
UNKNOWNrather than binary YES/NO and were counted as non-passing for this report.Note
Status: DEGRADED — evals artifacts are being produced consistently, but judge outcomes remain below the 60% HEALTHY threshold and include 23
UNKNOWNanswers.Key Metrics
Per-Workflow Pass Rates
bugand the relevant component label (orneeds-triageif no component is clear)?" (0.0% YES)Per-Question Breakdown per Workflow
PR Sous Chef
Issue Monster
Avenger
[aw] Failure Investigator (6h)
AI Moderator
Auto-Triage Issues
bugand the relevant component label (orneeds-triageif no component is clear)?CLI Version Checker
Code Scanning Fixer
Contribution Check
Copilot Session Insights
Daily AgentRx Trace Optimizer
Daily Assign Issue To User
Daily Cli Tools Tester
Daily Container Image Security Scan
Daily GitHub Docs SEO Optimizer
Daily Go Test Parallelizer
Daily Safe Outputs Conformance Checker
Daily VulnHunter Scan
Daily Workflow Updater
Design Decision Gate 🏗️
ESLint Refiner
Issue Arborist
Multi-Device Docs Tester
PR Code Quality Reviewer
PR Triage Agent
Sub-Issue Closer
Test Quality Sentinel
Tidy
Quality Signals
nudge-targetedis the weakest current question at 0.0% YES (0 YES / 6 NO) in PR Sous Chef.single_issue_scopedis the weakest current question at 0.0% YES (0 YES / 3 NO) in Issue Monster.ci_state_assessedis the weakest current question at 0.0% YES (0 YES / 2 NO; 2 UNKNOWN) in Avenger.Recommendations
PR Sous Chef,Issue Monster, andAvenger; those workflows account for the highest-volume zero-YES questions in the 7-day window.UNKNOWNin otherwise successful evals jobs. That usually points to missing agent context, ambiguous question wording, or prompts that do not surface enough observable evidence for the judge.push_evals_statereliability.References
§31469629212 §31468615174 §31468388964