Repository navigation
Fix evals accounting upload steps silently skipped after unrelated job failure - #60869
Conversation
|
@copilot mitigate this issue for legacy runs |
Co-authored-by: pelikhan <4175913+pelikhan@users.noreply.github.com>
…pload gating bug Co-authored-by: pelikhan <4175913+pelikhan@users.noreply.github.com>
Mitigated in 7238202: |
|
✅ Test Quality Sentinel completed test quality analysis. Test Quality Sentinel skipped because pre-fetch PR data was unavailable: unable to fetch test file diff
|
|
Warning Threat Detection Engine Failure — The analysis engine could not complete. This is a tooling failure, not a security finding. What happenedThe threat detection engine failed to produce results. Review the workflow run logs for details.
|
|
No ADR enforcement needed: PR does not have the 'implementation' label and has ≤100 new lines of code in business logic directories.
|
|
✅ Ponytail Reviewer completed successfully! Warning Threat Detection Engine Failure — The analysis engine could not complete. This is a tooling failure, not a security finding. What happenedThe threat detection engine failed to produce results. Review the workflow run logs for details. Warning Firewall blocked 1 domainThe following domain was blocked by the firewall during workflow execution:
To allow these domains, add them to the network:
allowed:
- defaults
- "ab.chatgpt.com"See Network Configuration for more information.
|
|
🧠 Matt Pocock Skills Reviewer has completed the skills-based review. ✅ Warning Firewall blocked 1 domainThe following domain was blocked by the firewall during workflow execution:
To allow these domains, add them to the network:
allowed:
- defaults
- "github.com"See Network Configuration for more information.
|
There was a problem hiding this comment.
🟡 Changes recommended
The legacy exemption can undercount consumed eval credits, and three upstream-managed lock files violate provenance rules.
Get a fresh assessment by requesting another Copilot review.
Pull request overview
Fixes eval accounting uploads being skipped after unrelated job failures, but also introduces an unsafe legacy-accounting exemption.
Changes:
- Adds
always()to eval summary and artifact upload conditions. - Updates accounting and compiler tests.
- Regenerates affected workflow lock files.
File summaries
| File | Description |
|---|---|
pkg/workflow/evals_steps.go |
Updates generated eval conditions. |
pkg/workflow/evals_steps_test.go |
Updates compiler expectations. |
actions/setup/js/daily_aic_component_coverage.cjs |
Adds legacy-run accounting handling. |
actions/setup/js/daily_aic_component_coverage.test.cjs |
Tests legacy-run handling. |
.github/workflows/ab-testing-advisor.lock.yml |
Regenerates eval conditions. |
.github/workflows/ace-editor.lock.yml |
Regenerates eval conditions. |
.github/workflows/agent-job-health.lock.yml |
Regenerates eval conditions. |
.github/workflows/agent-performance-analyzer.lock.yml |
Regenerates eval conditions. |
.github/workflows/agent-persona-explorer.lock.yml |
Regenerates eval conditions. |
.github/workflows/agentic-token-audit.lock.yml |
Regenerates eval conditions. |
.github/workflows/agentic-token-optimizer.lock.yml |
Regenerates eval conditions. |
.github/workflows/agentic-token-trend-audit.lock.yml |
Regenerates eval conditions. |
.github/workflows/ai-moderator.lock.yml |
Regenerates eval conditions. |
.github/workflows/api-consumption-report.lock.yml |
Regenerates eval conditions. |
.github/workflows/approach-validator.lock.yml |
Regenerates eval conditions. |
.github/workflows/archie.lock.yml |
Regenerates eval conditions. |
.github/workflows/architecture-guardian.lock.yml |
Regenerates eval conditions. |
.github/workflows/archivx-agentic-workflows-analyzer.lock.yml |
Regenerates eval conditions. |
.github/workflows/artifacts-summary.lock.yml |
Regenerates eval conditions. |
.github/workflows/audit-workflows.lock.yml |
Regenerates eval conditions. |
.github/workflows/auto-triage-issues.lock.yml |
Regenerates eval conditions. |
.github/workflows/avenger.lock.yml |
Regenerates eval conditions. |
.github/workflows/aw-failure-investigator.lock.yml |
Regenerates eval conditions. |
.github/workflows/blog-auditor.lock.yml |
Regenerates eval conditions. |
.github/workflows/bot-detection.lock.yml |
Regenerates eval conditions. |
.github/workflows/breaking-change-checker.lock.yml |
Regenerates eval conditions. |
.github/workflows/changeset.lock.yml |
Regenerates eval conditions. |
.github/workflows/ci-coach.lock.yml |
Regenerates eval conditions. |
.github/workflows/ci-doctor.lock.yml |
Regenerates eval conditions. |
.github/workflows/claude-code-user-docs-review.lock.yml |
Regenerates eval conditions. |
.github/workflows/cli-consistency-checker.lock.yml |
Regenerates eval conditions. |
.github/workflows/cli-version-checker.lock.yml |
Regenerates eval conditions. |
.github/workflows/cloclo.lock.yml |
Regenerates eval conditions. |
.github/workflows/code-scanning-fixer.lock.yml |
Regenerates eval conditions. |
.github/workflows/code-simplifier.lock.yml |
Regenerates eval conditions. |
.github/workflows/commit-changes-analyzer.lock.yml |
Regenerates eval conditions. |
.github/workflows/constraint-solving-potd.lock.yml |
Regenerates eval conditions. |
.github/workflows/contribution-check.lock.yml |
Regenerates eval conditions. |
.github/workflows/copilot-agent-analysis.lock.yml |
Regenerates eval conditions. |
.github/workflows/copilot-centralization-drilldown.lock.yml |
Regenerates eval conditions. |
.github/workflows/copilot-centralization-optimizer.lock.yml |
Regenerates eval conditions. |
.github/workflows/copilot-cli-deep-research.lock.yml |
Regenerates eval conditions. |
.github/workflows/copilot-opt.lock.yml |
Regenerates eval conditions. |
.github/workflows/copilot-pr-merged-report.lock.yml |
Regenerates eval conditions. |
.github/workflows/copilot-pr-nlp-analysis.lock.yml |
Regenerates eval conditions. |
.github/workflows/copilot-pr-prompt-analysis.lock.yml |
Regenerates eval conditions. |
.github/workflows/copilot-session-insights.lock.yml |
Regenerates eval conditions. |
.github/workflows/craft.lock.yml |
Regenerates eval conditions. |
.github/workflows/daily-action-setup-security-audit.lock.yml |
Regenerates eval conditions. |
.github/workflows/daily-agent-of-the-day-blog-writer.lock.yml |
Regenerates eval conditions. |
.github/workflows/daily-agentrx-trace-optimizer.lock.yml |
Regenerates eval conditions. |
.github/workflows/daily-ambient-context-optimizer.lock.yml |
Regenerates eval conditions. |
.github/workflows/daily-architecture-diagram.lock.yml |
Regenerates eval conditions. |
.github/workflows/daily-assign-issue-to-user.lock.yml |
Regenerates eval conditions. |
.github/workflows/daily-astrostylelite-markdown-spellcheck.lock.yml |
Regenerates eval conditions. |
.github/workflows/daily-aw-cross-repo-compile-check.lock.yml |
Regenerates eval conditions. |
.github/workflows/daily-awf-spec-compiler-surfacing.lock.yml |
Regenerates eval conditions. |
.github/workflows/daily-cache-strategy-analyzer.lock.yml |
Regenerates eval conditions. |
.github/workflows/daily-caveman-optimizer.lock.yml |
Regenerates eval conditions. |
.github/workflows/daily-cli-performance.lock.yml |
Regenerates eval conditions. |
.github/workflows/daily-cli-tools-tester.lock.yml |
Regenerates eval conditions. |
.github/workflows/daily-code-metrics.lock.yml |
Regenerates eval conditions. |
.github/workflows/daily-community-attribution.lock.yml |
Regenerates eval conditions. |
.github/workflows/daily-compiler-quality.lock.yml |
Regenerates eval conditions. |
.github/workflows/daily-compiler-threat-spec-optimizer.lock.yml |
Regenerates eval conditions. |
.github/workflows/daily-doc-healer.lock.yml |
Regenerates eval conditions. |
.github/workflows/daily-doc-updater.lock.yml |
Regenerates eval conditions. |
.github/workflows/daily-documentation-diagram.lock.yml |
Regenerates eval conditions. |
.github/workflows/daily-elixir-credo-snippet-audit.lock.yml |
Regenerates eval conditions. |
.github/workflows/daily-evals-report.lock.yml |
Regenerates eval conditions. |
.github/workflows/daily-experiment-report.lock.yml |
Regenerates eval conditions. |
.github/workflows/daily-fact.lock.yml |
Regenerates eval conditions. |
.github/workflows/daily-file-diet.lock.yml |
Regenerates eval conditions. |
.github/workflows/daily-firewall-report.lock.yml |
Regenerates eval conditions. |
.github/workflows/daily-function-namer.lock.yml |
Regenerates eval conditions. |
.github/workflows/daily-geo-optimizer.lock.yml |
Regenerates eval conditions. |
.github/workflows/daily-github-docs-seo-optimizer.lock.yml |
Regenerates eval conditions. |
.github/workflows/daily-go-test-parallelizer.lock.yml |
Regenerates eval conditions. |
.github/workflows/daily-grader-audit.lock.yml |
Regenerates eval conditions. |
.github/workflows/daily-graft-intelligence.lock.yml |
Regenerates eval conditions. |
.github/workflows/daily-harness-experiment-proposer.lock.yml |
Regenerates eval conditions. |
.github/workflows/daily-hippo-learn.lock.yml |
Regenerates eval conditions. |
.github/workflows/daily-issues-report.lock.yml |
Regenerates eval conditions. |
.github/workflows/daily-malicious-code-scan.lock.yml |
Regenerates eval conditions. |
.github/workflows/daily-mcp-concurrency-analysis.lock.yml |
Regenerates eval conditions. |
.github/workflows/daily-model-inventory.lock.yml |
Regenerates eval conditions. |
.github/workflows/daily-multi-device-docs-tester.lock.yml |
Regenerates eval conditions. |
.github/workflows/daily-news.lock.yml |
Regenerates eval conditions. |
.github/workflows/daily-observability-report.lock.yml |
Regenerates eval conditions. |
.github/workflows/daily-performance-summary.lock.yml |
Regenerates eval conditions. |
.github/workflows/daily-regulatory.lock.yml |
Regenerates eval conditions. |
.github/workflows/daily-reliability-review.lock.yml |
Regenerates eval conditions. |
.github/workflows/daily-rendering-scripts-verifier.lock.yml |
Regenerates eval conditions. |
.github/workflows/daily-repo-chronicle.lock.yml |
Regenerates eval conditions. |
.github/workflows/daily-safe-output-integrator.lock.yml |
Regenerates eval conditions. |
.github/workflows/daily-safe-output-optimizer.lock.yml |
Regenerates eval conditions. |
.github/workflows/daily-safe-outputs-conformance.lock.yml |
Regenerates eval conditions. |
.github/workflows/daily-safeoutputs-git-simulator.lock.yml |
Regenerates eval conditions. |
.github/workflows/daily-secrets-analysis.lock.yml |
Regenerates eval conditions. |
.github/workflows/daily-security-observability.lock.yml |
Regenerates eval conditions. |
.github/workflows/daily-security-red-team.lock.yml |
Regenerates eval conditions. |
.github/workflows/daily-semgrep-scan.lock.yml |
Regenerates eval conditions. |
.github/workflows/daily-spdd-spec-planner.lock.yml |
Regenerates eval conditions. |
.github/workflows/daily-spending-forecast.lock.yml |
Regenerates eval conditions. |
.github/workflows/daily-squid-image-scan.lock.yml |
Regenerates eval conditions. |
.github/workflows/daily-storify.lock.yml |
Regenerates eval conditions. |
.github/workflows/daily-syntax-error-quality.lock.yml |
Regenerates eval conditions. |
.github/workflows/daily-token-consumption-report.lock.yml |
Regenerates eval conditions. |
.github/workflows/daily-vulnhunter-scan.lock.yml |
Regenerates eval conditions. |
.github/workflows/daily-windows-terminal-integration-builder.lock.yml |
Regenerates eval conditions. |
.github/workflows/daily-workflow-updater.lock.yml |
Regenerates eval conditions. |
.github/workflows/daily-yamllint-fixer.lock.yml |
Regenerates eval conditions. |
.github/workflows/dataflow-pr-discussion-dataset.lock.yml |
Regenerates eval conditions. |
.github/workflows/deep-report.lock.yml |
Regenerates eval conditions. |
.github/workflows/deepsec-security-scan.lock.yml |
Regenerates eval conditions. |
.github/workflows/delight.lock.yml |
Regenerates eval conditions. |
.github/workflows/dependabot-burner.lock.yml |
Regenerates eval conditions. |
.github/workflows/dependabot-go-checker.lock.yml |
Regenerates eval conditions. |
.github/workflows/deployment-incident-monitor.lock.yml |
Regenerates eval conditions. |
.github/workflows/design-decision-gate.lock.yml |
Regenerates eval conditions. |
.github/workflows/detection-analysis-report.lock.yml |
Regenerates eval conditions. |
.github/workflows/dev-hawk.lock.yml |
Regenerates eval conditions. |
.github/workflows/dev.lock.yml |
Regenerates eval conditions. |
.github/workflows/developer-docs-consolidator.lock.yml |
Regenerates eval conditions. |
.github/workflows/dictation-prompt.lock.yml |
Regenerates eval conditions. |
.github/workflows/docs-noob-tester.lock.yml |
Regenerates eval conditions. |
.github/workflows/draft-pr-cleanup.lock.yml |
Regenerates eval conditions. |
.github/workflows/duplicate-code-detector.lock.yml |
Regenerates eval conditions. |
.github/workflows/eslint-miner.lock.yml |
Regenerates eval conditions. |
.github/workflows/eslint-monster.lock.yml |
Regenerates eval conditions. |
.github/workflows/eslint-refiner.lock.yml |
Regenerates eval conditions. |
.github/workflows/evoskill-evolver.lock.yml |
Regenerates eval conditions. |
.github/workflows/functional-pragmatist.lock.yml |
Regenerates eval conditions. |
.github/workflows/glossary-maintainer.lock.yml |
Regenerates eval conditions. |
.github/workflows/go-logger.lock.yml |
Regenerates eval conditions. |
.github/workflows/gpclean.lock.yml |
Regenerates eval conditions. |
.github/workflows/hourly-ci-cleaner.lock.yml |
Regenerates eval conditions. |
.github/workflows/issue-arborist.lock.yml |
Regenerates eval conditions. |
.github/workflows/issue-monster.lock.yml |
Regenerates eval conditions. |
.github/workflows/issue-triage-agent.lock.yml |
Regenerates eval conditions. |
.github/workflows/necromancer.lock.yml |
Regenerates eval conditions. |
.github/workflows/plan.lock.yml |
Regenerates eval conditions. |
.github/workflows/poem-bot.lock.yml |
Regenerates eval conditions. |
.github/workflows/pr-code-quality-reviewer.lock.yml |
Regenerates eval conditions. |
.github/workflows/pr-sous-chef.lock.yml |
Regenerates eval conditions. |
.github/workflows/pr-triage-agent.lock.yml |
Regenerates eval conditions. |
.github/workflows/purelock.lock.yml |
Regenerates eval conditions. |
.github/workflows/refiner.lock.yml |
Regenerates eval conditions. |
.github/workflows/release.lock.yml |
Regenerates eval conditions. |
.github/workflows/repo-audit-analyzer.lock.yml |
Regenerates eval conditions. |
.github/workflows/research.lock.yml |
Regenerates eval conditions. |
.github/workflows/security-review.lock.yml |
Regenerates eval conditions. |
.github/workflows/smoke-copilot-aoai-apikey.lock.yml |
Regenerates eval conditions. |
.github/workflows/smoke-copilot-aoai-entra.lock.yml |
Regenerates eval conditions. |
.github/workflows/smoke-copilot-sub-agents.lock.yml |
Regenerates eval conditions. |
.github/workflows/smoke-copilot.lock.yml |
Regenerates eval conditions. |
.github/workflows/smoke-gemini.lock.yml |
Regenerates eval conditions. |
.github/workflows/smoke-project.lock.yml |
Regenerates eval conditions. |
.github/workflows/smoke-temporary-id.lock.yml |
Regenerates eval conditions. |
.github/workflows/spec-enforcer.lock.yml |
Regenerates eval conditions. |
.github/workflows/stale-pr-cleanup.lock.yml |
Regenerates eval conditions. |
.github/workflows/stale-repo-identifier.lock.yml |
Regenerates eval conditions. |
.github/workflows/sub-issue-closer.lock.yml |
Regenerates eval conditions. |
.github/workflows/technical-doc-writer.lock.yml |
Regenerates eval conditions. |
.github/workflows/test-quality-sentinel.lock.yml |
Regenerates eval conditions. |
.github/workflows/tidy.lock.yml |
Regenerates eval conditions. |
.github/workflows/typist.lock.yml |
Regenerates eval conditions. |
.github/workflows/unbloat-docs.lock.yml |
Regenerates eval conditions. |
.github/workflows/weekly-blog-post-writer.lock.yml |
Regenerates eval conditions. |
Review details
- Files reviewed: 169/169 changed files
- Comments generated: 5
- Review effort level: Balanced
💡 Add a code-review agent skill for context-aware, tailored reviews. Learn more in the docs.
| if (succeeded("Collect evals token usage") && succeeded("Redact secrets in evals results") && conclusionOf("Upload evals results") === "skipped" && conclusionOf("Upload evals accounting after failure") === "skipped") { | ||
| return true; | ||
| } |
| if !strings.Contains(allSteps, "if: always() && steps.redact_evals_results.outcome == 'success'") { | ||
| t.Errorf("expected redact outcome gating for render/upload steps;\ngot:\n%s", allSteps) | ||
| } |
| await main(); | ||
| - name: Render evals results to step summary | ||
| if: steps.redact_evals_results.outcome == 'success' | ||
| if: always() && steps.redact_evals_results.outcome == 'success' |
| await main(); | ||
| - name: Render evals results to step summary | ||
| if: steps.redact_evals_results.outcome == 'success' | ||
| if: always() && steps.redact_evals_results.outcome == 'success' |
| await main(); | ||
| - name: Render evals results to step summary | ||
| if: steps.redact_evals_results.outcome == 'success' | ||
| if: always() && steps.redact_evals_results.outcome == 'success' |
There was a problem hiding this comment.
Reviewed the substantive changes (Go compiler fix, JS guardrail heuristic, and their tests); the rest of the diff is generated .lock.yml recompilation.
Findings:
- Fix is correct: adding
always() &&to the two evals upload step conditions matches the sibling failure-path step and resolves the false-negative guardrail failure described in the PR. - The new legacy-detection branch in
provesFailedEvalsHadNoUsage(daily_aic_component_coverage.cjs) is well-reasoned and covered by a new test (counts missing evals accounting as zero for a legacy run...), which passes along with the rest of the 52-test suite. - Go tests (
TestDailyAICEvalsAccountingTransport) pass with updated expected condition strings. - One minor doc-drift nit left as an inline comment: the step-name sync comment above
buildUploadEvalsArtifactStepdoesn't mention the new"Redact secrets in evals results"dependency introduced by this change.
No blocking issues found.
Warning
Firewall blocked 2 domains
The following domains were blocked by the firewall during workflow execution:
github.com/ghapigithub.laiyagushi.com
[!TIP]
github.com/ghapi is blocked because GitHub API access uses the built-in GitHub tools by default. Instead of adding github.com/ghapi to network.allowed, use tools.github.mode: gh-proxy for direct pre-authenticated GitHub CLI access without requiring network access to github.com/ghapi:
tools:
github:
mode: gh-proxySee GitHub Tools for more information on gh-proxy mode.
To allow these domains, add them to the network.allowed list in your workflow frontmatter:
network:
allowed:
- defaults
- "github.com/ghapi"
- "github.com"See Network Configuration for more information.
🧵 Reviewed using Impeccable skills by Impeccable Skills Reviewer · copilot · sonnet50 · 79.7 AIC · ⌖ 13.6 AIC · ⊞ 8.4K
Comments that could not be inline-anchored
pkg/workflow/evals_steps.go:438
The sync comment above (buildUploadEvalsArtifactStep, lines 434-438) says the JS-matched step names are limited to Collect evals token usage, Upload evals results, and Upload evals accounting after failure — but the new legacy-detection branch in provesFailedEvalsHadNoUsage also matches "Redact secrets in evals results" (generated by buildRedactEvalsSecretsStep, a different function). Renaming that step would now silently break the guardrail too, and the comment doesn't mention it…
There was a problem hiding this comment.
Skills-Based Review 🧠
Applied /diagnosing-bugs and /codebase-design. The core fix (adding always() && to the two gated evals upload steps in pkg/workflow/evals_steps.go) correctly addresses the root cause and is covered by updated unit tests in evals_steps_test.go.
📋 Key Themes & Highlights
Key Themes
- Legacy-run mitigation in
actions/setup/js/daily_aic_component_coverage.cjs(provesFailedEvalsHadNoUsage) is a reasonable stopgap for pre-fix runs still in the accounting window, backed by a new regression test, but it reuses the"failed_before_accounting"reason string for a fundamentally different case (compatibility shim vs. genuine zero-usage proof), which makes the two indistinguishable in guardrail logs. - The workaround also has no explicit expiry mechanism or tracking-issue reference, so it risks becoming permanent, silent debt — and could theoretically mask a future regression that happens to produce the same skip pattern for a different reason.
Positive Highlights
- ✅ Root cause correctly diagnosed and fixed at the source (compiler-generated
if:conditions), not just patched around in the guardrail script. - ✅ Good regression test added (
daily_aic_component_coverage.test.cjs) reproducing the exact legacy skip pattern with clear comments explaining the GitHub Actionssuccess()implicit-AND behavior. - ✅
.lock.ymlregeneration is consistent across all 169 affected workflows.
Note: only 2 non-lock source files (pkg/workflow/evals_steps.go, pkg/workflow/evals_steps_test.go) and the mitigation JS files were reachable for direct diff inspection in this sandbox — the pre-fetched diff patch was truncated to lock-file changes only, so source-level review used the GitHub API directly.
Warning
Firewall blocked 1 domain
The following domain was blocked by the firewall during workflow execution:
github.com
To allow these domains, add them to the network.allowed list in your workflow frontmatter:
network:
allowed:
- defaults
- "github.com"See Network Configuration for more information.
🧠 Reviewed using Matt Pocock's skills by Matt Pocock Skills Reviewer · copilot · sonnet50 · 93.5 AIC · ⌖ 14 AIC · ⊞ 10.4K
Comment /matt to run again
| // uploads), so observing it here proves this is a legacy pre-fix run. | ||
| // Once such runs age out of the rolling accounting window this branch | ||
| // becomes dead code, but it unblocks the guardrail for them in the meantime. | ||
| if (succeeded("Collect evals token usage") && succeeded("Redact secrets in evals results") && conclusionOf("Upload evals results") === "skipped" && conclusionOf("Upload evals accounting after failure") === "skipped") { |
There was a problem hiding this comment.
[/diagnosing-bugs] The legacy-run workaround reuses the same "failed_before_accounting" reason string as the genuine zero-usage path above it, so guardrail logs can no longer distinguish a real zero-usage failure from this compatibility shim for pre-fix runs.
💡 Suggested fix
Give this branch its own reason (e.g. "legacy_pre_fix_skipped_uploads") so [daily-workflow-aic] log lines and any downstream alerting can tell the two cases apart. This also makes it trivial to grep production logs later to see how often the workaround still fires and decide when it's safe to delete.
@copilot please address this.
| // the fix (redaction succeeding always triggers exactly one of the two | ||
| // uploads), so observing it here proves this is a legacy pre-fix run. | ||
| // Once such runs age out of the rolling accounting window this branch | ||
| // becomes dead code, but it unblocks the guardrail for them in the meantime. |
There was a problem hiding this comment.
[/codebase-design] This compatibility branch has no enforced expiry: the comment says it "becomes dead code" once legacy runs age out, but nothing actually checks run age/date, and there's no tracking issue reference to prompt removal.
💡 Suggested improvement
Consider either (a) bounding the condition with an explicit date/run-id check (e.g. only apply when job.started_at predates the fix's rollout), or (b) adding a // TODO(#issue): remove once pre-fix runs age out of the accounting window comment linking to a tracking issue, so this doesn't silently become permanent debt that also risks masking a future regression that happens to produce the same skip pattern for an unrelated reason.
@copilot please address this.
|
@copilot run pr-finished skill |
Co-authored-by: pelikhan <4175913+pelikhan@users.noreply.github.com>
Ran the PR finisher pass. Commit |
The daily AI-credits guardrail failed closed on the PR Sous Chef workflow, reporting missing accounting for the
evalscomponent becauseevals/token_usage.jsonlwas absent from a failed run.Root cause
In the evals job, the
Upload evals resultsandRender evals results to step summarysteps gated onif: steps.redact_evals_results.outcome == 'success'with no explicit status-check function. GitHub Actions implicitly ANDs such conditions withsuccess(), so when an earlier, unrelated step in the same job failed (e.g.Install AWF binary), both steps were skipped even though redaction itself succeeded — leaving no accounting artifact for the guardrail to verify.Changes
pkg/workflow/evals_steps.go: addedalways() &&to both step conditions so they run whenever redaction succeeded, regardless of earlier unrelated step failures — consistent with the siblingUpload evals accounting after failurestep, which already had this.pkg/workflow/evals_steps_test.go: updated expected condition strings to match..lock.ymlfiles.