Record not_started execution evidence when AWF fails before the engine harness - #61202
Conversation
…harness Co-authored-by: pelikhan <4175913+pelikhan@users.noreply.github.com>
…e harness Co-authored-by: pelikhan <4175913+pelikhan@users.noreply.github.com>
not_started execution evidence when AWF fails before the engine harness
|
@copilot resolve the merge conflicts on this branch. |
There was a problem hiding this comment.
🟡 Changes recommended
The new shell regression tests are not registered with the repository’s automated test targets.
Get a fresh assessment by requesting another Copilot review.
Pull request overview
Records trustworthy zero-usage evidence when AWF fails before launching an engine harness, preventing false guardrail failures such as #61184.
Changes:
- Downgrades execution evidence to
not_startedafter pre-harness failures. - Exports component metadata from compiler-generated execution steps.
- Adds tests, golden fixtures, dependency refreshes, and regenerated workflows.
File summaries
| File | Description |
|---|---|
actions/setup/sh/run_awf_with_startup_retries.sh |
Records pre-harness failure evidence. |
actions/setup/sh/run_awf_with_startup_retries_test.sh |
Tests evidence downgrade behavior. |
pkg/workflow/compiler_yaml_ai_execution.go |
Exports evidence metadata. |
.changeset/awf-startup-failure-zero-aic-evidence.md |
Documents the patch. |
pkg/workflow/testdata/TestWasmGolden_CompileFixtures/with-imports.golden |
Updates compiler fixture. |
pkg/workflow/testdata/TestWasmGolden_CompileFixtures/smoke-copilot.golden |
Updates compiler fixture. |
pkg/workflow/testdata/TestWasmGolden_CompileFixtures/playwright-cli-mode.golden |
Updates compiler fixture. |
pkg/workflow/testdata/TestWasmGolden_CompileFixtures/basic-copilot.golden |
Updates compiler fixture. |
pkg/workflow/testdata/TestWasmGolden_AllEngines/pi.golden |
Updates Pi output. |
pkg/workflow/testdata/TestWasmGolden_AllEngines/gemini.golden |
Updates Gemini output. |
pkg/workflow/testdata/TestWasmGolden_AllEngines/copilot.golden |
Updates Copilot output. |
pkg/workflow/testdata/TestWasmGolden_AllEngines/codex.golden |
Updates Codex output. |
pkg/workflow/testdata/TestWasmGolden_AllEngines/claude.golden |
Updates Claude output. |
.github/workflows/smoke-ci.lock.yml |
Refreshes generated dependencies. |
.github/workflows/notion-issue-summary.lock.yml |
Refreshes dependencies and evidence exports. |
.github/workflows/hippo-embed.lock.yml |
Refreshes dependencies and evidence exports. |
.github/workflows/firewall.lock.yml |
Refreshes dependencies and evidence exports. |
.github/workflows/example-permissions-warning.lock.yml |
Refreshes dependencies and evidence exports. |
.github/workflows/daily-team-evolution-insights.lock.yml |
Refreshes generated component evidence. |
.github/workflows/codex-github-remote-mcp-test.lock.yml |
Refreshes dependencies and evidence exports. |
Review details
- Files reviewed: 18/312 changed files
- Comments generated: 1
- Review effort level: Balanced
💡 Add a code-review agent skill for context-aware, tailored reviews. Learn more in the docs.
| @@ -70,4 +70,51 @@ set -e | |||
| [ "$STATUS" -eq 7 ] || fail "expected original post-harness failure status, got $STATUS" | |||
| [ "$(cat "$ATTEMPT_FILE")" = "1" ] || fail "expected no retry after harness marker" | |||
|
|
|||
| run_with_evidence_env() { | |||
…ai-credits-verification # Conflicts: # .github/workflows/ci-coach.lock.yml Co-authored-by: pelikhan <4175913+pelikhan@users.noreply.github.com>
Resolved the merge conflicts by merging |
|
Hey What looks good:
This is ready for review and merge. The large line count is expected given the 298 recompiled lock files. Warning Firewall blocked 1 domainThe following domain was blocked by the firewall during workflow execution:
To allow these domains, add them to the network:
allowed:
- defaults
- "github.com"See Network Configuration for more information.
|
|
🎉 This pull request is included in a new release. Release: |
The daily AI credits guardrail failed closed on run 34881527763: AWF preflight died (
cloud-hypervisor --versionerror) on both startup attempts, the Copilot harness never spawned, and therefore noapi-proxy-logs/token-usage.jsonlwas ever produced — leavingusage/agent/token_usage.jsonlmissing rather than empty, which none of the existing zero-usage proofs cover.The run provably consumed zero AI credits, but the component evidence file had already been flipped to
state: "started"immediately before the AWF invocation, soprovesExecutionNotStartedcould not vouch for it.Changes
actions/setup/sh/run_awf_with_startup_retries.sh— on final failure with no harness marker in the attempt log, rewrite the component execution evidence tostate: "not_started". Marker absence is already the signal this wrapper trusts to re-run a component without double-billing, so it is equally authoritative that no billable inference occurred. Guarded on the evidence env vars plus numericGITHUB_RUN_ID/GITHUB_RUN_ATTEMPT; atomic tmp+mvwrite, matching the compiler-generated evidence steps.pkg/workflow/compiler_yaml_ai_execution.go— thestarted-evidence injection also exports the component name and evidence path so the wrapper knows what to rewrite:Applies to
agent,detectionandevals.Tests —
run_awf_with_startup_retries_test.shcovers the downgrade (with exit status preserved) and the no-op case where the harness did start.Generated output — recompiled lock files, refreshed wasm/compile golden fixtures, changeset.
Notes for review
daily_aic_component_coverage.cjs: the existingprovesExecutionNotStartedpath already validates version/component/run id/attempt and rejects stale evidence from earlier rerun attempts, so the fail-closed posture is unchanged — only the set of runs that can prove zero usage grows.parallel_validationcould not complete — itsgit diffstep times out on the 298 recompiled.lock.ymlfiles.