Skip to content

Record not_started execution evidence when AWF fails before the engine harness - #61202

Merged
pelikhan merged 4 commits into
mainfrom
copilot/aw-fix-daily-ai-credits-verification
Sep 15, 2026
Merged

pelikhan merged 4 commits into
mainfrom
copilot/aw-fix-daily-ai-credits-verification

Conversation

Copilot AI commented Sep 15, 2026 •

Copy link
Copy Markdown
Contributor

The daily AI credits guardrail failed closed on run 34881527763: AWF preflight died (cloud-hypervisor --version error) on both startup attempts, the Copilot harness never spawned, and therefore no api-proxy-logs/token-usage.jsonl was ever produced — leaving usage/agent/token_usage.jsonl missing rather than empty, which none of the existing zero-usage proofs cover.

[ERROR] Fatal error: Error: ".../cloud-hypervisor --version" exited with code undefined
[copilot-awf-retry] AWF startup failed before copilot harness; retrying fresh (startup retry 1/1)
[ERROR] Fatal error: ...   # same failure, harness marker never emitted

The run provably consumed zero AI credits, but the component evidence file had already been flipped to state: "started" immediately before the AWF invocation, so provesExecutionNotStarted could not vouch for it.

Changes

  • actions/setup/sh/run_awf_with_startup_retries.sh — on final failure with no harness marker in the attempt log, rewrite the component execution evidence to state: "not_started". Marker absence is already the signal this wrapper trusts to re-run a component without double-billing, so it is equally authoritative that no billable inference occurred. Guarded on the evidence env vars plus numeric GITHUB_RUN_ID/GITHUB_RUN_ATTEMPT; atomic tmp+mv write, matching the compiler-generated evidence steps.

  • pkg/workflow/compiler_yaml_ai_execution.go — the started-evidence injection also exports the component name and evidence path so the wrapper knows what to rewrite:

    mv "$evidence_tmp" "/tmp/gh-aw/agent_execution.json"
    export GH_AW_AWF_EXECUTION_COMPONENT="agent"
    export GH_AW_AWF_EXECUTION_EVIDENCE_FILE="/tmp/gh-aw/agent_execution.json"
    GH_AW_AWF_ENGINE_NAME=copilot \
    ...

    Applies to agent, detection and evals.

  • Tests — run_awf_with_startup_retries_test.sh covers the downgrade (with exit status preserved) and the no-op case where the harness did start.

  • Generated output — recompiled lock files, refreshed wasm/compile golden fixtures, changeset.

Notes for review

  • No change to daily_aic_component_coverage.cjs: the existing provesExecutionNotStarted path already validates version/component/run id/attempt and rejects stale evidence from earlier rerun attempts, so the fail-closed posture is unchanged — only the set of runs that can prove zero usage grows.
  • Engines that don't use the retry wrapper still fail closed; they have no harness marker to reason about.
  • parallel_validation could not complete — its git diff step times out on the 298 recompiled .lock.yml files.

Copilot AI and others added 2 commits September 15, 2026 19:53
…harness

Co-authored-by: pelikhan <4175913+pelikhan@users.noreply.github.com>
…e harness

Co-authored-by: pelikhan <4175913+pelikhan@users.noreply.github.com>
Copilot AI changed the title [WIP] Fix verification of daily AI credits in tracking agent Record not_started execution evidence when AWF fails before the engine harness Sep 15, 2026
Copilot AI requested a review from pelikhan September 15, 2026 20:05
@pelikhan
pelikhan marked this pull request as ready for review September 15, 2026 20:05
Copilot AI balanced review requested due to automatic review settings September 15, 2026 20:05
@pelikhan

Copy link
Copy Markdown
Collaborator

@copilot resolve the merge conflicts on this branch.

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 Changes recommended

The new shell regression tests are not registered with the repository’s automated test targets.

Get a fresh assessment by requesting another Copilot review.

Pull request overview

Records trustworthy zero-usage evidence when AWF fails before launching an engine harness, preventing false guardrail failures such as #61184.

Changes:

  • Downgrades execution evidence to not_started after pre-harness failures.
  • Exports component metadata from compiler-generated execution steps.
  • Adds tests, golden fixtures, dependency refreshes, and regenerated workflows.
File summaries
File Description
actions/setup/sh/run_awf_with_startup_retries.sh Records pre-harness failure evidence.
actions/setup/sh/run_awf_with_startup_retries_test.sh Tests evidence downgrade behavior.
pkg/workflow/compiler_yaml_ai_execution.go Exports evidence metadata.
.changeset/awf-startup-failure-zero-aic-evidence.md Documents the patch.
pkg/workflow/testdata/TestWasmGolden_CompileFixtures/with-imports.golden Updates compiler fixture.
pkg/workflow/testdata/TestWasmGolden_CompileFixtures/smoke-copilot.golden Updates compiler fixture.
pkg/workflow/testdata/TestWasmGolden_CompileFixtures/playwright-cli-mode.golden Updates compiler fixture.
pkg/workflow/testdata/TestWasmGolden_CompileFixtures/basic-copilot.golden Updates compiler fixture.
pkg/workflow/testdata/TestWasmGolden_AllEngines/pi.golden Updates Pi output.
pkg/workflow/testdata/TestWasmGolden_AllEngines/gemini.golden Updates Gemini output.
pkg/workflow/testdata/TestWasmGolden_AllEngines/copilot.golden Updates Copilot output.
pkg/workflow/testdata/TestWasmGolden_AllEngines/codex.golden Updates Codex output.
pkg/workflow/testdata/TestWasmGolden_AllEngines/claude.golden Updates Claude output.
.github/workflows/smoke-ci.lock.yml Refreshes generated dependencies.
.github/workflows/notion-issue-summary.lock.yml Refreshes dependencies and evidence exports.
.github/workflows/hippo-embed.lock.yml Refreshes dependencies and evidence exports.
.github/workflows/firewall.lock.yml Refreshes dependencies and evidence exports.
.github/workflows/example-permissions-warning.lock.yml Refreshes dependencies and evidence exports.
.github/workflows/daily-team-evolution-insights.lock.yml Refreshes generated component evidence.
.github/workflows/codex-github-remote-mcp-test.lock.yml Refreshes dependencies and evidence exports.
Review details
  • Files reviewed: 18/312 changed files
  • Comments generated: 1
  • Review effort level: Balanced

💡 Add a code-review agent skill for context-aware, tailored reviews. Learn more in the docs.

@@ -70,4 +70,51 @@ set -e
[ "$STATUS" -eq 7 ] || fail "expected original post-harness failure status, got $STATUS"
[ "$(cat "$ATTEMPT_FILE")" = "1" ] || fail "expected no retry after harness marker"

run_with_evidence_env() {
…ai-credits-verification

# Conflicts:
#	.github/workflows/ci-coach.lock.yml

Co-authored-by: pelikhan <4175913+pelikhan@users.noreply.github.com>

Copilot AI commented Sep 15, 2026

Copy link
Copy Markdown
Contributor Author

@copilot resolve the merge conflicts on this branch.

Resolved the merge conflicts by merging origin/main and regenerating the conflicted ci-coach lock file from the merged workflow source. Addressed in f2d1df4.

@github-actions

Copy link
Copy Markdown
Contributor

Hey @copilot-swe-agent 👋 — great work on fixing the daily AI credits guardrail! This PR is well-documented and focused on solving the core issue where AWF preflight failures prevent proper evidence collection.

What looks good:

  • Clear problem statement with error logs and root cause
  • Focused fix: shell script wrapper upgrade + compiler changes
  • Tests included for downgrade and no-op cases
  • Excellent notes for reviewers
  • Well-reasoned fail-closed posture maintained

This is ready for review and merge. The large line count is expected given the 298 recompiled lock files.

Warning

Firewall blocked 1 domain

The following domain was blocked by the firewall during workflow execution:

  • github.com

To allow these domains, add them to the network.allowed list in your workflow frontmatter:

network:
  allowed:
    - defaults
    - "github.com"

See Network Configuration for more information.

Generated by ✅ Contribution Check · copilot · auto · 49.8 AIC · ⌖ 7.58 AIC · ⊞ 9.5K · ◷

@pelikhan
pelikhan merged commit 149befe into main Sep 15, 2026
@pelikhan
pelikhan deleted the copilot/aw-fix-daily-ai-credits-verification branch September 15, 2026 22:00
@github-actions

Copy link
Copy Markdown
Contributor

🎉 This pull request is included in a new release.

Release: v0.89.17

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[aw] Daily Code Metrics and Trend Tracking Agent could not verify daily AI credits

3 participants