You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Daily AI Credits guardrail rejects legacy successful --use-samples runs created before #61053 because deterministic sample replay produced no billable engine usage, left agent/token_usage.jsonl empty, and did not include the new agent/execution.json evidence. A later run compiled after #61053 treats that legacy zero-cost artifact as missing accounting, fails activation, and skips the agent and safe-output jobs.
This caused a cascade across the githubnext/gh-aw-test E2E suite on gh-aw main.
PR #61053 was merged as 976492cdd111415c7196bc85f03f0aa343368c45 at 2026-09-15T22:01:51Z and is an ancestor of the tested commit. It correctly preserves execution evidence for newly generated artifacts. The runs poisoning this guardrail window, such as 34927589119, 34927991036, and 34928383174, ran around 2026-09-15T04:07–04:23Z, before #61053 merged, and remained inside the rolling 24-hour window when this E2E ran.
The deterministic replay path consumes zero AI Credits, but the resulting usage artifact has an empty agent/token_usage.jsonl and no agent/execution.json evidence.
Daily workflow AI Credits are unknown: Missing accounting for executed agent component in run <id> (attempt 1, job <id>, conclusion success): agent/token_usage.jsonl is empty; agent_usage.jsonl is missing; agent_usage.json is missing
Expected: a legacy successful deterministic sample replay is recognized as zero-cost, or otherwise does not block all subsequent runs during the #61053 rollout window.
Actual: sumCoveredComponents treats the successful agent/replay job as billable but unaccounted, reports transient_error, and fails activation.
Additional workflows remained queued behind the cascade until the test harness timed out and disabled them.
Suggested fix
Add a narrowly scoped compatibility path for pre-#61053 successful sample artifacts, for example by recognizing legacy samples-mode workflow metadata alongside an empty authoritative agent accounting file. Avoid treating all empty successful accounting as zero, since that could hide real live-engine accounting failures. Newly generated artifacts should continue using the execution evidence added by #61053.
Deduplication
Related: #61053 fixes this case for newly generated artifacts by preserving execution evidence. #61137 covers a failed agent job whose accounting file is entirely missing. This report is the remaining rollout gap for successful pre-#61053 samples-mode artifacts where the primary file exists but is empty and the new evidence is unavailable. Exact searches for "conclusion success" "token_usage.jsonl is empty" and samples/daily-AIC terms found no open duplicate.
Summary
Daily AI Credits guardrail rejects legacy successful
--use-samplesruns created before #61053 because deterministic sample replay produced no billable engine usage, leftagent/token_usage.jsonlempty, and did not include the newagent/execution.jsonevidence. A later run compiled after #61053 treats that legacy zero-cost artifact as missing accounting, fails activation, and skips the agent and safe-output jobs.This caused a cascade across the
githubnext/gh-aw-testE2E suite on gh-aw main.PR #61053 was merged as
976492cdd111415c7196bc85f03f0aa343368c45at2026-09-15T22:01:51Zand is an ancestor of the tested commit. It correctly preserves execution evidence for newly generated artifacts. The runs poisoning this guardrail window, such as34927589119,34927991036, and34928383174, ran around2026-09-15T04:07–04:23Z, before #61053 merged, and remained inside the rolling 24-hour window when this E2E ran.Reproduction
gh aw compile --use-samplesand run it successfully.agent/token_usage.jsonland noagent/execution.jsonevidence.Expected: a legacy successful deterministic sample replay is recognized as zero-cost, or otherwise does not block all subsequent runs during the #61053 rollout window.
Actual:
sumCoveredComponentstreats the successful agent/replay job as billable but unaccounted, reportstransient_error, and fails activation.E2E evidence
mainatac594b6525e3--use-samples; AI engine was not calledRepresentative log evidence from run
35051877133, activation at2026-09-16T03:31:34Z:Affected started tests included:
test-copilot-add-discussion-commenttest-copilot-close-issuetest-copilot-add-labelstest-copilot-add-commenttest-copilot-assign-to-usertest-copilot-close-discussiontest-copilot-update-issuetest-copilot-replace-labeltest-copilot-remove-labelstest-copilot-update-discussiontest-copilot-unassign-from-usertest-copilot-assign-milestonetest-copilot-mark-pull-request-as-ready-for-reviewtest-copilot-add-reviewertest-copilot-close-pull-requesttest-copilot-update-pull-requesttest-copilot-nosandbox-add-labelsAdditional workflows remained queued behind the cascade until the test harness timed out and disabled them.
Suggested fix
Add a narrowly scoped compatibility path for pre-#61053 successful sample artifacts, for example by recognizing legacy samples-mode workflow metadata alongside an empty authoritative agent accounting file. Avoid treating all empty successful accounting as zero, since that could hide real live-engine accounting failures. Newly generated artifacts should continue using the execution evidence added by #61053.
Deduplication
Related: #61053 fixes this case for newly generated artifacts by preserving execution evidence. #61137 covers a failed agent job whose accounting file is entirely missing. This report is the remaining rollout gap for successful pre-#61053 samples-mode artifacts where the primary file exists but is empty and the new evidence is unavailable. Exact searches for
"conclusion success" "token_usage.jsonl is empty"and samples/daily-AIC terms found no open duplicate.