fix(engine): key chokepoint ledger paused off kill_switch eventType - #8965
fix(engine): key chokepoint ledger paused off kill_switch eventType#8965galuis116 wants to merge 1 commit into
Conversation
Budget-cap termination fires stage budget_cap with eventType kill_switch; denyResult previously keyed decision off stage, so the ledger recorded deny instead of paused. Key the ternary off eventType instead. Closes #8864
|
Superagent didn't find any vulnerabilities or security issues in this PR. |
|
Caution 🛑 LoopOver review result - fixes requiredReview updated: 2026-07-26 15:07:46 UTC
Review summary Nits — 3 non-blocking
CI checks failing
Decision drivers
Context & advisory signals — never blocks the verdict
Linked issue satisfactionAddressed Review context
Contributor next steps
Signal definitions
🧪 Chat with LoopOverAsk LoopOver a question about this PR directly in a comment — grounded only in the same cached, public-safe facts shown above, never a new claim.
Full command reference: https://loopover.ai/docs/loopover-commands 🧪 Experimental — new and may change. 🟩 Safe / merged · 🟦 Advisory · 🟨 Held for review · 🟥 Blocked / closed 💰 Earn for open-source contributions like this. Gittensor lets GitHub contributors earn for the work they already do — register to start earning →. Checked by LoopOver, a quiet PR intelligence layer for OSS maintainers.
|
|
LoopOver is closing this pull request on the maintainer's behalf (CI is failing (validate, validate-tests)). This is an automated maintenance action — to pursue this change, please open a new pull request with the issues resolved. Closed PRs may be analyzed later to improve review accuracy, but they are not automatically reopened or re-reviewed. |
…k, and release AI-review locks before a hard kill can (#9174) #8997 — a deploy restart can leave a PR wearing a decisive panel with no matching disposition. SIGTERM kills whatever pass is in flight at an arbitrary point, and the worst possible cut is between the public-surface publish (panel, gate check-run, CI aggregate — all committed) and maybeRunAgentMaintenance ever claiming its per-PR actuation lock for that head. Confirmed live on #8965: the panel published at 14:55:24Z from the dying container, and the matching close only landed at 15:07:53Z — ~12 minutes later, purely because an unrelated later sweep tick happened to re-run the whole pass from scratch. The existing outage-repair check (lastPublishedSurfaceSha !== headSha) cannot see this: the surface IS current, so it already looks healthy. maybeRunAgentMaintenance now records a marker (agent.maintenance.disposition_ considered) the moment it claims the actuation lock — the moment a real disposition attempt begins, independent of what it decides. A new standalone scan, reconcileSurfaceWithoutDisposition, rides the same sweep tick as the existing pr-outcome/pending-closure repair scans (#9026/#9031) and re-enqueues a regate for any open PR whose current-head surface carries no such marker. Deliberately NOT folded into the existing surfaceRepairPriorityPullNumbers: that function's return value drives the regular sweep's staleness-ordered fan-out and is asserted against by a large, already-green test surface (dispatch shape, ordering, backlog-restriction semantics) — an early attempt at injecting this check there flagged every "surface current, nothing else recorded" fixture in that file, which is a real, deliberate policy widening this fix makes, but one that belongs in its own independently-testable scan rather than retrofitted onto a heavily-relied-upon function blind. #8998 — an orphaned ai-review-lock (30-min TTL) starves every subsequent re-review, including an explicit maintainer re-run, with the "already in progress" placeholder for real time nobody is reviewing anything. Two of the three layers this needed were already shipped earlier this session: the boot-time flush (#9021, ORPHANED_LOCK_KEY_PATTERNS already includes ai-review-lock:*) and the explicit force-lock-steal for a maintainer's forced re-run (#9008, `steal: webhook.forceAiReview === true` already threaded into claimAiReviewLock). What was still missing: queue.stop() (#9007) lets an in-flight job finish naturally, releasing its own lock via its own finally block — but only if the orchestrator's SIGKILL grace period is long enough for that drain to complete, and a common 10-30s grace period is shorter than an AI-review LLM call legitimately runs. A new held-lock-registry tracks every real ai-review-lock claim this process currently holds; the shutdown handler now releases all of them FIRST, before anything else in the shutdown sequence, so the lock is gone the instant SIGTERM/SIGINT arrives regardless of whether the rest of shutdown gets time to finish. Complementary to the boot-time flush, not a replacement for it: a true `kill -9` delivers no signal at all, so that backstop still matters. #8999 — a lock-contended hold applied the sticky manual-review label, and the label then froze every LATER pass from running a FRESH AI review at all (only replay) — meaning the one thing that could produce the real verdict needed to auto-clear the label (already implemented, #9009) could never run, because the freeze it depends on being lifted is gated by the very label it's trying to clear. #9009 only handled the case where a fresh review DID run and showed contention had resolved; it could not reach that state on a repo requiring blocking AI review, because the freeze blocks the fresh attempt outright. Fixed by exempting the freeze specifically when the label's own provenance — per the same transient marker #9009 already reads/writes — is the lock-contention hold and nothing else. This does not reopen the gaming surface the freeze exists to close (a contributor iterating pushes to buy a fresh verdict): a lock-contention hold is an infra artifact of two ORB passes racing, not a contributor action, and if some OTHER reason also justifies the hold, that reason re-applies the label on this same pass's own disposition regardless of the exemption. #9013's largest remaining piece (a single per-PR mutex spanning the public-surface publish AND the disposition plan/execute, currently two separately-claimed critical sections) is NOT included here — an initial attempt showed it needs a moderate refactor of both giant call sites (the sweep/CI- completion path and the webhook path) with real risk of subtly changing span/catch/decisionOutcome semantics the existing test suite exercises extensively, and that deserves its own dedicated, carefully-tested pass rather than a rushed addition alongside three already-substantial fixes. Targeted tests only (no full local gate run this pass, per instruction): typecheck clean, and every touched/new test file passes locally — 39 new tests across surface-disposition-reconciler.test.ts, held-lock-registry.test.ts, and the extended job-dispatch.test.ts fan-out coverage — plus a broader targeted sweep of 967 tests across the queue/lifecycle/transient-lock/precision-breaker suites most likely to interact with the actuation-lock, freeze, and sweep- priority code paths this change touches. CI runs the full gate on push. Closes #8997 Closes #8998 Closes #8999
Summary
denyResult's ledgerdecisionternary offeventType === kill_switchinstead ofstage === kill_switch, so budget-cap termination (stagebudget_cap, eventTypekill_switch) recordspausedlike the top-level kill-switch path.ledgerEvent.decision === paused, and add a vitest suite that covers the kill-switch / soft-deny / termination branches for codecov patch coverage.Closes #8864
Scope
type(scope): short summaryConventional Commit format, for examplefix(api): restore profile access checks.CONTRIBUTING.mdand does not reintroduce GitHub Pages, VitePress,site/, orCNAME.Closes #123) — a linked open issue is required for every contributor PR.Validation
git diff --checknpm run actionlintnpm run typechecknpm run test:coveragelocally;codecov/patchrequires ≥99% coverage of the lines AND branches you changed (aim for 100% on your diff so CI variance does not fail near the threshold). Global coverage is a non-blocking trend with a loose 90% backstop, not the gate.npm run test:workersnpm run build:mcpnpm run test:mcp-packnpm run ui:openapi:checknpm run ui:lintnpm run ui:typechecknpm run ui:buildnpm audit --audit-level=moderateIf any required check was skipped, explain why:
packages/loopover-enginebuild +chokepoint.test.js(27 pass) andvitest run test/unit/engine-governor-chokepoint-kill-switch-decision.test.ts(3 pass). Unrelated UI/MCP/workers/actionlint checks not exercised.Safety
UI Evidencesection below with JPG/JPEG or PNG screenshots arranged as organized, captioned, clickable thumbnails. SVG screenshots are not used as review evidence. Review-only screenshots or recordings are not committed to the repository.UI Evidence
N/A — engine-only governor ledger decision fix; no visible UI.
Notes