Repository navigation
fix(acp): wake quiet hosts for respawns and queued retries - #7459
Conversation
🔐 Codex Security Review
Review SummaryOverall Risk: NONE
FindingsNo concrete security, correctness, or reliability findings were identified. Notes
Generated by Codex Security Review | |
|
@buzz-security-review e4b3cb9 |
Independent combined evidence (host + companion adapter) — exact heads, no source changesThis PR (host scheduling fix): independent host-only review of exact head Companion adapter PR (pic-suite): buzz://pr?id=7473c7f5bab6568886f33b56738613f3d4321ae22b727622f4859a4157ee9850&owner=efccd8ff4cab0cf2fc878d4ce288c336756c4fb2376ba98922002e6d5b7afd8a&d=pic-suite — repaired head Composed quiet-recovery results: each case (before retry eligibility, after eligibility, circuit reopen, normal typing) redelivered the original interrupted input exactly once with one durable worker run, one worker write, one terminal, one callback; startup children retire within bounded grace; the hard turn cap is preserved (the new progress path does not bypass it). Gate status at posting: exact-range security review SUCCESS, zero findings (run 34156878857); CI success (run 34156407833). This post is independent evidence metadata only — not a formal GitHub review approval and not merge/activation authorization. Activation additionally requires the effective max-turn/idle policy check for the two originally affected managed agents (an activation-only gate, not a source blocker). |
jedwards27
left a comment
There was a problem hiding this comment.
:bot: Jude’s code review agent
Verdict: APPROVE
Reviewed: 3c7f288c60d67df78577b237e27c3dfc8831aaa1..e4b3cb9057b33554aae8a1769a17a34dd3eb96ce (exact head e4b3cb9057b33554aae8a1769a17a34dd3eb96ce)
Risk: high — this changes ACP host scheduling across respawn completion, retry eligibility, maintenance/refill, cancellation, and shutdown.
Behavior/contracts traced:
- The biased production
select!preserves shutdown priority while independently waking for respawn completion, retry eligibility, and maintenance (crates/buzz-acp/src/lib.rs:3106-3125). - Retry wakes arm only when the pool is ready and has idle capacity;
next_retry_deadlineexcludes empty, removed, and in-flight scopes (crates/buzz-acp/src/queue.rs:497-513). Existing requeue/mark_completeordering retains retry budget, backoff, and timestamps. - Consuming a respawn result clears the single-slot in-flight state, restores a successful replacement to its indexed slot, and immediately dispatches pending work (
crates/buzz-acp/src/lib.rs:3713-3758). Maintenance compacts state, refills circuit-eligible empty slots, and dispatches flushable work (crates/buzz-acp/src/lib.rs:3049-3093). - Closed respawn receivers disable only that wake arm without spinning or suppressing timers; cancellation does not consume a later result (
crates/buzz-acp/src/recovery_wake.rs:21-32,65-104). Shutdown remains first priority and aborts/drains respawns while explicitly shutting down returned children (crates/buzz-acp/src/lib.rs:3118-3125,4122-4135). - Adversarial review found no changed path that bypasses retry budgets or the hard turn cap, duplicates dispatch, strands an eligible quiet-host batch, or leaks a returned respawn child. Batching, dedup/delivery state, removed-scope handling, and circuit policy remain unchanged.
Findings: no blocking or non-blocking defect found in the reviewed range.
Author action: none.
Verification owner: integration/deployment owner for composed companion-adapter runtime replay and confirmation of effective max-turn/idle activation policy for affected managed agents.
Validation at matching exact head:
- Clean detached worktree and
git rev-parse HEADconfirmede4b3cb9057b33554aae8a1769a17a34dd3eb96ceimmediately before submission. git diff --check 3c7f288c60d67df78577b237e27c3dfc8831aaa1..HEAD— pass.- Independent exact-head
cargo test -p buzz-acp— 920 library tests + 9 integration tests, 0 failures. - Required exact-head CI results inspected: Rust lint/unit tests, Linux cross-compiles, Windows Rust, Desktop core/build/smoke/integration, relay/backend integration, PostgreSQL, security, Semgrep, zizmor, and DCO are successful. Path-inapplicable jobs are skipped.
Runtime evidence: the PR documents loopback stock-host/adaptor fixture results for fast and delayed replacement, with unmodified-base mutation evidence. A later exact-head composed report covers before/after retry eligibility, circuit reopen/refill, normal typing, bounded child retirement, exactly-once durable delivery, and preserved hard cap. These are supporting author-posted artifacts; I did not independently replay the original live incident or run the final composed process suite.
Residual risk: adapter/runtime configuration and deployment activation outside this three-file host patch. No conventional UI or native visual surface changed, so screenshot/native UI evidence is not applicable.
jedwards27
left a comment
There was a problem hiding this comment.
:bot: Jude’s code review agent
Verdict: APPROVE
Reviewed: 3c7f288c60d67df78577b237e27c3dfc8831aaa1..e4b3cb9057b33554aae8a1769a17a34dd3eb96ce (exact head e4b3cb9057b33554aae8a1769a17a34dd3eb96ce)
Risk: high — host scheduler wakeups, retry timing, respawn state, shutdown ordering, and child cleanup.
Behavior/contracts traced: Both lanes traced the production biased select!, retry-deadline eligibility, respawn result consumption and slot replacement, maintenance/refill/dispatch, closed-channel behavior, cancellation, and shutdown cleanup (crates/buzz-acp/src/lib.rs:3049-3125,3713-3758,4122-4135; queue.rs:497-513; recovery_wake.rs:21-32,65-104). Quiet-host recovery no longer depends on incidental relay/presence/typing traffic. Shutdown remains first priority; retry wakes only arm with a ready pool and idle capacity; removed/empty/in-flight scopes do not generate retry deadlines; returned children are explicitly shut down during teardown.
Findings: no concrete author-actionable defect found. The lanes found no path in the searched three-file host patch that bypasses retry budgets, duplicates dispatch, strands eligible quiet-host work, suppresses timers after a closed respawn receiver, or leaks a returned respawn child.
Author action: none.
Verification owner: integration/deployment owner for composed companion-adapter replay and activation-policy confirmation.
Validation
At matching exact head:
git diff --check 3c7f288...HEAD— pass.cargo test -p buzz-acp— pass in shared exact-head evidence: 920 library + 9 integration tests, 0 failures.- Exact-head GitHub Rust lint/unit, Windows Rust, Linux cross-compile, security, relay/backend integration, DCO, and aggregate required gates — success.
Manual/native evidence: no live incident replay or composed companion-adapter process run.
Residual risk: activation/configuration outside this three-file patch—effective max-turn/idle policy and companion-adapter runtime behavior—was not independently exercised. This is an integration/deployment confidence gap, not author rework.
e4b3cb9 to
122a205
Compare
122a205 to
76e60b8
Compare
jedwards27
left a comment
There was a problem hiding this comment.
:bot: Jude’s code review agent
Verdict: APPROVE
Reviewed: 2ebfde9147140a4036b11bfc2c6b0da68688a05c..76e60b8a8e3a53f2d7c85dbb64d84ab1a97e6fbd (exact head 76e60b8a8e3a53f2d7c85dbb64d84ab1a97e6fbd)
Risk: high — ACP host scheduling across respawn completion, retry eligibility, maintenance/refill, cancellation, shutdown, and child cleanup.
Behavior/contracts traced: Shutdown remains the first branch in the biased production select!; respawn, retry, and maintenance independently wake a quiet host (crates/buzz-acp/src/lib.rs:3135-3158). Retry timers arm only for a ready pool with idle capacity, and next_retry_deadline exposes only nonempty, non-in-flight queued scopes while retaining expired deadlines to close the crossed-deadline race (lib.rs:3136-3143; queue.rs:526-542). Respawn and retry wakes converge on the existing dispatch_pending owner; successful replacements return to their indexed slot, failures reopen that slot’s circuit, and no alternate delivery path was introduced (lib.rs:3747-3793). Closed respawn receivers disable only that source; cancellation does not consume a later result (recovery_wake.rs:21-32,65-104). Shutdown aborts respawn work and explicitly shuts down successfully returned children already queued for collection (lib.rs:4171-4184).
The production-seam regressions preserve the original event ID and received_at, assert one redelivery, and prove no second flush; empty, removed, held, and in-flight scopes do not hot-loop (queue.rs:2355-2441). The reviewed patch does not bypass retry budgets, batching, dead-letter/failure handling, or the hard turn cap.
Findings: no concrete blocking or non-blocking defect found in the three-file patch.
Author action: none.
Verification owner: base/CI owner for the pre-existing repository compile failure, followed by an exact-head required-gate rerun; integration/deployment owner for final-adapter composed runtime replay and effective max-turn/idle activation-policy confirmation.
Validation at matching exact head:
- Clean detached worktree; live PR head and local
HEADboth confirmed as76e60b8a8e3a53f2d7c85dbb64d84ab1a97e6fbdimmediately before submission. git diff --check 2ebfde9147140a4036b11bfc2c6b0da68688a05c..HEAD— pass.- Independent lanes recorded exact-head passes for
cargo fmt --all -- --checkandjust file-size-check. - Local
cargo test -p buzz-acpcould not execute because the macOS linker exited 69 until the host’s Xcode license is accepted; this is reviewer-tooling failure, not source evidence. - Exact-head Rust lint/unit/cross-compile and relay-artifact lanes fail on private
http_denialvisibility incrates/buzz-relay, outside this PR range. The failing source files are byte-identical between base and head, and matching base CI runs fail the same way. This is a pre-existing required-gate blocker owned by the base/CI repair, not PR-author rework. - Exact-range automated security review, security lanes, Semgrep, zizmor, and DCO passed.
Manual/native evidence: none. This patch has no visual/native UI surface. Prior composed adapter/runtime evidence belongs to stale head e4b3cb9 and is not promoted as exact-head proof.
Residual risk: exact-head composed adapter/runtime recovery and activation configuration remain unobserved. Required repository gates must return green after the unrelated base compile repair; that external gate, not this approval, owns merge readiness.
Signed-off-by: Logan Johnson <loganj@squareup.com>
76e60b8 to
6faaaa9
Compare
Brings 16 upstream block/buzz commits (a14107a) into the fork integration branch: ACP mention/edit steering (block#6131, block#6132), quiet-host recovery wakes (block#7459), relay NIP-FI shadow mode (block#8034, block#8062), writer lock foundations (block#7706), Goose MCP handshake (block#8037), Claude model names (block#8053), summarized thinking (block#8051) and mobile iOS changes. Merged cleanly without textual conflicts. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Signed-off-by: Arnoldinh0 <arnaudlafosse92100@gmail.com>
## Problem With `--no-presence --no-typing --heartbeat-interval 0`, a crashed ACP transport can leave work queued indefinitely. Background respawn results were only drained before `select!`; neither their arrival nor queue retry eligibility woke the host. A respawn-only wake is insufficient: replacement initialization normally precedes the independent queue backoff deadline. ## Change - Select on respawn completion, reusing the existing return-to-pool and dispatch paths. - Select on the earliest queued retry throttle, including already-expired deadlines, only while the pool is ready and has idle capacity. - Give the existing 30-second maintenance/refill check its own wake source. Circuit cooldown/refill policy is unchanged; it no longer depends on presence/typing traffic. - Disable closed respawn receivers without spinning or disabling timers; prioritize shutdown. - Keep retry backoff/budget, batching, timestamps, dedup, delivery state, and intentional host max-turn cap unchanged. Held/no-slot dispatch releases clear expired retry throttles through the existing `mark_complete` path. No reconnect/gateway changes or adapter retry engine. Based directly on main, independent of other ACP plumbing work. Closest related open PR found: block#7317 (usage-limit retry policy); this is a scheduling fix, not that policy change. ## Validation - `cargo test -p buzz-acp`: 920 library tests + 9 integration tests pass. - Five new Rust tests exercise the production wake helper and queue boundary: respawn before/after retry eligibility, backoff preservation, one-shot redelivery, empty/removed/held/in-flight scopes, closed receivers, and shutdown cancellation. - `cargo clippy -p buzz-acp --all-targets --all-features -- -D warnings`, `cargo fmt --all -- --check`, `just file-size-check`, and `git diff --check`: pass. - Actual standalone stock host binary built from this tree + installed/generated Pi adapter fixture, pinned adapter source `801d23374a588fd1154bb3da8783f6dcc9084e15` and upstream pi-acp 0.0.33. Loopback-only synthetic relay/provider; no live agent/config changes. - All optional presence/typing/heartbeat wakes disabled; 2.5-second adapter inactivity budget, 45-second callback bound, host hard cap unchanged. - Fast replacement: initialization 1.347s after requeue, retry eligible at 4.447s. Automatic original-batch recovery, one saved worker terminal and one callback, five dispatch cycles. - Delayed replacement: initialization 9.118s after requeue, retry eligible at 4.157s. Same recovery invariants, four dispatch cycles. - Both use exactly one relay input; both settle within 2.84s of the later of replacement/eligibility (including a 1.5s duplicate-observation window). This also rejects waiting for the 30s maintenance tick. - Same fixture against freshly built unmodified main `3c7f288c60d67df78577b237e27c3dfc8831aaa1`: expected failure at the 45-second callback bound. Full workspace/Desktop/mobile `just ci` not run locally; no Tauri files touched. Relevant ACP package checks above are complete. ## Review / integration gate Draft for independent review. The fixture uses the pinned pre-repair adapter to isolate this host defect, not concurrent adapter shutdown-fencing changes. Composed testing with the final adapter candidate and independent review remain required before claiming the shared lifecycle fix complete. No activation, merge, direct main push, or incident replay performed. Existing all-agents-dead exit and max-turn safety policies remain unchanged. Signed-off-by: Logan Johnson <loganj@squareup.com>
Problem
With
--no-presence --no-typing --heartbeat-interval 0, a crashed ACP transport can leave work queued indefinitely. Background respawn results were only drained beforeselect!; neither their arrival nor queue retry eligibility woke the host. A respawn-only wake is insufficient: replacement initialization normally precedes the independent queue backoff deadline.Change
mark_completepath.No reconnect/gateway changes or adapter retry engine. Based directly on main, independent of other ACP plumbing work. Closest related open PR found: #7317 (usage-limit retry policy); this is a scheduling fix, not that policy change.
Validation
cargo test -p buzz-acp: 920 library tests + 9 integration tests pass.cargo clippy -p buzz-acp --all-targets --all-features -- -D warnings,cargo fmt --all -- --check,just file-size-check, andgit diff --check: pass.801d23374a588fd1154bb3da8783f6dcc9084e15and upstream pi-acp 0.0.33. Loopback-only synthetic relay/provider; no live agent/config changes.3c7f288c60d67df78577b237e27c3dfc8831aaa1: expected failure at the 45-second callback bound.Full workspace/Desktop/mobile
just cinot run locally; no Tauri files touched. Relevant ACP package checks above are complete.Review / integration gate
Draft for independent review. The fixture uses the pinned pre-repair adapter to isolate this host defect, not concurrent adapter shutdown-fencing changes. Composed testing with the final adapter candidate and independent review remain required before claiming the shared lifecycle fix complete. No activation, merge, direct main push, or incident replay performed. Existing all-agents-dead exit and max-turn safety policies remain unchanged.