fix(daemon): detach managed-session viewers without reaping providers - #119
Conversation
A managed Actor's cccc_runtime session is its viewer attachment (`claude attach <job>`), not the provider job — yet a reaped attach exit drove reconcile_one -> local_headless::stop -> claude stop, and any observer-channel end drove stop_after_process_exit -> claude stop (SIGTERM, exit 143). Both conflated observer death with provider exit: workers were killed ~13-40s after every spawn, and record_managed's unconditional resume_eligible resumed the dead session on every respawn (the recurring managed-session death loop). - reconcile_one now detaches the viewer for managed-session actors via local_headless::detach_after_viewer_exit; no claude stop, no actor.stop record. Untracked exits still take the record path. - stop_after_process_exit releases instead of stopping: Claude jobs are supervisor-owned and must not be killed on observer teardown (Codex/Acp keep protocol close). Released sessions carry a `released` flag: explicit stop still performs the confirmed kill via stop_after_release -> kill_and_confirm, while respawn's stop_for_replacement skips it so find_live_job can re-adopt the live job. - When the provider job is confirmed absent from Agent View's job list, the runtime_session binding is invalidated (status=exited, resume_eligible=false, failure_count+1) so the next spawn is fresh instead of resuming a killed session. - submit_batch re-attaches a detached viewer on demand — the viewer attachment is also the delivery channel, so detaching must not blackhole deliveries. Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> (cherry picked from commit ed3ebe57611a68e9f07f148517a5484a1227c77e)
The fake supervisor exited on the first post-kill 'list', but the session's 1s liveness poll and kill_and_confirm's 50ms confirmation poll share that socket — whichever lost the race read a dead socket until STOP_TIMEOUT (flake: 'Claude Agent View task did not stop: Connection refused'). Serve stragglers for a bounded 1.5s window, then exit so teardown can still await the task. Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
…tach Only a job Agent View positively reports gone is released without a confirmed stop; a job that may still be working keeps the confirmed-stop path and its retryable error state. This removes the released-session state and restores the reader test's original assertions. Add regression tests for a reaped viewer detaching without a provider kill, viewer reattach, and invalidated resume bindings.
|
Thanks for splitting this out of #118 — the focused scope made it much easier to review. What looks right to us. A managed Claude Actor's runtime PTY is only the Our concern: the managed-reader change. Smaller points:
Could you share how you reproduced D1? Specifically:
On the We checked the ledgers on one of our long-running installs and could not attribute any early Claude exits to this path. That doesn't mean it can't happen; we'd just like to understand the trigger before changing stop semantics. We have a narrower version ready. It keeps your viewer detach and reattach and the resume invalidation. It releases without a confirmed stop only when Agent View positively reports the job absent; otherwise it keeps the confirmed stop and its retryable error. It also adds tests for viewer detach (no control request is sent to the job), reattach, and invalidated bindings. Happy to push it to this branch or open it separately — whichever you prefer. |
|
Thank you so much, @waterbang — and thank you @ChesterRa for CCCC. I'm fairly new to contributing to open source and genuinely thrilled to be helping. Your review was more careful than my PR, and your concerns are fair. The honest answer is that we don't know what made For context on where this comes from: we run CCCC every day across two machines, with 14 working groups and about 59 agents on a mix of runtimes (Claude Code, Codex, Hermes, Devin, opencode, pi, Grok), running continuously through the 10 days this data covers. 1. Versions. CCCC 0.4.41 on both hosts; Claude Code 2.1.280 (Linux) and 2.1.283 (macOS). 2. The exits. We went through every Claude actor stop in both installs' ledgers, 17–27 Sep:
3. What we ruled out. No OOM or kernel kills on either host, no Claude Code upgrade near the exits, and no scheduled job lining up with them. 4. Was the job still working when stopped? In the one case we traced end to end, no. The Claude daemon's job went 5. Resume. All 17 managed bindings on our Linux host still read If it's useful, we can share the full data (per-exit CSV plus method). We also have a small, additive change that records how the process ended ( Thank you again, @ChesterRa, for all the work you've put into CCCC. I have more to come. I love CCCC, and it's been a real pleasure building on it. |
|
Thank you very much for your reply. We welcome your continued feedback. We will take the time to review it carefully. Thank you again. |
Managed-session actors ran the claude worker under a viewer PTY (
claude attach); when the viewer exited,reconcile_oneand the managed reader reaped the provider job viaclaude stopand wrote a poisoned resume binding, so fleet-spawned actors died ~13–40s in and respawned--resumeonto a dead session.This change detaches the viewer without killing the provider: jobs survive observer teardown and are re-adopted via
find_live_job,resume_eligibleis invalidated on confirmed provider exit, viewers re-attach on demand insubmit_batch, and the confirmedclaude stopis kept for explicitactor.stop/stop_all.Fixes
claude stopon observer teardown) and poisonsresume_eligiblebindingsTesting
cargo check -p cccc-pair-daemonandcargo test -p cccc-pair-daemon— all suites pass (578 unit tests + integration/fixture suites; control-socket fixtures prove the job survives observer release and still dies on explicit stop).Split out of #118 to keep review scope narrow.