Skip to content

fix(daemon): detach managed-session viewers without reaping providers - #119

Merged
waterbang merged 3 commits into
ChesterRa:mainfrom
chriscoveries:fix/managed-session-detach
Sep 28, 2026
Merged

waterbang merged 3 commits into
ChesterRa:mainfrom
chriscoveries:fix/managed-session-detach

Conversation

@chriscoveries

Copy link
Copy Markdown
Contributor

Managed-session actors ran the claude worker under a viewer PTY (claude attach); when the viewer exited, reconcile_one and the managed reader reaped the provider job via claude stop and wrote a poisoned resume binding, so fleet-spawned actors died ~13–40s in and respawned --resume onto a dead session.

This change detaches the viewer without killing the provider: jobs survive observer teardown and are re-adopted via find_live_job, resume_eligible is invalidated on confirmed provider exit, viewers re-attach on demand in submit_batch, and the confirmed claude stop is kept for explicit actor.stop/stop_all.

Fixes

  • D1: viewer detach/PTY reap kills the live claude worker (claude stop on observer teardown) and poisons resume_eligible bindings

Testing

cargo check -p cccc-pair-daemon and cargo test -p cccc-pair-daemon — all suites pass (578 unit tests + integration/fixture suites; control-socket fixtures prove the job survives observer release and still dies on explicit stop).

Split out of #118 to keep review scope narrow.

chriscoveries and others added 2 commits September 26, 2026 09:23
A managed Actor's cccc_runtime session is its viewer attachment
(`claude attach <job>`), not the provider job — yet a reaped attach
exit drove reconcile_one -> local_headless::stop -> claude stop, and
any observer-channel end drove stop_after_process_exit -> claude stop
(SIGTERM, exit 143). Both conflated observer death with provider exit:
workers were killed ~13-40s after every spawn, and record_managed's
unconditional resume_eligible resumed the dead session on every
respawn (the recurring managed-session death loop).

- reconcile_one now detaches the viewer for managed-session actors via
  local_headless::detach_after_viewer_exit; no claude stop, no
  actor.stop record. Untracked exits still take the record path.
- stop_after_process_exit releases instead of stopping: Claude jobs
  are supervisor-owned and must not be killed on observer teardown
  (Codex/Acp keep protocol close). Released sessions carry a
  `released` flag: explicit stop still performs the confirmed kill via
  stop_after_release -> kill_and_confirm, while respawn's
  stop_for_replacement skips it so find_live_job can re-adopt the live
  job.
- When the provider job is confirmed absent from Agent View's job
  list, the runtime_session binding is invalidated
  (status=exited, resume_eligible=false, failure_count+1) so the next
  spawn is fresh instead of resuming a killed session.
- submit_batch re-attaches a detached viewer on demand — the viewer
  attachment is also the delivery channel, so detaching must not
  blackhole deliveries.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
(cherry picked from commit ed3ebe57611a68e9f07f148517a5484a1227c77e)
The fake supervisor exited on the first post-kill 'list', but the
session's 1s liveness poll and kill_and_confirm's 50ms confirmation
poll share that socket — whichever lost the race read a dead socket
until STOP_TIMEOUT (flake: 'Claude Agent View task did not stop:
Connection refused'). Serve stragglers for a bounded 1.5s window,
then exit so teardown can still await the task.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
…tach

Only a job Agent View positively reports gone is released without a confirmed
stop; a job that may still be working keeps the confirmed-stop path and its
retryable error state. This removes the released-session state and restores
the reader test's original assertions.

Add regression tests for a reaped viewer detaching without a provider kill,
viewer reattach, and invalidated resume bindings.
@waterbang

Copy link
Copy Markdown
Collaborator

Thanks for splitting this out of #118 — the focused scope made it much easier to review.

What looks right to us. A managed Claude Actor's runtime PTY is only the claude attach viewer, so tearing down the Agent View job when that viewer exits (reconcile_one → local_headless::stop) is a real bug. Detaching the terminal and re-attaching on the next delivery is the right direction.

Our concern: the managed-reader change. stop_after_process_exit now releases the session without a confirmed stop on every reader exit, not only viewer exits. But run_client also emits MANAGED_AGENT_DISCONNECTED while the job is still alive — for example, a prompt that never starts, a transcript correlation timeout, or sustained control-socket failure. Before this change those cases got a confirmed claude stop. Now CCCC marks the Actor stopped and records process_exit while the job may still be running a turn, often with permission bypass, and nobody is watching it. The PR also removes the assertion that protected this (observer failure must actually stop the job) and changes the rejected-control case from a retryable error state to stopped.

Smaller points:

  • stop_locked_inner clears released before stop_after_release succeeds. If that kill fails, a retried actor.stop sees stopped == true, drops the registry entry, and the job is never stopped.
  • detach_after_viewer_exit, reattach_viewer, invalidate_managed and stop_for_replacement have no tests. The existing reader test was rewritten rather than extended.
  • cargo fmt --all --check fails on the fake control server change in claude.rs.

Could you share how you reproduced D1? Specifically:

  1. The Claude Code version.
  2. What made claude attach exit while the job was healthy (crash, CLI upgrade, a detach key, something else).
  3. Any logs showing the job was still working when it was stopped.

On the --resume side: Agent View can normally resume an ended session from its transcript. What exactly went wrong after respawning onto a killed session?

We checked the ledgers on one of our long-running installs and could not attribute any early Claude exits to this path. That doesn't mean it can't happen; we'd just like to understand the trigger before changing stop semantics.

We have a narrower version ready. It keeps your viewer detach and reattach and the resume invalidation. It releases without a confirmed stop only when Agent View positively reports the job absent; otherwise it keeps the confirmed stop and its retryable error. It also adds tests for viewer detach (no control request is sent to the job), reattach, and invalidated bindings. Happy to push it to this branch or open it separately — whichever you prefer.

@chriscoveries

Copy link
Copy Markdown
Contributor Author

Thank you so much, @waterbang — and thank you @ChesterRa for CCCC. I'm fairly new to contributing to open source and genuinely thrilled to be helping. Your review was more careful than my PR, and your concerns are fair.

The honest answer is that we don't know what made claude attach exit. I overstated D1; I saw the effect, not the cause. You're very welcome to push your narrower version to this branch, and we're glad to run it on our installs and report back. Here's everything we could establish, in case it helps.

For context on where this comes from: we run CCCC every day across two machines, with 14 working groups and about 59 agents on a mix of runtimes (Claude Code, Codex, Hermes, Devin, opencode, pi, Grok), running continuously through the 10 days this data covers.

1. Versions. CCCC 0.4.41 on both hosts; Claude Code 2.1.280 (Linux) and 2.1.283 (macOS).

2. The exits. We went through every Claude actor stop in both installs' ledgers, 17–27 Sep:

  • Linux: 83 process_exit stops across 8 Claude actors in 5 groups, all by: system. None were initiated by a user or another actor. 59 had no exit code, 23 had exit code 1, and 1 had code 0.
  • macOS: 22 process_exit stops, all with no exit code, including a cluster of 9 in 35 minutes.
  • As far as we can find, a code-less process_exit only comes from record_process_exit(…, None) at managed_reader.rs:194, the managed-reader path. So most of our Claude exits went through the reader path your concern is about, not an observed process exit. What we can't tell from the ledger is whether the job was still running each time.
  • The code-1 exits are mostly one actor: 22 exits in 9 minutes. SessionStatus keeps only exit_code, and portable-pty reports a signal death as code 1, so we can't tell a kill from an ordinary exit(1).

3. What we ruled out. No OOM or kernel kills on either host, no Claude Code upgrade near the exits, and no scheduled job lining up with them.

4. Was the job still working when stopped? In the one case we traced end to end, no. The Claude daemon's job went blocked at 09:07:25, the daemon logged bg settled <job> (killed) at 09:07:30.011, and CCCC recorded process_exit 85 ms later. The same job was claimed again about 21 minutes later, with its session IDs unchanged. That's the daemon's state classification, not an independent liveness probe, but it suggests CCCC was reacting to a kill on the Claude side, not causing one. That fits your version.

5. Resume. All 17 managed bindings on our Linux host still read resume_eligible: true, failure_count: 0, including the Claude ones that went through the exits above, so a failed resume never seems to be recorded. On macOS, during one exit cluster, the transcript actually being written was not the session CCCC had bound for that actor.

If it's useful, we can share the full data (per-exit CSV plus method). We also have a small, additive change that records how the process ended (exit_detail: "Terminated by Killed" versus "Exited with code 1") and which path recorded the stop (source: runtime or managed_reader) on actor.stop. That would make the next occurrence diagnosable for both of us. Happy to open it as a separate PR if you'd like it.

Thank you again, @ChesterRa, for all the work you've put into CCCC. I have more to come. I love CCCC, and it's been a real pleasure building on it.

@waterbang

Copy link
Copy Markdown
Collaborator

Thank you very much for your reply. We welcome your continued feedback. We will take the time to review it carefully. Thank you again.

@waterbang
waterbang merged commit 63367db into ChesterRa:main Sep 28, 2026
6 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants