You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
bug: agent-server.py spins at ~83% CPU after Claude CLI subprocess becomes defunct #405
When a scheduled task finishes (or terminates) on a running agent container, the claude CLI subprocess can become a defunct (zombie) child of python3 /app/agent-server.py. The parent agent-server is not reaping it and spins at ~83% CPU indefinitely, making the entire agent HTTP server unresponsive — every endpoint times out (including /, /api/session, /api/files/download) even though the container remains in running state. Workaround is a container restart.
Component
Agent Runtime / Base Image (docker/base-image/agent_server/services/claude_code.py)
Priority
P2 — one agent becomes unreachable (scheduled and ad-hoc tasks fail against it), but there is a workaround (container restart) and no cross-agent impact.
Error
No exception in the agent-server logs — the process is simply stuck. Last log line before the hang:
Likely area: the asyncio subprocess reader / reaper that handles the claude --print --output-format stream-json --verbose … child. docker/base-image/agent_server/services/process_registry.py cleanup hooks may also be involved.
Root Cause (hypothesis)
The agent-server spawns claude as a subprocess and streams its stdout. When the subprocess exits (normally or abnormally) without its stdout/stderr being fully drained, the reader loop appears to busy-spin instead of detecting EOF + waitpid-ing the child. The child becomes a zombie, the loop keeps running, and the event loop is starved from handling new HTTP requests.
Reproduction Steps
Deploy an agent container from trinity-agent-base:latest (as of 2026-04-19).
Trigger a scheduled task that invokes a slash-command workflow (e.g. a long-running digest-style command with a multi-thousand-char system-prompt append, 50-turn cap).
Let Claude Code finish / be cancelled / hit the turn cap.
Claude CLI invocation: claude --print --output-format stream-json --verbose …
Process table inside container (anonymised):
UID PID %CPU STAT COMMAND
dev 1 0.0 Ss /bin/bash /app/startup.sh
root 17 0.0 S sudo /usr/sbin/sshd -D
dev 18 82.8 Sl python3 /app/agent-server.py <-- spinning
dev 35 0.6 Z \_ [claude] <defunct> <-- zombie
dev 22 0.0 S tail -f /dev/null
Workaround
docker restart agent-<name>
After restart: CPU drops to ~0.2%, endpoints return 200, backend logs Circuit CLOSED for agent <agent> (recovered).
Suggested investigation / fix
Audit the subprocess read loop in claude_code.py — ensure:
EOF on stdout breaks the loop
await proc.wait() (or equivalent) is called in all exit paths
The reader doesn't loop on an empty read when the pipe is closed
Add a watchdog: if the subprocess is not alive but the reader hasn't returned in N seconds, force-break.
Confirm process_registry cleanup always runs (finally block), even when the CLI exits non-zero or is killed by the turn cap.
Consider reaping any defunct children on a periodic health check.
Base image files changed in the affecting deploy (docker/base-image/):
agent_server/routers/files.py
agent_server/routers/git.py
agent_server/utils/git_conflict.py
startup.sh
(None of these are the obvious culprit — hypothesis is the issue is latent in claude_code.py and was triggered by timing/task shape, not by a change in this release.)
Docker base image: trinity-agent-base:latest / :0.9.0
Summary
When a scheduled task finishes (or terminates) on a running agent container, the
claudeCLI subprocess can become a defunct (zombie) child ofpython3 /app/agent-server.py. The parent agent-server is not reaping it and spins at ~83% CPU indefinitely, making the entire agent HTTP server unresponsive — every endpoint times out (including/,/api/session,/api/files/download) even though the container remains inrunningstate. Workaround is a container restart.Component
Agent Runtime / Base Image (
docker/base-image/agent_server/services/claude_code.py)Priority
P2 — one agent becomes unreachable (scheduled and ad-hoc tasks fail against it), but there is a workaround (container restart) and no cross-agent impact.
Error
No exception in the agent-server logs — the process is simply stuck. Last log line before the hang:
Backend side logs the symptom repeatedly:
(Circuit re-opens every ~35 s as the backend polls; two pollers run in parallel, so you see pairs of lines.)
Location
docker/base-image/agent_server/services/claude_code.pyclaude --print --output-format stream-json --verbose …child.docker/base-image/agent_server/services/process_registry.pycleanup hooks may also be involved.Root Cause (hypothesis)
The agent-server spawns
claudeas a subprocess and streams its stdout. When the subprocess exits (normally or abnormally) without its stdout/stderr being fully drained, the reader loop appears to busy-spin instead of detecting EOF + waitpid-ing the child. The child becomes a zombie, the loop keeps running, and the event loop is starved from handling new HTTP requests.Reproduction Steps
trinity-agent-base:latest(as of 2026-04-19).docker stats agent-<name>→ ~80–100% CPUdocker exec agent-<name> ps auxf→ showspython3 /app/agent-server.pypegged,[claude] <defunct>childObserved environment at repro
--append-system-prompt~5.5k chars, 30-minute timeoutclaude-opus-4-6claude --print --output-format stream-json --verbose …Workaround
After restart: CPU drops to ~0.2%, endpoints return 200, backend logs
Circuit CLOSED for agent <agent> (recovered).Suggested investigation / fix
claude_code.py— ensure:await proc.wait()(or equivalent) is called in all exit pathsprocess_registrycleanup always runs (finally block), even when the CLI exits non-zero or is killed by the turn cap.Environment
13a0018(feat: branch ownership enforcement S7, Branch ownership enforcement (S7) #382, feat: branch ownership enforcement (S7, #382) #396)docker/base-image/):agent_server/routers/files.pyagent_server/routers/git.pyagent_server/utils/git_conflict.pystartup.sh(None of these are the obvious culprit — hypothesis is the issue is latent in
claude_code.pyand was triggered by timing/task shape, not by a change in this release.)trinity-agent-base:latest/:0.9.0Related
docker/base-image/agent_server/services/claude_code.pydocker/base-image/agent_server/services/process_registry.pydocker/base-image/agent_server/routers/chat.py(task kickoff path)