Skip to content

bug: agent-server.py spins at ~83% CPU after Claude CLI subprocess becomes defunct #405

Description

@vybe

Summary

When a scheduled task finishes (or terminates) on a running agent container, the claude CLI subprocess can become a defunct (zombie) child of python3 /app/agent-server.py. The parent agent-server is not reaping it and spins at ~83% CPU indefinitely, making the entire agent HTTP server unresponsive — every endpoint times out (including /, /api/session, /api/files/download) even though the container remains in running state. Workaround is a container restart.

Component

Agent Runtime / Base Image (docker/base-image/agent_server/services/claude_code.py)

Priority

P2 — one agent becomes unreachable (scheduled and ad-hoc tasks fail against it), but there is a workaround (container restart) and no cross-agent impact.

Error

No exception in the agent-server logs — the process is simply stuck. Last log line before the hang:

INFO:agent_server.services.process_registry:[ProcessRegistry] Registered execution <exec-id>

Backend side logs the symptom repeatedly:

Failed to get session info for <agent>:
Circuit OPENED for agent <agent> after N failures

(Circuit re-opens every ~35 s as the backend polls; two pollers run in parallel, so you see pairs of lines.)

Location

  • File: docker/base-image/agent_server/services/claude_code.py
  • Likely area: the asyncio subprocess reader / reaper that handles the claude --print --output-format stream-json --verbose … child. docker/base-image/agent_server/services/process_registry.py cleanup hooks may also be involved.

Root Cause (hypothesis)

The agent-server spawns claude as a subprocess and streams its stdout. When the subprocess exits (normally or abnormally) without its stdout/stderr being fully drained, the reader loop appears to busy-spin instead of detecting EOF + waitpid-ing the child. The child becomes a zombie, the loop keeps running, and the event loop is starved from handling new HTTP requests.

Reproduction Steps

  1. Deploy an agent container from trinity-agent-base:latest (as of 2026-04-19).
  2. Trigger a scheduled task that invokes a slash-command workflow (e.g. a long-running digest-style command with a multi-thousand-char system-prompt append, 50-turn cap).
  3. Let Claude Code finish / be cancelled / hit the turn cap.
  4. Observe (after a few minutes):
    • docker stats agent-<name> → ~80–100% CPU
    • docker exec agent-<name> ps auxf → shows python3 /app/agent-server.py pegged, [claude] <defunct> child
    • All HTTP endpoints on the agent time out from the backend container
    • Backend repeatedly opens the circuit breaker for the agent

Observed environment at repro

  • Task: 50-turn cap, --append-system-prompt ~5.5k chars, 30-minute timeout
  • Model: claude-opus-4-6
  • Claude CLI invocation: claude --print --output-format stream-json --verbose …
  • Process table inside container (anonymised):
    UID   PID %CPU STAT  COMMAND
    dev   1   0.0  Ss    /bin/bash /app/startup.sh
    root  17  0.0  S     sudo /usr/sbin/sshd -D
    dev   18  82.8 Sl    python3 /app/agent-server.py        <-- spinning
    dev   35  0.6  Z      \_ [claude] <defunct>               <-- zombie
    dev   22  0.0  S     tail -f /dev/null
    

Workaround

docker restart agent-<name>

After restart: CPU drops to ~0.2%, endpoints return 200, backend logs Circuit CLOSED for agent <agent> (recovered).

Suggested investigation / fix

  1. Audit the subprocess read loop in claude_code.py — ensure:
    • EOF on stdout breaks the loop
    • await proc.wait() (or equivalent) is called in all exit paths
    • The reader doesn't loop on an empty read when the pipe is closed
  2. Add a watchdog: if the subprocess is not alive but the reader hasn't returned in N seconds, force-break.
  3. Confirm process_registry cleanup always runs (finally block), even when the CLI exits non-zero or is killed by the turn cap.
  4. Consider reaping any defunct children on a periodic health check.

Environment

  • Trinity version: 13a0018 (feat: branch ownership enforcement S7, Branch ownership enforcement (S7) #382, feat: branch ownership enforcement (S7, #382) #396)
  • Base image files changed in the affecting deploy (docker/base-image/):
    • agent_server/routers/files.py
    • agent_server/routers/git.py
    • agent_server/utils/git_conflict.py
    • startup.sh
      (None of these are the obvious culprit — hypothesis is the issue is latent in claude_code.py and was triggered by timing/task shape, not by a change in this release.)
  • Docker base image: trinity-agent-base:latest / :0.9.0
  • Host OS: Linux (Ubuntu-class VM)

Related

  • docker/base-image/agent_server/services/claude_code.py
  • docker/base-image/agent_server/services/process_registry.py
  • docker/base-image/agent_server/routers/chat.py (task kickoff path)

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions