Skip to content

A runner's remaining subscription usage is not part of eligibility — an exhausted machine still takes dispatches, and a reset time is a schedulable fact nothing can express #858

Description

@mkreyman

A runner's remaining subscription usage is not part of eligibility, so a machine that cannot start a session for two days is still handed dispatches — and a fleet with no capacity now but capacity at 10:00 tomorrow looks identical to a fleet with none at all.

Raised by Mark, 2026-09-15.

What is wrong today

Runners.accepts?/5 decides eligibility on: connected, not draining, repository declared, kind declared, and in_flight < max_sessions. Every session runs on the machine's claude.ai subscription — the runner unsets ANTHROPIC_API_KEY and refuses to start unless its launcher strips it — so the scarce resource is the subscription WINDOW, and nothing in that list can see it.

Two consequences, and the second is the one worth building for:

  1. A runner out of quota is still "available". Socket up, not draining, slots free — so available_runner/2 picks it and the dispatch goes out. The session then fails immediately or stalls, the ledger row burns a slot until Capacity.heal/3 releases it, and the story goes back to the queue to be handed to the same machine a minute later.

  2. "No capacity now" and "no capacity until 10:00 tomorrow" are the same answer. DispatchDriver and TriageDispatcher both answer :no_runner and retry every minute. With a reset time they could report when capacity returns, which is a fact an operator can act on and a scheduler could eventually use. Today a fleet that will be fine in the morning is indistinguishable from one that is broken.

The open question, which decides the shape

Does the CLI emit a reset time when a session exhausts the limit? claude --help lists no usage subcommand, so there is no documented surface to poll; the failure itself is the honest source. Asked of the loopctl-runner maintaining session, since only that side can observe it:

  • what the runner sees on exhaustion — exit code, stderr, a structured error on the stream-json output — and whether it carries a reset timestamp;
  • whether that is distinguishable from a transient 429, which wants the opposite handling (a five-hour window excludes the machine for hours; a rate limit is a retry);
  • whether anything under ~/.claude holds it readably, asked only to know if a poll is possible — building on an undocumented file is not the preferred route.

Proposed shape, if a reset time is available

  • The runner records the reset it was GIVEN, never one it computed, and reports exhausted_until on its next status.
  • RunnerStatus gains that optional field. Additive, and silence MUST mean available — the same asymmetry as Kinds.implied_by_silence/0. A control plane that excluded quiet runners would idle the whole fleet on a field nobody had shipped yet.
  • Runners.accepts?/5 excludes a runner whose exhausted_until is in the future, so the rule lives in the ONE reader every dispatcher already consults rather than being re-implemented in each.
  • The flag clears when a session succeeds, so a wrong or stale reset self-corrects rather than stranding a machine.
  • :no_runner gains the earliest reset across the eligible-but-exhausted set, so the driver's log and the operator can see when the queue will move.

What this is NOT

Not a scheduler. "Defer until T" is the whole of it; a calendar, a priority queue across machines, or holding work for a machine that might reset sooner are all out of scope and should be argued separately if ever wanted.

If the CLI emits no reset time, the fallback is weaker but still worth having: derive exhaustion from repeated immediate failures on one machine and exclude it for a fixed backoff, with no claim about when it returns.

Acceptance

  • A runner reporting exhausted_until in the future is not selected by available_runner/2 or by the driver, proved by a test.
  • A runner that reports NOTHING is selected exactly as today, proved by a test — this is the clause that must not regress.
  • :no_runner carries the earliest known reset when one is known, and nothing when none is.
  • The contract's new field is additive with a version bump, and a runner at the previous version keeps working unchanged.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions