Repository navigation
Live operator limits: tighten-only GPU ceiling and attempt width - #436
Conversation
An operator sharing a cluster account with a fleet could only slow it by editing a target's contract (a reviewed PR) or by HOLD_LAUNCHES, which lets in-flight runs keep launching. `<state root>/limits.toml` now sets a GPU ceiling (fleet-wide and per target) and a per-target max_active_attempts, re-read at every admission. Values only tighten the contract; a malformed file blocks new GPU admissions and attempts and is reported. Usage comes from one scheduler snapshot per check: jobs are attributed to a target by the run id in their name, so the operator's own jobs never count and requeued jobs count as the scheduler reports them. Every GPU-bearing submission path is checked; an author's over-ceiling launch is refused with a message and no launch charge, while evaluations wait for capacity. Two simultaneous admissions can overshoot by at most one batch. With no limits file, admission is unchanged. `outerloop limits` shows the effective values. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
There was a problem hiding this comment.
Round 1 — reviewed head 7aa87294 — reviewer summarizer:hermes/gpt-5.6-terra over coverage+credentials+deployment+general+lifecycle+prose.
terra
Advisory findings from outerloop — the code owner decides. Reply to disagree; the outerloop:no-review label opts this PR out.
Verdict: 4 blocking, 0 advisory.
3 findings attached to the lines below.
Intake submissions do not reserve an attempt slot. [coverage] After submit succeeds, this path returns without writing a pending marker, while its width check counts only existing run records, so another tick can submit a different intake issue before the first queued climb creates its record and exceed max_active_attempts. (src/outerloop/tick.py:3066; high confidence)
Blocking findings: intake admission can exceed max_active_attempts because queued submissions do not reserve capacity; launch, evaluation, and wake scheduler job names can exceed Slurm limits for valid long identifiers. Rejected findings: none; the four reported findings address distinct code locations and submission paths.
- Launch, evaluation and wake names stay within 128 characters. Names that fit are unchanged; longer ones carry a short stable run key, and usage attributes a job by run id or run key. - Intake counts queued attempt jobs toward max_active_attempts. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
# Conflicts: # CHANGELOG.md
|
Compatibility statement (added for the 0.3.0rc1 release audit, per RELEASING.md).
|
An operator who shares a cluster account with a fleet had two ways to slow it down: edit a target's contract (a reviewed PR in the target repo) or create
HOLD_LAUNCHES, which blocks new runs but lets in-flight runs keep launching. Neither frees GPUs quickly, and a contract change is the wrong tool for a temporary need.Change
<state root>/limits.toml:[defaults]and[targets."owner/repo"], withmax_gpus(a GPU ceiling for everything the fleet submits) andmax_active_attempts. Re-read at every admission; no restart.outerloop limits.squeue --me --jsonsnapshot per check. A job counts toward a target only when its name carries a run id that exists under<root>/runs/, so the operator's own jobs never count and requeued jobs count as the scheduler reports them. Handles the wrapped numeric fields of recent Slurm JSON as well as the older plain form.outerloop limits: effective limits and current fleet GPU usage per target (read-only).The contract schema is unchanged.
Verified
🤖 Generated with Claude Code