Skip to content

Live operator limits: tighten-only GPU ceiling and attempt width - #436

Merged
renmengye merged 3 commits into
mainfrom
feat/operator-limits
Sep 28, 2026
Merged

renmengye merged 3 commits into
mainfrom
feat/operator-limits

Conversation

@renmengye

Copy link
Copy Markdown
Member

An operator who shares a cluster account with a fleet had two ways to slow it down: edit a target's contract (a reviewed PR in the target repo) or create HOLD_LAUNCHES, which blocks new runs but lets in-flight runs keep launching. Neither frees GPUs quickly, and a contract change is the wrong tool for a temporary need.

Change

  • <state root>/limits.toml: [defaults] and [targets."owner/repo"], with max_gpus (a GPU ceiling for everything the fleet submits) and max_active_attempts. Re-read at every admission; no restart.
  • Tighten-only: the effective value is the minimum of the contract and the operator's value. No operator value can loosen a contract budget. A malformed file blocks new GPU admissions and attempts and is reported in the log and by outerloop limits.
  • Usage from the scheduler: one squeue --me --json snapshot per check. A job counts toward a target only when its name carries a run id that exists under <root>/runs/, so the operator's own jobs never count and requeued jobs count as the scheduler reports them. Handles the wrapped numeric fields of recent Slurm JSON as well as the older plain form.
  • Enforced at every GPU-bearing submission path: author experiments and sweeps, evaluations and verification, fresh/steward/self-initiated attempts, and wake/session jobs. An author's over-ceiling launch is refused with a message and no launch charge; evaluations wait for capacity; a combined submit refuses only the over-cap launches. Control-plane tick jobs are CPU-only (asserted by a test).
  • Two simultaneous admissions can overshoot the ceiling by at most one batch (documented).
  • No limits file: admission is unchanged and no scheduler query is made.
  • outerloop limits: effective limits and current fleet GPU usage per target (read-only).

The contract schema is unchanged.

Verified

  • Full gate: 2218 passed, 2 skipped; ruff, format, mypy clean; mutation checks fail all 8 submission-path tests when admission is bypassed.
  • Read-only against a live Slurm 25.05 cluster: the usage query parsed real output and attributed only the fleet's sweep jobs to their target, ignoring the operator's own GPU jobs.

🤖 Generated with Claude Code

An operator sharing a cluster account with a fleet could only slow it by
editing a target's contract (a reviewed PR) or by HOLD_LAUNCHES, which lets
in-flight runs keep launching. `<state root>/limits.toml` now sets a GPU
ceiling (fleet-wide and per target) and a per-target max_active_attempts,
re-read at every admission. Values only tighten the contract; a malformed
file blocks new GPU admissions and attempts and is reported.

Usage comes from one scheduler snapshot per check: jobs are attributed to a
target by the run id in their name, so the operator's own jobs never count
and requeued jobs count as the scheduler reports them. Every GPU-bearing
submission path is checked; an author's over-ceiling launch is refused with a
message and no launch charge, while evaluations wait for capacity. Two
simultaneous admissions can overshoot by at most one batch. With no limits
file, admission is unchanged. `outerloop limits` shows the effective values.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

@github-actions github-actions Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Round 1 — reviewed head 7aa87294 — reviewer summarizer:hermes/gpt-5.6-terra over coverage+credentials+deployment+general+lifecycle+prose.

terra
Advisory findings from outerloop — the code owner decides. Reply to disagree; the outerloop:no-review label opts this PR out.

Verdict: 4 blocking, 0 advisory.

3 findings attached to the lines below.

Intake submissions do not reserve an attempt slot. [coverage] After submit succeeds, this path returns without writing a pending marker, while its width check counts only existing run records, so another tick can submit a different intake issue before the first queued climb creates its record and exceed max_active_attempts. (src/outerloop/tick.py:3066; high confidence)

Blocking findings: intake admission can exceed max_active_attempts because queued submissions do not reserve capacity; launch, evaluation, and wake scheduler job names can exceed Slurm limits for valid long identifiers. Rejected findings: none; the four reported findings address distinct code locations and submission paths.

Comment thread src/outerloop/attempt.py Outdated
Comment thread src/outerloop/dispatch.py
Comment thread src/outerloop/tick.py Outdated
- Launch, evaluation and wake names stay within 128 characters. Names that
  fit are unchanged; longer ones carry a short stable run key, and usage
  attributes a job by run id or run key.
- Intake counts queued attempt jobs toward max_active_attempts.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@renmengye renmengye added the outerloop:review re-request the advisory review label Sep 28, 2026

@github-actions github-actions Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Round 1 — reviewed head 02f794fb — reviewer hermes/gpt-5.6-terra.

terra
Advisory findings from outerloop — the code owner decides. Reply to disagree; the outerloop:no-review label opts this PR out.

Verdict: no defects found.

@renmengye renmengye added outerloop:review re-request the advisory review and removed outerloop:review re-request the advisory review labels Sep 28, 2026

@github-actions github-actions Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Round 2 — reviewed head a097791f — reviewer hermes/gpt-5.6-terra.

terra
Advisory findings from outerloop — the code owner decides. Reply to disagree; the outerloop:no-review label opts this PR out.

Verdict: no defects found.

@renmengye
renmengye merged commit 4098fb4 into main Sep 28, 2026
5 checks passed
@renmengye
renmengye deleted the feat/operator-limits branch September 28, 2026 04:06
@renmengye

Copy link
Copy Markdown
Member Author

Compatibility statement (added for the 0.3.0rc1 release audit, per RELEASING.md).

  • Surfaces: the capacity-wait stage flag; scheduler job-name attribution and deduplication; intake pending markers.
  • Oldest state read: v0.2.1 pending markers and records without the flag; job names from older submitters are attributed with documented tolerance.
  • First tick after upgrade: in-flight jobs keep their names; new submissions use the new names; existing pending markers are honored.
  • Rollback: older kernels read the markers but may not attribute new-format job names; drain pending intake first.
  • Fixtures: a prior-release run and job layout across the changed surfaces is added on the release branch for 0.3.0rc1, alongside F15.

@renmengye renmengye mentioned this pull request Oct 3, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

outerloop:review re-request the advisory review

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant