Per-target GPU lanes from the deployment's settings - #433
Conversation
OUTERLOOP_GPU_LANES maps a target to a partition, account, GPU type and extra sbatch flags, so one repo's GPU jobs (evals, author launches, sweeps) can go to a dedicated node while others keep the fleet lane. A typed lane requests --gres=gpu:<type>:N. Extra flags may not repeat any flag the kernel sets, since sbatch lets the later one win. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
There was a problem hiding this comment.
Round 1 — reviewed head fd217dd8 — reviewer summarizer:hermes/gpt-5.6-terra over coverage+credentials+deployment+general+lifecycle+prose.
terra
Advisory findings from outerloop — the code owner decides. Reply to disagree; the outerloop:no-review label opts this PR out.
Verdict: nothing blocking — 1 advisory note.
Advisory (non-blocking):
- Suggestion: The GPU request documentation omits typed lanes. [prose] GPU jobs request GPUs per node: typed lanes use --gres=gpu:<type>:N; other lanes use --gpus-per-node=N. (
docs/install.md:497; high confidence)
One advisory documentation finding: the install guide omits the typed GPU request syntax. Rejected findings: none.
…es the code Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
|
Advisory batched in 8ee5e07: the install guide now says typed lanes request GPUs with |
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
There was a problem hiding this comment.
Round 2 — reviewed head dc6e3f30 — reviewer hermes/gpt-5.6-terra.
terra
Advisory findings from outerloop — the code owner decides. Reply to disagree; the outerloop:no-review label opts this PR out.
Verdict: 1 blocking, 0 advisory.
1 finding attached to the lines below.
One blocking defect: lane extras can increase nodes and bypass GPU ceilings.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Today every GPU job the kernel submits goes to one fleet-wide lane (
OUTERLOOP_GPU_PARTITION/OUTERLOOP_GPU_ACCOUNT) with an untyped--gpus-per-node. A deployment that has a dedicated partition for one target, needing its own account, a typed GPU request or extra scheduler flags, has no way to route that target's jobs there while other targets keep the fleet lane.What changes
OUTERLOOP_GPU_LANES: JSON mapping a target to a lane. It is validated once at startup; a mistake is one clear error naming the setting.max_concurrent_gpus,niceand the queue view are unchanged.--gres=gpu:<type>:N; without one, jobs keep--gpus-per-node=N.--name=valueand may not repeat any flag the kernel sets (account, partition, gres, gpus*, cpus*, mem*, time, qos, nice, array, dependency, begin, job-name, output, error, wrap, parsable, chdir): sbatch lets the later flag win, so a repeat would silently override the kernel's value.Tests
31 new cases: parsing, valid and invalid; forbidden flags; placement per target, including resume; eval and author-array submissions for a laned target; typed gres argv; one startup error for bad JSON; CPU jobs ignore the lane; tick preflight accepts a lane without a fleet GPU partition; byte-identical argv for targets with no lane. Full suite 2180 passed.
Compatibility
No persisted state changes. The lane config is kept out of the persisted wake recipe; wake jobs read it from the environment. The setting is optional. Upgrading: no action needed.
🤖 Generated with Claude Code