Repository navigation
GPU job sizing as operator settings - #459
Conversation
There was a problem hiding this comment.
Round 1 — reviewed head c99f0785 — reviewer summarizer:hermes/gpt-5.6-terra over coverage+credentials+deployment+general+lifecycle+prose.
terra
Advisory findings from outerloop — the code owner decides. Reply to disagree; the outerloop:no-review label opts this PR out.
Verdict: 2 blocking, 0 advisory.
GPU sizing settings in .env never reach the tick. [credentials] The new settings are documented as deployment configuration, but they are absent from TICK_ENV_KEYS, so env_file_values drops them and the deploy script cannot export them to resident, login, or local ticks; GPU jobs keep the default sizing. (src/outerloop/cli.py:56; high confidence)
The new GPU sizing settings are not exported from the deployment .env file. [deployment+general] tick_deploy.sh exports only its fixed allowlist from the deployment .env, which omits both documented OUTERLOOP_EVAL_*_PER_GPU settings; values therefore never reach eval_job_spec, leaving dispatched GPU evals and author launches at the 8-CPU/64-GB floors. (scripts/tick_deploy.sh:166; high confidence)
Two blocking propagation gaps: src/outerloop/cli.py omits the settings from TICK_ENV_KEYS, and scripts/tick_deploy.sh omits them from its deployment-env allowlist. Rejected as duplicate: the deployment finding at scripts/tick_deploy.sh:152 is the same shell allowlist omission as the merged deployment/general finding at line 166; its evidence is retained there. Coverage, lifecycle, and prose supplied no findings.
GPU evals and author launches always requested at least 8 cores and 64 GB per GPU (a hard-coded floor). On nodes with fewer cores per GPU (a single-GPU cloud VM with 6 schedulable cores, for example) every GPU job was rejected by Slurm as an unsatisfiable node configuration, so an attempt failed at its first experiment launch.
This makes the floor two operator settings,
OUTERLOOP_EVAL_CPUS_PER_GPUandOUTERLOOP_EVAL_MEM_GB_PER_GPU, defaulting to 8 and 64 (no behavior change when unset). Invalid or non-positive values keep the default with a warning. CPU evals are untouched; an explicit larger request is still never shrunk.Docs: docs/install.md, next to the GPU lane settings. CHANGELOG entry with an
Upgrading:line (no action).Verification: new test for the settings (including invalid values); full gate (ruff check, ruff format --check, mypy, pytest).
Written and reviewed by Claude.