Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
4 changes: 4 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -118,6 +118,8 @@ Versions follow [SemVer](https://semver.org).

### Added

- Optional per-target GPU lanes route evals and author launches to deployment-specific partitions, accounts, GPU types, and sbatch flags.

- Packaged `harnesses.toml` owns Claude, Codex, and Hermes pins. `outerloop harness status` reports installed versions, paths, drift, and operator overrides; `harness upgrade [name...]` verifies versioned installations before atomically recording their paths. Successful kernel deploys upgrade only configured backends; failures retain the previous installation.

- Live, tighten-only `<root>/limits.toml` GPU and active-attempt ceilings, with global defaults and per-target sections. Scheduler-reported GPU usage covers pending and running experiments, sweeps, evaluations, and GPU-bearing sessions. Authors receive uncharged launch refusals; evaluations wait for capacity. Lowering a ceiling does not cancel jobs.
Expand All @@ -133,6 +135,8 @@ Versions follow [SemVer](https://semver.org).

### Upgrading

- No action needed; OUTERLOOP_GPU_LANES is optional.

- Upgrading: full run-ID names and legacy 60-character queue names remain readable; shortened names use a derived run key without changing run records. Intake adds `@intake-<issue>` files in the existing pending directory; legacy unsuffixed and agent-slot markers remain readable. Drain queued intake jobs from older submitters (which wrote no marker) before relying on attempt ceilings. Upgrade all kernels together; older kernels do not recognize shortened names or intake markers, so drain those jobs before rollback.

- Upgrading: the optional `stage.capacity_wait` flag tolerates missing fields; existing state records need only their target for scheduler attribution. No contract schema change or admission journal. Drain older jobs whose names omit the full run ID (and older local jobs without scheduler metadata), and upgrade all submitters before relying on ceilings. Concurrent admissions may overshoot by one batch for two simultaneous checks; no cross-node admission lock.
Expand Down
20 changes: 19 additions & 1 deletion docs/install.md
Original file line number Diff line number Diff line change
Expand Up @@ -542,6 +542,23 @@ shares the author's backend, and must name an explicit model on any other
backend); the author backend is
`OUTERLOOP_AUTHOR_BACKEND`/`OUTERLOOP_AUTHOR_MODEL`.

For a target-specific GPU lane, add a JSON mapping to the deployment's `.env`:

```bash
OUTERLOOP_GPU_LANES='{"owner/repo":{"partition":"gpu-large","account":"my-account","gpu_type":"a100","extra":["--comment=reserved"]}}'
```

This lane submits `--account=my-account --partition=gpu-large --gres=gpu:a100:N
--comment=reserved` for that target's GPU evals and author launches (including
arrays and re-measures). Other targets keep the fleet GPU lane; CPU jobs are
unchanged. Only `partition` is required; an omitted `account` uses the CPU/default
account, and omitted `gpu_type` preserves untyped per-node GPU requests. `extra`
is a list of `--name=value` flags. A flag the kernel sets itself (account,
partition, gres, gpus*, cpus*, mem*, time, qos, nice, array, dependency, begin,
job-name, output, error, wrap, parsable, chdir) is rejected, since sbatch lets the
later flag win; so are unknown keys and malformed JSON.
These are cluster settings, not target contract fields.

`OUTERLOOP_AUTHOR_OVERRIDES` optionally selects an author for individual targets
and agent slots, without changing judges or other targets:

Expand Down Expand Up @@ -580,7 +597,8 @@ so it also needs the image
cluster, evals run inside the Apptainer image at `OUTERLOOP_IMAGE` (default
`~/outerloop-images/agent-py312.sif`) in a jail that binds only the
checked-out tree — an eval that needs data must fetch it into the tree, and
GPU jobs are requested per node (`--gpus-per-node`). The tick has three
GPU jobs are requested per node: `--gpus-per-node=N`, or `--gres=gpu:<type>:N`
for a lane with a GPU type (see `OUTERLOOP_GPU_LANES`). The tick has three
scheduling knobs. `OUTERLOOP_CADENCE_MIN`, read from the `.env`, is how often the
chain ticks (minutes; default 30). Two finer ones are read from the tick's own
environment (set at launch, not the per-tick `.env`): `OUTERLOOP_MIN_TICK_MINUTES`
Expand Down
2 changes: 1 addition & 1 deletion scripts/tick_deploy.sh
Original file line number Diff line number Diff line change
Expand Up @@ -154,7 +154,7 @@ if [ -n "$ENV_TRUSTED" ]; then
OUTERLOOP_VERTEX_ADC OUTERLOOP_VERTEX_SMALL_MODEL \
OUTERLOOP_TARGET \
OUTERLOOP_GITHUB_APP_FILE OUTERLOOP_BOT_LOGIN OUTERLOOP_BOT_ALIASES \
OUTERLOOP_GPU_PARTITION OUTERLOOP_GPU_ACCOUNT \
OUTERLOOP_GPU_PARTITION OUTERLOOP_GPU_ACCOUNT OUTERLOOP_GPU_LANES \
OUTERLOOP_QOS OUTERLOOP_APPTAINER_BIN \
OUTERLOOP_IMAGE \
OUTERLOOP_PANEL OUTERLOOP_PANEL_KEY_FILE \
Expand Down
11 changes: 10 additions & 1 deletion src/outerloop/attempt.py
Original file line number Diff line number Diff line change
Expand Up @@ -914,6 +914,8 @@ def _dispatch_settings(args: argparse.Namespace) -> DispatchSettings:
gpu_partition=getattr(args, "gpu_partition", ""),
gpu_account=getattr(args, "gpu_account", ""),
seed_cache=seed_dir(Path(args.run_root), target) if target else None,
target=target,
gpu_lanes=getattr(args, "gpu_lanes", {}),
)


Expand All @@ -931,6 +933,8 @@ def with_seed(dispatch: DispatchSettings, run_root: Path, target: str) -> Dispat
"""These settings with the target's seed cache filled in from the record's
target when the CLI gave none (wake and follow-up jobs carry the run id,
not the target)."""
if isinstance(dispatch, DispatchSettings):
dispatch = dc_replace(dispatch, target=target)
# tolerant of any settings object: a backend that knows no seed (or a
# test double) is left exactly as it is
if not target or getattr(dispatch, "seed_cache", "unknown") is not None:
Expand All @@ -949,7 +953,8 @@ def _make_launcher(
the sealed snapshot (write_eval_job's copy-out handles artifacts), and a
partially-submitted batch is reaped rather than orphaned. `gpus` is the
benchmark's: an author's experiments run on the same lane as its evals."""
account, partition = dispatch.placement(gpus)
lane = dispatch.lane(gpus)
account, partition = lane.account, lane.partition

def launcher(sha: str, request: SyscallRequest) -> str:
from outerloop.dispatch import eval_job_spec, write_eval_job
Expand Down Expand Up @@ -983,6 +988,8 @@ def launcher(sha: str, request: SyscallRequest) -> str:
partition=partition,
eval_minutes=launch.minutes,
gpus=gpus,
gpu_type=lane.gpu_type,
extra=lane.extra,
nice=LAUNCH_NICE,
array=array_spec(launch),
)
Expand Down Expand Up @@ -5072,8 +5079,10 @@ def _run_id(value: str) -> str:
parser.add_argument("--author-overridden", action="store_true", help=argparse.SUPPRESS)
args = parser.parse_args()
from outerloop.author_overrides import select_override
from outerloop.gpu_lanes import gpu_lanes_from_env

try:
args.gpu_lanes = gpu_lanes_from_env()
if not args.resume and not args.author_bound:
selected = select_override(args.target, args.agent_id)
if selected:
Expand Down
1 change: 1 addition & 0 deletions src/outerloop/cli.py
Original file line number Diff line number Diff line change
Expand Up @@ -84,6 +84,7 @@
"OUTERLOOP_BOT_ALIASES",
"OUTERLOOP_GPU_PARTITION",
"OUTERLOOP_GPU_ACCOUNT",
"OUTERLOOP_GPU_LANES",
"OUTERLOOP_QOS",
"OUTERLOOP_APPTAINER_BIN",
"OUTERLOOP_IMAGE",
Expand Down
9 changes: 8 additions & 1 deletion src/outerloop/compute.py
Original file line number Diff line number Diff line change
Expand Up @@ -141,6 +141,8 @@ class JobSpec:
array: str = ""
extra: tuple[str, ...] = ()

gpu_type: str = ""

def to_argv(self) -> list[str]:
if bool(self.command) == bool(self.script):
raise ValueError("exactly one of command/script must be set")
Expand All @@ -162,7 +164,12 @@ def to_argv(self) -> list[str]:
# and Slurm submit plugins commonly classify a job by its
# per-node GRES — the per-job form has been rejected on a GPU
# partition as "CPU job setup is not valid"
argv.append(f"--gpus-per-node={self.gpus}")
# Typed --gres is also per-node.
argv.append(
f"--gres=gpu:{self.gpu_type}:{self.gpus}"
if self.gpu_type
else f"--gpus-per-node={self.gpus}"
)
if self.qos:
argv.append(f"--qos={self.qos}")
if self.nice:
Expand Down
4 changes: 4 additions & 0 deletions src/outerloop/dispatch.py
Original file line number Diff line number Diff line change
Expand Up @@ -532,6 +532,8 @@ def eval_job_spec(
gpus: int = 0,
nice: int = 0,
array: str = "",
gpu_type: str = "",
extra: tuple[str, ...] = (),
) -> JobSpec:
"""The JobSpec for one dispatched eval: the hint CLAMPED to our ceiling
plus setup slack — a contract value above EVAL_JOB_MINUTES_CEILING must
Expand Down Expand Up @@ -561,6 +563,8 @@ def eval_job_spec(
gpus=gpus,
nice=nice,
array=array,
gpu_type=gpu_type,
extra=extra,
)


Expand Down
92 changes: 92 additions & 0 deletions src/outerloop/gpu_lanes.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,92 @@
"""Deployment-owned GPU lanes; targets never choose scheduler flags."""

from __future__ import annotations

import json
import os
import re
from dataclasses import dataclass

KERNEL_FLAGS = frozenset(
{
"account",
"partition",
"gres",
"time",
"qos",
"nice",
"array",
"dependency",
"begin",
"job-name",
"output",
"error",
"wrap",
"parsable",
"chdir",
}
)


@dataclass(frozen=True)
class GpuLane:
partition: str
account: str = ""
gpu_type: str = ""
extra: tuple[str, ...] = ()


def gpu_lanes_from_env() -> dict[str, GpuLane]:
"""Read and validate once at each composition root, before submitting work."""
return parse_gpu_lanes(os.environ.get("OUTERLOOP_GPU_LANES", "{}"))


def parse_gpu_lanes(raw: str) -> dict[str, GpuLane]:
"""Reject malformed config with one setting-labelled startup error."""
try:
data = json.loads(raw)
if not isinstance(data, dict):
raise ValueError("expected an object mapping owner/repo to lanes")
lanes = {}
for target, lane in data.items():
if not re.fullmatch(r"[^/\s]+/[^/\s]+", target):
raise ValueError(f"invalid target {target!r}; expected owner/repo")
if not isinstance(lane, dict) or lane.keys() - {
"partition",
"account",
"gpu_type",
"extra",
}:
raise ValueError(
f"{target}: expected a lane object with only partition/account/gpu_type/extra"
)
if not isinstance(lane.get("partition"), str) or not lane["partition"].strip():
raise ValueError(f"{target}: partition must be a nonempty string")
for key in ("account", "gpu_type"):
if key in lane and not isinstance(lane[key], str):
raise ValueError(f"{target}: {key} must be a string")
gpu_type = lane.get("gpu_type", "")
if gpu_type and not re.fullmatch(r"[A-Za-z0-9_.-]+", gpu_type):
raise ValueError(f"{target}: gpu_type must be a single GPU type")
extra = lane.get("extra", [])
if not isinstance(extra, list):
raise ValueError(f"{target}: extra must be a list of strings")
for flag in extra:
if not isinstance(flag, str) or not re.fullmatch(
r"--[a-z][a-z0-9-]*=[^\x00\r\n]*", flag
):
raise ValueError(f"{target}: extra flags must have the form --name=value")
name = flag[2:].split("=", 1)[0]
# later flags win in sbatch, so an extra must not repeat one the kernel sets
# node, task and exclusivity counts multiply a per-node GPU request past
# what the caps charge, so they are the kernel's too
if name in KERNEL_FLAGS or name.startswith(
("gpus", "cpus", "mem", "nodes", "ntasks", "exclusive", "tres-per")
):
raise ValueError(f"{target}: kernel-owned extra flag --{name}")
lanes[target] = GpuLane(
lane["partition"], lane.get("account", ""), gpu_type, tuple(extra)
)
return lanes
except (ValueError, TypeError) as exc:
raise ValueError(f"OUTERLOOP_GPU_LANES: {exc}") from None
35 changes: 32 additions & 3 deletions src/outerloop/measure.py
Original file line number Diff line number Diff line change
Expand Up @@ -21,7 +21,7 @@
import json
import logging
import time
from dataclasses import dataclass
from dataclasses import dataclass, field
from pathlib import Path
from typing import Any

Expand All @@ -31,6 +31,7 @@
read_eval_result,
write_eval_job,
)
from outerloop.gpu_lanes import GpuLane
from outerloop.job_names import run_job_name
from outerloop.orchestrator import EvalError

Expand Down Expand Up @@ -284,6 +285,9 @@ class DispatchedMeasurer:
seed_cache: Path | None = None
qos: str = ""

gpu_type: str = ""
gpu_extra: tuple[str, ...] = ()

def _placement(self, m: Measure) -> tuple[str, str]:
if m.gpus <= 0 or not self.compute.has_lanes:
return self.account, self.partition
Expand Down Expand Up @@ -403,6 +407,8 @@ def _dispatch(self, m: Measure) -> str:
partition=partition,
eval_minutes=self.eval_minutes,
gpus=m.gpus,
gpu_type=self.gpu_type if m.gpus and self.compute.has_lanes else "",
extra=self.gpu_extra if m.gpus and self.compute.has_lanes else (),
)
from outerloop.operator_limits import run_target, state_root, submit_batch

Expand Down Expand Up @@ -549,7 +555,23 @@ class DispatchSettings:
seed_cache: Path | None = None
qos: str = ""

gpu_lanes: dict[str, GpuLane] = field(default_factory=dict)
target: str = ""

def lane(self, gpus: int) -> GpuLane:
if gpus <= 0 or not self.compute.has_lanes:
return GpuLane(self.partition, self.account)
lane = self.gpu_lanes.get(self.target)
if lane is not None:
return GpuLane(lane.partition, lane.account or self.account, lane.gpu_type, lane.extra)
account, partition = self._fleet_placement(gpus)
return GpuLane(partition, account)

def placement(self, gpus: int) -> tuple[str, str]:
lane = self.lane(gpus)
return lane.account, lane.partition

def _fleet_placement(self, gpus: int) -> tuple[str, str]:
"""(account, partition) for a job needing `gpus` GPUs. Raises when a
GPU job has no lane — a queue that can never run is worse than a
loud refusal. A backend without lanes (local compute) runs every job,
Expand All @@ -571,6 +593,11 @@ def measurer(
out; `eval_minutes` is the benchmark's contract hint (clamped in the
job spec). GPUs are per MEASURE (Measure.gpus): the measurer carries
the lane and places each measure when it dispatches it."""
lane = (
self.lane(1)
if self.target in self.gpu_lanes
else GpuLane(self.gpu_partition, self.gpu_account)
)
return DispatchedMeasurer(
compute=self.compute,
run_dir=run_dir,
Expand All @@ -581,8 +608,10 @@ def measurer(
partition=self.partition,
eval_minutes=eval_minutes,
run_tag=run_tag,
gpu_partition=self.gpu_partition,
gpu_account=self.gpu_account,
gpu_partition=lane.partition,
gpu_account=lane.account,
gpu_type=lane.gpu_type,
gpu_extra=lane.extra,
# target-wide, beside the run dirs: every attempt on one base
# shares its cached baseline measurement (Benchmark.baseline)
baseline_cache=run_dir.parent / "baselines",
Expand Down
Loading
Loading