Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
44 changes: 38 additions & 6 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -6,6 +6,35 @@ Versions follow [SemVer](https://semver.org).

## [Unreleased]

### Upgrading

Operator actions (everything else needs no action; details in each entry):

- Before selecting a chat-only Codex endpoint profile, install the bridge:
`outerloop harness upgrade --used`.
- Existing Hermes source-only installs: run `bash scripts/install_hermes.sh
"$REVIEW_HERMES_REPO"` (or full `outerloop init`) before Hermes sessions launch.
- Harness version overrides now need matching SHA-256 settings.
- Evals and baselines recorded by the previous kernel are measured again once
under the new cache key, with no extra charge.
- Parked authors keep the instructions they started with; judges use the new
rubric at once.
- Runs parked by the previous kernel resume after the upgrade.
- Before rolling back: consume pending rebind requests, finish capacity-parked
runs, runs with extended session limits, overridden and endpoint-routed runs,
and chat-only Codex sessions, and stop additional instances.

- A PR tests one idea. It may include the few changes that idea needs, and the
report states the effect of each change. Authors test several values in one
array launch and report the results around the chosen value; the panel may
block a single-value tuning change without them (new `landscape` finding
category). Upgrading: no action. Judges use the new rubric at once. An author
parked across the upgrade keeps the instructions it started with. Older readers
treat `landscape` as `other`. (This was listed under 0.2.1 by mistake; it
shipped after that tag.)

- Include GPU count and resolved GPU type in dispatched eval and baseline cache identity, including budget discounts. Legacy eval slots and baseline entries are cache misses. Upgrading: no action. An eval or baseline recorded by the previous kernel, in flight or finished, is measured again once under the new cache key, with no extra budget charge; pausing or draining before the upgrade does not avoid it.

- Codex sessions never get Codex's built-in sub-agent tools; multi-agent work
goes through the kernel's channels. `OUTERLOOP_CODEX_WEB_SEARCH` controls
Codex's built-in web search: `auto` (default) leaves it to Codex on native
Expand Down Expand Up @@ -86,7 +115,11 @@ Versions follow [SemVer](https://semver.org).
delivers the refusal and a retry note, resuming the same session when supported
or starting a fresh session with the run context otherwise. Meters are
preserved and capacity waits do not exhaust stuck retries.
Existing records need no migration.
Upgrading: no action; existing records need no migration. A run parked by the
previous kernel in author-sleep with no session and no pending job starts a
fresh author leg on its next wake instead of ending as a session error.
Finish capacity-parked runs before rolling back; older kernels cannot resume
them.

- Add operator `outerloop rebind <run-id> [--root <root>] [--note <text>]` to
explicitly move an existing run to its slot's current author at the next leg,
Expand Down Expand Up @@ -162,7 +195,9 @@ Versions follow [SemVer](https://semver.org).
dropped from the seal (tracked paths retain their parent content), while
admitted work and line memory survive normal endings and crashes after
scope refusals. Filtering leaves working files and the real index untouched
and logs a bounded list of dropped paths. No persisted-state format changes.
and logs a bounded list of dropped paths. Upgrading: no action; existing line
branches are left as they are, and the next snapshot drops out-of-scope paths.
Rollback is safe.

- Support separate instances on one cluster account: process-only absolute
`OUTERLOOP_ENV_FILE`, with the existing ownership/write-permission checks,
Expand Down Expand Up @@ -301,7 +336,7 @@ Versions follow [SemVer](https://semver.org).

- Hermes installs a standalone Python and venv once per pinned commit in a sibling runtime, then launches Python directly. Full `init` provisions configured Hermes judges and records their source path; `--no-install-harness` opts out.

### Upgrading
### Upgrading notes for the entries above

- No action needed; OUTERLOOP_GPU_LANES is optional.

Expand All @@ -314,7 +349,6 @@ Versions follow [SemVer](https://semver.org).

### Upgrading

- No action needed; the panel's new `landscape` category is additive, and a reader that does not know it treats it as `other`.
- No action needed; an author's report-only answer to the panel now updates the PR, and a PR whose panel clears is marked ready for review.
- No action needed; a PR held only by a base-moved blessing heals on the next tick.
- No contract change; existing contract files need no edits.
Expand Down Expand Up @@ -350,8 +384,6 @@ Versions follow [SemVer](https://semver.org).

### Changed

- A PR is one idea, not one knob: an idea may bring the few changes it needs when the report gives each change's own effect, and a larger idea touching several places is welcome.
- Authors are told to sweep a hyperparameter or size in one array launch and show the landscape around the chosen value; the panel may block a single-point tuning change that gives no such picture.
- An idea with a clear mechanism that does not yet beat the best is reported as a success and kept on the author's research line.
- The sweep rechecks ancestry for PRs held only by a base-moved blessing.
- Merges performed by the sweep are observed and confirmed in the same tick.
Expand Down
4 changes: 2 additions & 2 deletions CITATION.cff
Original file line number Diff line number Diff line change
Expand Up @@ -13,5 +13,5 @@ authors:
repository-code: "https://github.com/outerloop-science/outerloop"
url: "https://outerloop.science"
license: Apache-2.0
version: 0.2.1
date-released: "2026-09-25"
version: 0.3.0rc1
date-released: "2026-10-03"
2 changes: 1 addition & 1 deletion src/outerloop/__init__.py
Original file line number Diff line number Diff line change
@@ -1,3 +1,3 @@
"""Outerloop: autonomous research agents that improve a benchmark on your own code."""

__version__ = "0.2.1"
__version__ = "0.3.0rc1"
20 changes: 18 additions & 2 deletions src/outerloop/attempt.py
Original file line number Diff line number Diff line change
Expand Up @@ -1611,6 +1611,16 @@ def _end(result: AttemptResult, drop_refs: list[str]) -> AttemptOutcome:
drop_snapshot(ws, Snapshot(commit="", tree="", ref=ref))
return outcome

# Old jobless checkpoints can predate capacity_wait and have no native
# session. Their saved workspace/inbox is enough to start a fresh author.
sessionless_checkpoint = (
record.stage.get("phase") == "author-sleep"
and not record.resume_session_id
and not record.stage.get("afterany")
and not record.stage.get("syscall_launches")
and not stage_launch_job_ids(record)
and not record.experiment_job_id
)
# A capacity park starts fresh before the first session or without resume support.
# Other wakes NEED the author harness and saved session. Fail as a
# named ending, not a crash: the run cannot proceed and re-waking will not
Expand All @@ -1623,8 +1633,13 @@ def _end(result: AttemptResult, drop_refs: list[str]) -> AttemptOutcome:
not record.resume_session_id
and not record.author_rebind_id
and not record.stage.get("capacity_wait")
and not sessionless_checkpoint
)
or (
not getattr(harness, "supports_resume", True)
and not record.stage.get("capacity_wait")
and not sessionless_checkpoint
)
or (not getattr(harness, "supports_resume", True) and not record.stage.get("capacity_wait"))
):
return _end(
AttemptResult(
Expand Down Expand Up @@ -2923,7 +2938,8 @@ def resume_run(

base_sha = str(stage["base_sha"])
candidate_sha = str(stage["candidate_sha"])
candidate_ref = str(stage["candidate_ref"])
# v0.2.1 author-sleep parks carry no candidate_ref; the current writer uses ""
candidate_ref = str(stage.get("candidate_ref") or "")
issue_number = record.issue_number
# the run's target branch rides the stage, so a wake opens its PR against
# the branch the ORIGINAL climb selected — not the CLI's default (the wake
Expand Down
16 changes: 16 additions & 0 deletions tests/conftest.py
Original file line number Diff line number Diff line change
Expand Up @@ -82,3 +82,19 @@ def _no_result_settle(monkeypatch):
tests that want the wait set it explicitly."""
monkeypatch.setattr("outerloop.measure.RESULT_SETTLE_S", 0.0)
monkeypatch.setattr("outerloop.measure.RESULT_POLL_S", 0.0)


@pytest.fixture
def rc1_record(tmp_path):
"""Install unmodified output of the v0.2.1 record writer."""
from pathlib import Path

source = Path(__file__).parent / "fixtures/rc1_v021"

def install(kind="open"):
directory = tmp_path / "runs/one"
directory.mkdir(parents=True, exist_ok=True)
(directory / "state.json").write_bytes((source / f"{kind}.json").read_bytes())
return directory, source

return install
48 changes: 48 additions & 0 deletions tests/fixtures/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -46,3 +46,51 @@ byte; explicit claims continue to carry provenance.
pre-session-evidence shape (numeric cost and turns, no tokens or artifact
paths). The production parked-session reader preserves that cost and report,
tolerates absent new fields, and accepts new null costs without a migration.

## 0.3.0rc1 upgrade matrix (v0.2.1 → release branch)

`rc1_v021/` was produced from tag `v0.2.1`
(`21dfe925b4ac25fd8db24c82fe243ca5fd3c25b4`), not by deleting fields from
current-writer output. The producer is retained as `rc1_v021/generate.py.txt`.
The relevant sources were inspected with
`git show v0.2.1:src/outerloop/<module>.py` for `runstate`, `measure`,
`dispatch`, `launchlog`, `tick`, `brief`, and `attempt`. To reproduce:

```sh
mkdir -p /tmp/rc1-v021-src
git archive v0.2.1 src | tar -x -C /tmp/rc1-v021-src
uv run python tests/fixtures/rc1_v021/generate.py.txt /tmp/rc1-v021-src/src /tmp/rc1-fixtures
```

The script imports those archived modules in a separate interpreter. All
identities, credentials, queue replies, metrics, and file contents are synthetic.
Git commit dates are fixed; snapshot ref UUIDs may vary on regeneration.
Job scripts' output-root coordinates alone are normalized to `/fixture/state`;
tests never execute archived job scripts. The Git bundle contains real objects,
line ancestry and dispatch refs, without an environment-specific remote URL.
Tests substitute temporary-repository SHAs/refs and target coordinates where a
live Git workspace is necessary; they do not seed legacy records with current
`save_record`.

| Surface | Fixture and production | Tests and passes |
| --- | --- | --- |
| #452 records/request | `open.json`, `merged.json`, `ended.json`: archived `RunRecord` + `save_record`, with absent author-history/rebind/candidate-author fields. `rebind.json` is a **new request overlay**, since v0.2.1 did not have rebind. | `test_v021_rebind_retry`: absent request no-op; apply once with lazy history/candidate credit; second pass byte-stable; interrupt before record write or after write/before request removal, reload and retry. Ended records stay byte-identical. |
| #452 eval/launch credit | `eval-provenance/{command.txt,job.sh}`: archived `write_eval_job` (no provenance file existed). `launches.jsonl`: archived `append_submitted` (no author/commit). | `test_v021_rebind_eval_and_launch_provenance`: legacy ledger prefix preserved; new provenance credits the sealed candidate's original author; interrupted provenance replacement and append retry; repeats add neither duplicate rows nor changed provenance. |
| #449 terminal request | The three lifecycle records above, with `end-request.json` as a **new request overlay**. `pr-states.json` is synthetic GitHub response data, not a kernel persistence format; open/merged records intentionally differ only in the external PR state. | `test_v021_operator_end_retry`: absent request preserves bytes; present request ends once; interrupt after report creation/before terminal write, retry; stale writes cannot reopen; ended records ignore requests. `test_v021_pr_lifecycle_without_operator_request`: actual open/merged PR responses through current `close_if_done`, then repeat without another terminal transition. |
| #453 / #436 parks | `sessionless.json`: archived record writer, jobless `author-sleep`, empty native session ID, no `capacity_wait`. This is a synthetic record representable by the release writer, not a claim that v0.2.1 had capacity admission. `open.json` supplies the retained-session variant. | `test_v021_sessionless_author_sleep_resumes_directly`: current wake must start fresh, preserve the candidate, retry an interruption before the author leg, and reuse its pending gate on repeat. `test_v021_capacity_error_becomes_durable_park`: an actual `CapacityError` through the orchestrator retries once then parks, dedups the capacity note, and resumes without charging. `test_v021_capacity_refusal_park_resume`: current refusal/park layered onto either old record; interrupt after inbox append/before park write; retry and repeat dedup the note and preserve meters; wake starts/resumes appropriately; another gate pass does not rerun the author or charge again. |
| #444 branches | `line.bundle`: archived `_checkout_line`, `_push_line_snapshot`, `snapshot_tree` over a synthetic Git repository. The old line intentionally includes a protected-path change that the old writer allowed. | `test_v021_line_snapshot_upgrade_retry`: retain old contents/ancestry and dispatch refs; filter only new protected changes while retaining admitted work and memory; repeat leaves line tip unchanged; crash after successful push/before local acknowledgement then retry produces no extra seal. |
| #436 queue/intake | `queue.json`, `identity.json`: queue replies derived from archived launch/wake naming and `DispatchedMeasurer._job_name`; `pending/*.json`: archived `write_pending`, unsuffixed and agent-slot forms. | `test_v021_job_layout_and_pending_retry`: repeat attribution/queue ownership and queued-slot reservation without writes; terminal jobs free capacity; interruption publishing the new intake marker preserves old markers; retry and repeat produce one new marker and reserve capacity. Eval scheduler recovery also runs against the release slot below. |
| #455 eval slots | `eval-run/eval-*/{command.txt,job.sh,submitted,exit-code,stdout}` and `identity.json`: archived `DispatchedMeasurer._dispatch`, fake submit returning 101, then synthetic completed output. In-flight variant removes only simulated completion outputs. | `test_legacy_slot_is_miss_and_redispatch_is_idempotent[v021-*]`: completed and in-flight old slots miss; new submit once; crash between submit/marker write adopts live new job; repeated pending reads do not submit; new completed result reused; legacy bytes preserved. `test_legacy_eval_redispatch_preserves_run_gpu_meter` checks the resumed gate's charged meter. |
| #455 baseline | `baselines/main@aaaa….json`: archived `write_baseline_cache`. v0.2.1 **already wrote `gpus`**, but omitted `gpu_type`; its dispatched slot key omitted both GPU determinants. | `test_v021_baseline_cache_miss_retry_reuse`: old 0.1 misses, real dispatched reader measures new baseline/candidate exactly once; interrupt baseline-cache atomic replacement after results land; retry reuses those slots; second decision uses new baseline cache without submissions. |
| #430 retained instructions | `brief.txt`: archived `build_brief` + `render`, fixed timestamp and synthetic task, paired with `open.json`'s saved session. | `test_v021_parked_brief_uses_current_rubric`: resume the saved session without calling `build_brief` or redelivering old text; interrupt author before completion then retry; judge through current verifier brief/parser, including aggregation/landscape rubric; repeat on same measured candidate reuses results and gives same judgment. |

**Owner decision for #430:** explicitly tolerate previously delivered v0.2.1
instructions in parked sessions. No correction or initial brief is redelivered;
new judging uses the current rubric. This is tolerance, not a claim that changing
a fresh brief changes an existing session's history.

**Upgrade defect exposed:** the unmarked, sessionless, jobless author-sleep
fixture ended as `session-error` before reaching its author. The release fix is
limited to recognizing that checkpoint as eligible for a fresh author leg; it
adds no persisted fields. Parks with outstanding launches retain the existing
resumability guard.
Original file line number Diff line number Diff line change
@@ -0,0 +1 @@
{"value": 0.1, "seed": 7, "run": "v021", "base_sha": "aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa", "image": "/img.sif", "command": "run main", "metric": "score", "seed_env": "", "gpus": 0}
Loading
Loading