Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
22 changes: 22 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -20,6 +20,28 @@ Versions follow [SemVer](https://semver.org).
sessions auto-trust project config, so repository model settings can apply;
project provider settings are filtered. The external containment boundary
remains necessary and is unchanged.
- Retain private per-invocation session evidence: delivered prompt, bounded native
log, timestamps, identifiers, usage, and SHA-256 hashes beside existing captures.
Record missing logs and truncation without failing the author. Capture Codex and
Hermes tokens and price them with `OUTERLOOP_TOKEN_PRICES`; unpriced cost is
unknown, while explicit zero prices remain zero. Status exposes known spend
and unpriced session counts. Bound native discovery by depth, entries, and
elapsed time; withhold unresolved secret prefixes at truncation boundaries.
Incomplete Codex turn usage remains unpriced, and final-message reads refuse
symlinks at every path component and apply a byte cap. Redact provisioned
credential file contents before retention, leave failed Codex invocations
unpriced, and hash Hermes's delivered resume prompt. Cost totals persist in a
kernel-written, atomically replaced index outside the workspace and session
HOME. Separate readers and restarted kernels retain totals; any uncontained
invocation marks them unverified, also shown in status.
Upgrading: legacy sidecars are not imported into cost totals; pre-index history
has unknown coverage.
Upgrading: new session sidecars and stage usage fields are additive; older
records without them remain readable, with unknown evidence coverage and no
fabricated token backfill. Numeric legacy costs retain their meaning. New
unknown costs use JSON null; older kernels that coerce parked costs with
`float()` cannot resume those records without an update. Drain parked runs
before rolling back to such a kernel.

- Include GPU count and resolved GPU type in dispatched eval and baseline cache identity, including budget discounts. Legacy eval slots and baseline entries are cache misses. Upgrading: An eval dispatched by the previous kernel and still in flight at upgrade is measured again once under the new cache key (no extra budget charge). To avoid the extra run, upgrade when no evals are in flight: `touch <root>/PAUSE` (stops wakes, so no new evals are dispatched), wait until no eval jobs remain in the queue, upgrade, then `rm <root>/PAUSE` and run `outerloop start`.

Expand Down
2 changes: 1 addition & 1 deletion docs/design/architecture.md
Original file line number Diff line number Diff line change
Expand Up @@ -289,7 +289,7 @@ grep + recency + distillation until that provably fails.
| Module | Job |
| --- | --- |
| `contract` | Schema + loader for `.outerloop.yaml`, incl. the hard-coded invariants |
| `harness` | Run one agent session in a scrubbed environment (no PAT, no billing keys; per-run HOME, fresh per run; brief on stdin, never argv); capture transcript and cost (the orchestrator captures the diff); secret-scan transcripts before storage. See "Harness and context engineering" below |
| `harness` | Run one agent session in a scrubbed environment (no PAT, no billing keys; per-run HOME, fresh per run; brief on stdin, never argv); capture private session evidence and usage (the orchestrator captures the diff); redact known secrets before storage. See [session capture](../session-capture.md). See "Harness and context engineering" below |
| `orchestrator` | Tick logic: sentinel, lease, task selection, session dispatch, state sync |
| `compute` | sbatch/squeue submit-and-poll behind one interface |
| `github` | Bot auth and push (orchestrator-side, after sessions end), PR/issue ops |
Expand Down
102 changes: 102 additions & 0 deletions docs/session-capture.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,102 @@
# Session capture

Each harness invocation, including a resume, retains evidence beside the existing
stdout capture in the run directory (`state/runs/<run-id>/`). The files are outside
`ws/` and `ws-home/`, are not mounted writable in the author's container, and
survive the default 24-hour workspace cleanup. They follow the run record's
retention policy; capture adds no separate expiry. Uncontained sessions still
have the same-user filesystem access described in the threat model. Advisory
callers without a run record retain files beside their supplied workspace; their
caller controls that directory's lifetime.

These private artifacts must never be committed, attached to PRs, or published
(SECURITY.md). Capture applies the same **known-secret redaction** as stdout,
including issued application tokens and pre-launch snapshots of provisioned credential
files (Vertex ADC, Codex auth, and mounted key contents). JSON string values,
escaped forms, and PEM key lines are included; deleting a credential during the
session does not remove it from the redaction set. This is not a general secret scanner.
Artifacts use owner-only permissions. Native reads refuse symlinks and special
files. Capture failures are logged and never change the author's outcome.

A unique `<workspace>-<backend>-<invocation-id>.session.json` records backend,
configured model, backend session ID, resume ID, start/end timestamps, turns when
known, token counters, cost, and the SHA-256 of the delivered prompt before redaction, including rehydrated Hermes
context on resume.
The invocation ID distinguishes wakes sharing a backend session ID. Artifact
entries record paths and SHA-256 hashes of the **stored, redacted bytes** for the
stdout capture, delivered prompt, and native log. Missing native logs have
`status: "missing"`; failed writes have `status: "write-error"`. Hermes also
retains the exact query pointing to its brief file, and its delivered prompt
includes the rehydrated conversation on resume.

Native log lookup uses identifiers, never newest-file selection:

| Backend | Native evidence | Usage |
| --- | --- | --- |
| Claude Code | `$HOME/.claude/projects/*/<session-id>.jsonl`; fresh invocations supply a UUID with `--session-id`, allowing lookup after timeout | Final JSON `usage.input_tokens`, `cache_read_input_tokens`, `cache_creation_input_tokens`, `output_tokens`; reported `total_cost_usd` |
| Codex | `$HOME/.codex/sessions/**/rollout-*-<thread-id>.jsonl`, using `thread.started.thread_id` | Sum `turn.completed.usage.input_tokens`, `cached_input_tokens`, `output_tokens` across this invocation |
| Hermes | `$HOME/evidence-<session-id>-<invocation-id>.json`, an explicit transcript export from the existing `run_conversation` result | Reported agent `session_prompt_tokens`, `session_completion_tokens`, `session_cache_read_tokens`, `session_cache_write_tokens`; absent usage remains unknown |

Hermes's pinned legacy CLI omits usage from its sample. A small wrapper invokes
that same CLI, retaining all returned messages (including tools), the native
Hermes session ID, and the counters before returning the original result.
It does not estimate tokens. The existing sample remains the author-result parsing channel,
not an identifier for evidence. Interrupted sessions may have no native transcript
or only partial usage; timeouts retain unknown dollars. Claude/Codex resumed native
logs can contain prior turns; their usage counters come from the current
invocation's stdout, so history is not billed again.

`OUTERLOOP_NATIVE_LOG_MAX_BYTES` caps the stored native log (default 33554432,
32 MiB). A cutoff sets `truncated: true` and records `cap_bytes`; truncated logs
may no longer be valid JSON/JSONL. Redaction looks beyond the cutoff to avoid
leaking a secret split by the boundary. An invalid setting uses the default;
zero retains an empty file marked truncated when the source is nonempty.
Unresolved secret prefixes at the lookahead boundary are withheld after replacing complete secrets.
Native discovery refuses directory symlinks and stops at 8 directory levels,
10,000 entries, or 0.25 seconds of elapsed time (checked between filesystem
operations). Reaching a limit records `status: "unavailable"` with
`reason: "discovery-limit"`.

`SessionResult.tokens` is a mapping of reported counters; missing keys mean
unknown. Normalized `input_tokens` includes cache reads and cache writes.
`cost_usd` retains its existing name and numeric values, with `None`/JSON null
for unknown cost. Claude reports dollars directly. For Codex/Hermes, operators
can supply a JSON model price table through `OUTERLOOP_TOKEN_PRICES`. Rates are
USD per million tokens, keyed by the exact configured model:

```sh
export OUTERLOOP_TOKEN_PRICES='{"example-model":{"input_tokens":2,"cached_input_tokens":0.5,"output_tokens":10,"cache_creation_input_tokens":2.5}}'
```

Uncached input is total input minus cache reads and cache writes. Each nonzero
bucket needs a finite, nonnegative rate. Missing model/rates, malformed prices,
or absent input/output usage leave dollars unknown. Codex also leaves invocation
usage and dollars unknown if any completed turn lacks required counters or
reports a different set of counters from the other turns, or the invocation fails. Explicit zero rates support
free/self-hosted models. This is recorded usage pricing, not invoice reconciliation;
existing budget enforcement is unchanged.

Parked records retain tokens and their session record path. Legacy numeric costs
remain readable; absent usage stays unknown and is not backfilled. Status totals
come exclusively from the fixed `session-index.json` in the run directory under
the state root, never from a glob or files in the workspace or session HOME.
Only the kernel writes this index, serializing updates and atomically replacing
it after retaining each session record. Separate status processes and restarted
kernels read the same persisted totals. Containment keeps this directory outside
the author's mounts; owner-only permissions alone are not the trust boundary.
Each record and index entry carries `verified`, true for contained invocations.
If any indexed invocation ran uncontained, totals carry `verified: false` and
text status labels them `unverified`. Local mode shares the operator's filesystem
authority, so those persisted totals can be altered by the author.

`known_session_cost_usd` sums priced sessions, `unpriced_sessions` counts unknown
costs, and `session_cost_usd` is null if any captured session is unpriced or no
indexed records exist. Missing or unreadable indexes report unknown totals and
`verified: false`; status never reconstructs them from sidecars. Pre-upgrade
sidecars are not imported, so history before the index has unknown coverage.
Do not interpret the known subtotal as total historical spend.

Validate fresh and resumed sessions with each pinned CLI on a real deployment,
including timeout, native paths, actual usage fields, Hermes wrapper imports and
container read-only bind, model aliases in the price table, and post-cleanup
artifact access. Synthetic tests do not establish provider invoice parity.
14 changes: 4 additions & 10 deletions src/outerloop/attempt.py
Original file line number Diff line number Diff line change
Expand Up @@ -66,7 +66,6 @@
ClaudeModelUnset,
Harness,
ResumeContextBlocked,
SessionResult,
default_binary,
default_claude_model,
redact,
Expand Down Expand Up @@ -128,6 +127,7 @@
from outerloop.runstate import (
run_dir as run_dir_of,
)
from outerloop.session_evidence import restore_session
from outerloop.syscall import (
MAX_ARTIFACT_BYTES,
MAX_REPLY_CHARS,
Expand Down Expand Up @@ -738,6 +738,8 @@ def _park_run(
)[:MAX_CLAIM_CHARS],
"session_cost_usd": parked.session.cost_usd if parked.session else 0.0,
"session_turns": parked.session.num_turns if parked.session else 0,
"session_tokens": parked.session.tokens if parked.session else {},
"session_record_path": parked.session.session_record_path if parked.session else "",
}
if parked.submitted:
# a SUBMITTED candidate park (buildout Phase B): the wake delivers the
Expand Down Expand Up @@ -2903,15 +2905,7 @@ def resume_run(

# rebuild the session from what the park saved: the (redacted) write-up and
# its real spend, so the report shows true cost/turns. It is never re-run.
session = SessionResult(
stop_reason="resumed",
is_error=False,
cost_usd=float(stage.get("session_cost_usd", 0.0)), # type: ignore[arg-type]
num_turns=int(stage.get("session_turns", 0)), # type: ignore[call-overload]
session_id=record.resume_session_id,
final_text=str(stage.get("report", "")),
transcript_path="",
)
session = restore_session(stage, record.resume_session_id)
try:
result = resume_attempt(
contract,
Expand Down
Loading
Loading