Skip to content

Hermes as an author backend, with a bounded resume replay - #439

Merged
renmengye merged 2 commits into
mainfrom
feat/hermes-author
Sep 28, 2026
Merged

renmengye merged 2 commits into
mainfrom
feat/hermes-author

Conversation

@renmengye

Copy link
Copy Markdown
Member

Hermes was limited to judging because it could not resume a session, and authors sleep and are woken with results. Resume arrived later (saved-transcript rehydration, supports_resume = True), but the author allowlists in init, attempt and the author config check were never updated, and the design doc still said Hermes could not resume.

Changes

  • Hermes is an author backend alongside Claude Code and Codex: init (selection, runtime discovery, recorded paths), attempt validation, tick preflight, author key selection, fresh and wake wiring, and absolute syscall paths. Sleep/wake, submit, the depth loop, branches and commits already used shared paths. A Hermes author always runs contained.
  • Bounded resume replay. Each wake starts a fresh Hermes invocation and replays the saved transcript (Hermes's own context compression works within one invocation, not across wakes), so a long-running author's prompt grew without bound. The replay keeps the original brief and the latest results, fills the rest with the most recent turns within OUTERLOOP_HERMES_RESUME_MAX_CHARS (default 120000), and states how many turns were omitted. The saved transcript stays complete.
  • If the brief and latest results alone exceed the budget, the run parks as blocked on configuration without consuming wake retries, visible in the tick log and run status, and resumes once the budget is raised.
  • Author and judge credentials are compared by resolved path and by value. init and preflight validate a Hermes author's model or endpoint profile, image, runtime, and the replay budget.
  • Docs updated (install guide, the stale design-doc line) with compatibility and rollback notes; legacy transcript fixture included.

Verified: full gate 2384 passed, 2 skipped; ruff, format, mypy clean. Cost reporting for Hermes stays zero because its trajectory carries no billing usage (documented).

🤖 Generated with Claude Code

Hermes was limited to judging because it could not resume a session; resume
arrived later (saved-transcript rehydration) but the author allowlists were
never updated. Hermes is now an author backend alongside Claude Code and
Codex: init, attempt validation, tick preflight, author key selection, fresh
and wake wiring, and the syscall path treat it as a peer. It always runs
contained.

Each wake starts a fresh Hermes invocation and replays the saved transcript,
so a long-running author's prompt grew without bound. The replay now keeps the
original brief and the latest results, fills the rest with the most recent
turns within OUTERLOOP_HERMES_RESUME_MAX_CHARS (default 120000), and states how
many turns were omitted; the saved transcript stays complete. If the brief and
latest results alone exceed the budget, the run parks as blocked on
configuration without consuming wake retries, and resumes once the budget is
raised.

Author and judge credentials are compared by resolved path and value, and init
and preflight validate a Hermes author's model or endpoint, image and runtime.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

@github-actions github-actions Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Round 1 — reviewed head 20670590 — reviewer summarizer:hermes/gpt-5.6-terra over coverage+credentials+deployment+general+lifecycle+prose.

terra
Advisory findings from outerloop — the code owner decides. Reply to disagree; the outerloop:no-review label opts this PR out.

Verdict: 2 blocking, 1 advisory.

2 findings attached to the lines below.

Advisory (non-blocking):

  • Do not validate the Hermes replay budget for non-Hermes setup. [deployment+general] This unconditional validation of OUTERLOOP_HERMES_RESUME_MAX_CHARS runs before the selected author is known, so a Claude or Codex deployment exits with status 2 when this Hermes-only setting is invalid even though no Hermes author resume can use it. (src/outerloop/init.py:786; high confidence)

Merged three findings: two blocking defects in blocked-resume recovery and endpoint credential preflight, plus one advisory init-validation defect. Rejected findings: none; the deployment and general reports about Hermes replay validation were merged as the same finding.

Comment thread src/outerloop/attempt.py
Comment thread src/outerloop/tick.py Outdated
…dential

- The configuration-blocked marker is cleared when a resume succeeds or the run
  ends, so a later normal sleep is not re-woken.
- Attempt and tick preflight share one helper for the author's effective
  credential, including endpoint profiles.
- The Hermes replay budget is validated only when Hermes is the author.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@renmengye renmengye added the outerloop:review re-request the advisory review label Sep 28, 2026

@github-actions github-actions Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Round 1 — reviewed head d0a5c315 — reviewer hermes/gpt-5.6-terra.

terra
Advisory findings from outerloop — the code owner decides. Reply to disagree; the outerloop:no-review label opts this PR out.

Verdict: no defects found.

@renmengye
renmengye merged commit 0c187ac into main Sep 28, 2026
5 checks passed
@renmengye
renmengye deleted the feat/hermes-author branch September 28, 2026 17:21
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

outerloop:review re-request the advisory review

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant