Repository navigation
docs: plan the role-CLI rollout (one tool surface per agentic role) - #134
Merged
Merged
Conversation
The launch/sleep tool (#133) generalizes: the agent-facing surface for every kernel interaction is a kernel-installed CLI; files/schemas behind it are internal ABI. This plan sequences that across roles and the orchestrator. Two wins, named per application: (A) crossing the sandbox/trust boundary (the kernel performs privileged acts the agent directs — GitHub writes, curated egress); (B) interactive validation replacing the judges' parse-and-repair loop (verdicts well-formed by construction; the repair tax deleted). Invariants proven on #132/#133 and held everywhere: the CLI is never the trust boundary (kernel re-validates all verbs); sandbox tools standalone with parity tests, orchestrator-side roles import real validators; no legacy on migration (a landed verb deletes the surface it replaces); verbs are RoleSpec-gated. Phases: (0) the author-sleep wake + flip the interlock + advertise (completes Phase A, prerequisite); (1) `submit` verb = buildout Phase B, retiring the orchestrator panel-revision loop; (2) judge-verdict CLI, deleting parse-and- repair for migrated roles; (3) the finish as verbs (pr open/comment, ledger update; agent-driven finish realized); (4) planner `propose` (policy encoded as validation: no --mechanism field exists) + `retrieve` (allowlist egress) when those land. Non-goals: no DSL, no MCP dependency, no trust migration. Buildout doc cross-referenced (Phase B = the submit verb).
There was a problem hiding this comment.
Round 1 — reviewed head 51560611 — reviewer hermes/gpt-5.6-terra.
terra
Advisory findings from autoresearch — the code owner decides. Reply to disagree; the autoresearch:no-review label opts this PR out.
Verdict: nothing blocking — 1 advisory note.
1 finding attached to the lines below.
Judges deliberately have no Bash/execute, so CLI-over-Bash cannot naively apply to Phase 2. The plan now specifies a new RoleSpec capability distinct from can_execute: single-command allowlisted invocation (the session may run exactly the installed verdict binary and nothing else; per-backend, like supports_resume; a backend that cannot restrict stays message-based). Running verdict writes only the kernel-side verdict file -- no repo mutation, no general execution -- so the read-only posture holds in substance. Acceptance now includes proving an attempted other command is refused. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
There was a problem hiding this comment.
Round 2 — reviewed head 9a329db6 — reviewer hermes/gpt-5.6-terra.
terra
Advisory findings from autoresearch — the code owner decides. Reply to disagree; the autoresearch:no-review label opts this PR out.
Verdict: nothing blocking — 1 advisory note.
1 finding attached to the lines below.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What
Design-only. Plans the rollout of the tool/CLI paradigm proven by the launch/sleep tool (#133): one kernel-installed CLI surface per agentic role; files/schemas behind it are internal ABI.
The two wins (named per application)
Invariants (held on #132/#133, non-negotiable)
CLI never the trust boundary (kernel re-validates everything) · sandbox tools standalone + parity-tested, orchestrator-side roles import real validators · no legacy on migration (a landed verb deletes the surface it replaces, same PR) · verbs RoleSpec-gated.
Phases
AUTHOR_SLEEP_WAKE_READY+ advertise (completes Phase A — prerequisite).submitverb (= buildout Phase B), retiring the orchestrator panel-revision loop.verdict add-finding/verdict conclude; deletes parse-and-repair for migrated roles. Highest reliability leverage.pr open/pr comment/ledger update; the agent-driven finish given its concrete surface; PAT stays kernel-side.propose(the phenomena-not-mechanisms rule enforced by the surface — no--mechanismfield exists) +retrieve(allowlist egress), when those roles/tools land.2 and 3 are order-independent; everything rests on Phase 0 only.
Non-goals
No flow DSL (rejected in the buildout doc — the syscall is the abstraction; extension = new verbs). No MCP/backend-specific tooling (CLI-over-Bash keeps claude and codex identical). No trust migration. Not a role-plumbing rewrite.
Docs-only; suite + gate green. No secrets / no large files — confirmed.