Skip to content

docs: plan the role-CLI rollout (one tool surface per agentic role) - #134

Merged
renmengye merged 3 commits into
mainfrom
docs/role-cli-plan
Aug 24, 2026
Merged

renmengye merged 3 commits into
mainfrom
docs/role-cli-plan

Conversation

@renmengye

Copy link
Copy Markdown
Member

What

Design-only. Plans the rollout of the tool/CLI paradigm proven by the launch/sleep tool (#133): one kernel-installed CLI surface per agentic role; files/schemas behind it are internal ABI.

The two wins (named per application)

  • A — crossing the sandbox/trust boundary: the kernel performs privileged acts the agent directs (GitHub writes, curated egress, cluster jobs).
  • B — interactive validation replaces parse-and-repair: judges currently emit one schema-constrained message that the role-runner parses/validates/repairs-once; a verdict CLI validates each call on the spot, so the artifact is well-formed by construction and the repair tax disappears.

Invariants (held on #132/#133, non-negotiable)

CLI never the trust boundary (kernel re-validates everything) · sandbox tools standalone + parity-tested, orchestrator-side roles import real validators · no legacy on migration (a landed verb deletes the surface it replaces, same PR) · verbs RoleSpec-gated.

Phases

  1. Author-sleep wake + flip AUTHOR_SLEEP_WAKE_READY + advertise (completes Phase A — prerequisite).
  2. submit verb (= buildout Phase B), retiring the orchestrator panel-revision loop.
  3. Judge-verdict CLI — verdict add-finding / verdict conclude; deletes parse-and-repair for migrated roles. Highest reliability leverage.
  4. The finish as verbs — pr open / pr comment / ledger update; the agent-driven finish given its concrete surface; PAT stays kernel-side.
  5. Planner propose (the phenomena-not-mechanisms rule enforced by the surface — no --mechanism field exists) + retrieve (allowlist egress), when those roles/tools land.

2 and 3 are order-independent; everything rests on Phase 0 only.

Non-goals

No flow DSL (rejected in the buildout doc — the syscall is the abstraction; extension = new verbs). No MCP/backend-specific tooling (CLI-over-Bash keeps claude and codex identical). No trust migration. Not a role-plumbing rewrite.

Docs-only; suite + gate green. No secrets / no large files — confirmed.

The launch/sleep tool (#133) generalizes: the agent-facing surface for every
kernel interaction is a kernel-installed CLI; files/schemas behind it are
internal ABI. This plan sequences that across roles and the orchestrator.

Two wins, named per application: (A) crossing the sandbox/trust boundary (the
kernel performs privileged acts the agent directs — GitHub writes, curated
egress); (B) interactive validation replacing the judges' parse-and-repair loop
(verdicts well-formed by construction; the repair tax deleted).

Invariants proven on #132/#133 and held everywhere: the CLI is never the trust
boundary (kernel re-validates all verbs); sandbox tools standalone with parity
tests, orchestrator-side roles import real validators; no legacy on migration
(a landed verb deletes the surface it replaces); verbs are RoleSpec-gated.

Phases: (0) the author-sleep wake + flip the interlock + advertise (completes
Phase A, prerequisite); (1) `submit` verb = buildout Phase B, retiring the
orchestrator panel-revision loop; (2) judge-verdict CLI, deleting parse-and-
repair for migrated roles; (3) the finish as verbs (pr open/comment, ledger
update; agent-driven finish realized); (4) planner `propose` (policy encoded as
validation: no --mechanism field exists) + `retrieve` (allowlist egress) when
those land. Non-goals: no DSL, no MCP dependency, no trust migration.

Buildout doc cross-referenced (Phase B = the submit verb).

@github-actions github-actions Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Round 1 — reviewed head 51560611 — reviewer hermes/gpt-5.6-terra.

terra
Advisory findings from autoresearch — the code owner decides. Reply to disagree; the autoresearch:no-review label opts this PR out.

Verdict: nothing blocking — 1 advisory note.

1 finding attached to the lines below.

Comment thread docs/design/role-cli.md
Judges deliberately have no Bash/execute, so CLI-over-Bash cannot naively apply
to Phase 2. The plan now specifies a new RoleSpec capability distinct from
can_execute: single-command allowlisted invocation (the session may run exactly
the installed verdict binary and nothing else; per-backend, like
supports_resume; a backend that cannot restrict stays message-based). Running
verdict writes only the kernel-side verdict file -- no repo mutation, no general
execution -- so the read-only posture holds in substance. Acceptance now
includes proving an attempted other command is refused.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

@github-actions github-actions Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Round 2 — reviewed head 9a329db6 — reviewer hermes/gpt-5.6-terra.

terra
Advisory findings from autoresearch — the code owner decides. Reply to disagree; the autoresearch:no-review label opts this PR out.

Verdict: nothing blocking — 1 advisory note.

1 finding attached to the lines below.

Comment thread docs/design/role-cli.md
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
@renmengye
renmengye merged commit 586c1f0 into main Aug 24, 2026
1 check passed
@renmengye
renmengye deleted the docs/role-cli-plan branch August 24, 2026 11:34
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant