spec: v0.20 — The Agentic Turn - #53
Conversation
| │ ├── messages.jsonl # append-only event log | ||
| │ ├── claims.jsonl # who owns what subgoal | ||
| │ ├── locks/<resource>.lock # advisory file locks | ||
| │ └── inbox/<agent_id>/ # per-agent unread pointer |
There was a problem hiding this comment.
🟢 Low spec.md:45
The .agent/bus/inbox/<agent_id>/ directory is listed in the layout (line 45) as a "per-agent unread pointer," but section 2.2's bus subsystem spec never defines what files belong in this directory, how unread pointers are structured, read, or written, or which CLI commands interact with them. This leaves an unimplemented component in the specified directory layout that implementers cannot build from the spec.
Consider either adding inbox semantics to section 2.2 (file format, read/write protocol, CLI commands) or removing inbox/<agent_id>/ from the directory layout if it's deferred to a future release.
🚀 Reply "fix it for me" or copy this AI Prompt for your agent:
In file spec.md around line 45:
The `.agent/bus/inbox/<agent_id>/` directory is listed in the layout (line 45) as a "per-agent unread pointer," but section 2.2's bus subsystem spec never defines what files belong in this directory, how unread pointers are structured, read, or written, or which CLI commands interact with them. This leaves an unimplemented component in the specified directory layout that implementers cannot build from the spec.
Consider either adding inbox semantics to section 2.2 (file format, read/write protocol, CLI commands) or removing `inbox/<agent_id>/` from the directory layout if it's deferred to a future release.
Retargets the draft from v0.19 to v0.20 and reconciles it with what v0.19.0 actually shipped. The original was written against a v0.18 baseline; the loop supervisor has since landed several primitives the draft proposed to build from scratch. Adds section 0.1 mapping every shipped primitive (owned worktrees, budgets, deny-path gates, approval gates, stagnation breaker, atomic checkpoints, content-free event journal, deterministic verification, autonomy tiers) to the subsystem that must now build on it rather than duplicate it. Substantive changes beyond renumbering: - Bus messages drop the inline 4KB payload for a payload_ref path. v0.19 shipped a field whitelist for the loop event journal that excludes task, prompt, command, and output by construction; two append-only journals under .agent/ with opposite privacy guarantees means the weaker one becomes the leak. Both now share one whitelist helper. - Speculative execution becomes N loop runs plus a scoreboard instead of a second worktree/budget/gate implementation. Renamed .agent/spec/ to .agent/trials/ — 'spec' already means the loop contract schema, and .agent/spec/ reads as 'the spec for .agent'. - Background actions become L1 report-only loop contracts; act.py is a scheduler and proposal store, not a second supervisor. Carries over the v0.19 caveat that the supervisor is not an operating-system sandbox. - Eval quarantine reuses the recall exclusion path v0.19.1 added for superseded lessons, so retrieval keeps one filter chain. - Migration section is v0.19.x to v0.20; the Python floor is no longer restated, since the spec must not silently raise it. Moves the document from the repo root to docs/specs/ so the root stays README/CHANGELOG/LICENSE plus entry points.
69aac89 to
5b5d769
Compare
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 5b5d769613
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
| ### Design invariants | ||
| - **Local-first, file-based.** Every new subsystem is a directory of JSONL + small Python tools. No daemons required for the MVP. File watchers are optional. | ||
| - **Append-only audit.** Every state change writes a record. Nothing is destructive except explicit `compact` commands. | ||
| - **Permissioned.** Every new tool goes through the existing `protocols/permissions.md` pre-tool-call hook. |
There was a problem hiding this comment.
Enforce permissions inside direct runtime CLI paths
The proposed agentic-stack commands and cron-invoked act.py paths do not automatically traverse the existing pre-tool-call hook: the v0.19.1 adapters generally wire post-tool/session hooks, and the standalone adapter explicitly does not enforce permissions itself. Consequently, merely adding global_write and act_apply rules leaves direct federate.py promote and act.py approve invocations outside the promised gate. Require each mutating CLI to invoke the permission/approval checker internally, or specify concrete pre-tool wiring for every entry point.
Useful? React with 👍 / 👎.
| **Execution model.** MVP runs each approach sequentially in its own loop run on the local machine — one bounded command at a time, matching the v0.19 scheduler guidance. v0.21: optional Modal/E2B adapter to fan out in parallel. | ||
|
|
||
| **Invariants.** | ||
| - Worktree lifetime and TTL are the loop supervisor's, not a second policy. |
There was a problem hiding this comment.
Add loop TTL support before delegating trial pruning
The v0.19.1 loop supervisor has no worktree TTL or bulk pruning mechanism: loop cleanup --help requires run_id [target], and cmd_cleanup only removes that explicitly selected terminal/paused run. Therefore trial prune and the nightly cleanup described later cannot delegate expiration to the current supervisor, so abandoned trial worktrees will remain indefinitely. The spec needs to add TTL/bulk cleanup to the loop supervisor or assign this lifecycle policy to the trial subsystem.
Useful? React with 👍 / 👎.
|
|
||
| **Invariants.** | ||
| - Worktree lifetime and TTL are the loop supervisor's, not a second policy. | ||
| - Losing approaches are never lost: written to `.agent/flywheel/approved-runs.jsonl` with `verdict=lost` so the flywheel learns from them too. |
There was a problem hiding this comment.
Keep losing trials out of the approved-run corpus
Automatically writing every loser to approved-runs.jsonl conflicts with the existing flywheel contract, which accepts only human-approved, redacted runs. The current exporter does not interpret verdict=lost; it still emits traces and aggregates the record into context cards and metrics, while excluding it from eval/training only when its human-review fields are non-approved. This either contaminates the approved corpus or fails to provide the claimed learning signal, so losers should use a dedicated redacted failure-mode store or the flywheel schema/exporter must explicitly support them.
Useful? React with 👍 / 👎.
| ``` | ||
|
|
||
| - Triggered by a git pre-commit hook on `.agent/memory/semantic/lessons.jsonl`. | ||
| - A failing eval flips the lesson's `recall_status` to `quarantined`. Quarantined lessons are excluded from `recall.py` until released — the same exclusion path v0.19.1 added for superseded lessons (`render_lessons.superseded_by_map`), so retrieval keeps exactly one filter chain. |
There was a problem hiding this comment.
Preserve the two-failure quarantine threshold
This says any single failing eval immediately changes recall_status to quarantined, but the invariant below requires two consecutive failures on different commits. A runner implemented from this behavior would remove an accepted lesson from recall after the first transient failure, defeating the churn mitigation stated in the same spec. Defer the status transition until the second qualifying failure.
Useful? React with 👍 / 👎.
| { | ||
| "plans.enabled": true, | ||
| "bus.enabled": true, | ||
| "evals.enabled": true, | ||
| "retriever.enabled": false, |
There was a problem hiding this comment.
Use the existing feature-toggle object shape
Existing .features.json consumers expect entries such as "mission_control": {"enabled": false} and call .get("enabled") on each entry. These boolean values—e.g. "plans.enabled": true—would instead make the shared loader call .get on a boolean, causing enabled runtime features to fail at startup. Define each subsystem as an object containing enabled (and use the established key naming convention), or update every shared feature reader as part of the migration.
Useful? React with 👍 / 👎.
Retargets the draft from v0.19 to v0.20 and reconciles it with what v0.19.0 actually shipped. The original was written against a v0.18 baseline; the loop supervisor has since landed several primitives the draft proposed to build from scratch. Adds section 0.1 mapping every shipped primitive (owned worktrees, budgets, deny-path gates, approval gates, stagnation breaker, atomic checkpoints, content-free event journal, deterministic verification, autonomy tiers) to the subsystem that must now build on it rather than duplicate it. Substantive changes beyond renumbering: - Bus messages drop the inline 4KB payload for a payload_ref path. v0.19 shipped a field whitelist for the loop event journal that excludes task, prompt, command, and output by construction; two append-only journals under .agent/ with opposite privacy guarantees means the weaker one becomes the leak. Both now share one whitelist helper. - Speculative execution becomes N loop runs plus a scoreboard instead of a second worktree/budget/gate implementation. Renamed .agent/spec/ to .agent/trials/ — 'spec' already means the loop contract schema, and .agent/spec/ reads as 'the spec for .agent'. - Background actions become L1 report-only loop contracts; act.py is a scheduler and proposal store, not a second supervisor. Carries over the v0.19 caveat that the supervisor is not an operating-system sandbox. - Eval quarantine reuses the recall exclusion path v0.19.1 added for superseded lessons, so retrieval keeps one filter chain. - Migration section is v0.19.x to v0.20; the Python floor is no longer restated, since the spec must not silently raise it. Moves the document from the repo root to docs/specs/ so the root stays README/CHANGELOG/LICENSE plus entry points.
Design spec for converting agentic-stack from a static portable brain into an active multi-agent runtime.
Retargeted from v0.19 to v0.20. This draft was written in May against a v0.18 baseline. v0.19.0 shipped in July and landed the bounded loop supervisor, which already implements several primitives the original draft proposed to build from scratch. The spec now builds on them instead.
Lives at
docs/specs/v0.20-agentic-turn.md(was a root-levelspec.md).Eight subsystems
.agent/plans/) — a fifth memory layer for intent..agent/bus/) — append-only JSONL coordination across harnesses..agent/evals/) — accepted lessons become regression-gated behavior contracts..agent/retriever/) — fastembed + sqlite-vec + BM25 with context-pack assembly.requires/providesfrontmatter, resolver, typedcall.~/.agentic-stack/global/) — promote lessons accepted in 3+ projects..agent/trials/) — N approaches scored, winner merged..agent/act/) — sandboxed proposals with human approval.Plans, retriever, skill graph, and federation are untouched by v0.19 and stand as originally specced.
What changed in this revision
Section 0.1 is new: a table mapping each v0.19 primitive (owned worktrees, budgets, deny-path gates, approval gates, stagnation breaker, atomic checkpoints, content-free event journal, deterministic verification, autonomy tiers) to the subsystem that must now build on it rather than duplicate it. Beyond that:
payloadcapped at 4 KB. v0.19 shipped the opposite rule for.agent/runtime/loops/events.jsonl— a field whitelist excluding task, prompt, command, and output by construction. Two append-only journals under.agent/with opposite guarantees means the weaker one becomes the leak, so bus messages now carrypayload_ref(a path) and both journals share one whitelist helper..agent/spec/→.agent/trials/: "spec" already means the loop contract schema, and.agent/spec/reads as "the spec for.agent".act.pyis a scheduler and proposal store, not a second supervisor. Carries over the v0.19 caveat that the supervisor bounds and audits child processes but is not an operating-system sandbox.agentic-stack upgrade, with loop contracts and runtime state untouched. The Python floor is no longer restated so the spec can't silently raise it.Design invariants are unchanged: local-first, append-only, permissioned through
protocols/permissions.md, cross-platform, and backward compatible behind.features.jsonflags.Still a design surface, not an implementation. The 6-week rollout in §9 decomposes into per-phase PR series.
Note
Add v0.20 'The Agentic Turn' specification document
Adds a 672-line spec at v0.20-agentic-turn.md defining the target architecture and interfaces for the v0.20 agentic stack release.
.agent/directory layout and eight new subsystems: plans, bus, evals, retriever, skill graph, federation, trials, and actMacroscope summarized 5b5d769.