Skip to content

spec: v0.20 — The Agentic Turn - #53

Merged
codejunkie99 merged 1 commit into
masterfrom
spec/v0.19-agentic-turn
Aug 6, 2026
Merged

codejunkie99 merged 1 commit into
masterfrom
spec/v0.19-agentic-turn

Conversation

@codejunkie99

@codejunkie99 codejunkie99 commented May 25, 2026 •

Copy link
Copy Markdown
Owner

Design spec for converting agentic-stack from a static portable brain into an active multi-agent runtime.

Retargeted from v0.19 to v0.20. This draft was written in May against a v0.18 baseline. v0.19.0 shipped in July and landed the bounded loop supervisor, which already implements several primitives the original draft proposed to build from scratch. The spec now builds on them instead.

Lives at docs/specs/v0.20-agentic-turn.md (was a root-level spec.md).

Eight subsystems

  1. Plans layer (.agent/plans/) — a fifth memory layer for intent.
  2. Multi-agent bus (.agent/bus/) — append-only JSONL coordination across harnesses.
  3. Eval runner (.agent/evals/) — accepted lessons become regression-gated behavior contracts.
  4. Hybrid retriever (.agent/retriever/) — fastembed + sqlite-vec + BM25 with context-pack assembly.
  5. Skill graph — typed requires/provides frontmatter, resolver, typed call.
  6. Cross-project federation (~/.agentic-stack/global/) — promote lessons accepted in 3+ projects.
  7. Speculative execution (.agent/trials/) — N approaches scored, winner merged.
  8. Background autonomous actions (.agent/act/) — sandboxed proposals with human approval.

Plans, retriever, skill graph, and federation are untouched by v0.19 and stand as originally specced.

What changed in this revision

Section 0.1 is new: a table mapping each v0.19 primitive (owned worktrees, budgets, deny-path gates, approval gates, stagnation breaker, atomic checkpoints, content-free event journal, deterministic verification, autonomy tiers) to the subsystem that must now build on it rather than duplicate it. Beyond that:

  • Bus privacy rule inverted. The original carried an inline payload capped at 4 KB. v0.19 shipped the opposite rule for .agent/runtime/loops/events.jsonl — a field whitelist excluding task, prompt, command, and output by construction. Two append-only journals under .agent/ with opposite guarantees means the weaker one becomes the leak, so bus messages now carry payload_ref (a path) and both journals share one whitelist helper.
  • Speculative execution folded onto loops. A trial is now N loop runs of one contract plus a scoreboard, not a second worktree/budget/gate implementation. Renamed .agent/spec/ → .agent/trials/: "spec" already means the loop contract schema, and .agent/spec/ reads as "the spec for .agent".
  • Background actions folded onto loops. Each allowed action is an L1 report-only loop contract; act.py is a scheduler and proposal store, not a second supervisor. Carries over the v0.19 caveat that the supervisor bounds and audits child processes but is not an operating-system sandbox.
  • Eval quarantine reuses the v0.19.1 recall filter added for superseded lessons, so retrieval keeps one exclusion chain rather than two.
  • Migration is now v0.19.x → v0.20, add-only via agentic-stack upgrade, with loop contracts and runtime state untouched. The Python floor is no longer restated so the spec can't silently raise it.

Design invariants are unchanged: local-first, append-only, permissioned through protocols/permissions.md, cross-platform, and backward compatible behind .features.json flags.

Still a design surface, not an implementation. The 6-week rollout in §9 decomposes into per-phase PR series.

Note

Add v0.20 'The Agentic Turn' specification document

Adds a 672-line spec at v0.20-agentic-turn.md defining the target architecture and interfaces for the v0.20 agentic stack release.

  • Defines a new .agent/ directory layout and eight new subsystems: plans, bus, evals, retriever, skill graph, federation, trials, and act
  • Specifies data shapes (JSON/YAML), CLI surfaces, permission classes, tool schemas, and a delegation contract
  • Includes migration strategy, feature flags, rollout plan, acceptance criteria, and open questions
  • This is a documentation-only change with no modifications to executable code

Macroscope summarized 5b5d769.

Comment thread spec.md Outdated
│ ├── messages.jsonl # append-only event log
│ ├── claims.jsonl # who owns what subgoal
│ ├── locks/<resource>.lock # advisory file locks
│ └── inbox/<agent_id>/ # per-agent unread pointer

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟢 Low spec.md:45

The .agent/bus/inbox/<agent_id>/ directory is listed in the layout (line 45) as a "per-agent unread pointer," but section 2.2's bus subsystem spec never defines what files belong in this directory, how unread pointers are structured, read, or written, or which CLI commands interact with them. This leaves an unimplemented component in the specified directory layout that implementers cannot build from the spec.

Consider either adding inbox semantics to section 2.2 (file format, read/write protocol, CLI commands) or removing inbox/<agent_id>/ from the directory layout if it's deferred to a future release.

🚀 Reply "fix it for me" or copy this AI Prompt for your agent:
In file spec.md around line 45:

The `.agent/bus/inbox/<agent_id>/` directory is listed in the layout (line 45) as a "per-agent unread pointer," but section 2.2's bus subsystem spec never defines what files belong in this directory, how unread pointers are structured, read, or written, or which CLI commands interact with them. This leaves an unimplemented component in the specified directory layout that implementers cannot build from the spec.

Consider either adding inbox semantics to section 2.2 (file format, read/write protocol, CLI commands) or removing `inbox/<agent_id>/` from the directory layout if it's deferred to a future release.

Retargets the draft from v0.19 to v0.20 and reconciles it with what v0.19.0
actually shipped. The original was written against a v0.18 baseline; the loop
supervisor has since landed several primitives the draft proposed to build
from scratch.

Adds section 0.1 mapping every shipped primitive (owned worktrees, budgets,
deny-path gates, approval gates, stagnation breaker, atomic checkpoints,
content-free event journal, deterministic verification, autonomy tiers) to the
subsystem that must now build on it rather than duplicate it.

Substantive changes beyond renumbering:

- Bus messages drop the inline 4KB payload for a payload_ref path. v0.19
  shipped a field whitelist for the loop event journal that excludes task,
  prompt, command, and output by construction; two append-only journals under
  .agent/ with opposite privacy guarantees means the weaker one becomes the
  leak. Both now share one whitelist helper.
- Speculative execution becomes N loop runs plus a scoreboard instead of a
  second worktree/budget/gate implementation. Renamed .agent/spec/ to
  .agent/trials/ — 'spec' already means the loop contract schema, and
  .agent/spec/ reads as 'the spec for .agent'.
- Background actions become L1 report-only loop contracts; act.py is a
  scheduler and proposal store, not a second supervisor. Carries over the
  v0.19 caveat that the supervisor is not an operating-system sandbox.
- Eval quarantine reuses the recall exclusion path v0.19.1 added for
  superseded lessons, so retrieval keeps one filter chain.
- Migration section is v0.19.x to v0.20; the Python floor is no longer
  restated, since the spec must not silently raise it.

Moves the document from the repo root to docs/specs/ so the root stays
README/CHANGELOG/LICENSE plus entry points.
@codejunkie99
codejunkie99 force-pushed the spec/v0.19-agentic-turn branch from 69aac89 to 5b5d769 Compare August 6, 2026 23:31
@codejunkie99 codejunkie99 changed the title spec: v0.19 — The Agentic Turn spec: v0.20 — The Agentic Turn Aug 6, 2026
@codejunkie99
codejunkie99 marked this pull request as ready for review August 6, 2026 23:32
@codejunkie99
codejunkie99 merged commit 9424c58 into master Aug 6, 2026
1 check passed

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 5b5d769613

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

### Design invariants
- **Local-first, file-based.** Every new subsystem is a directory of JSONL + small Python tools. No daemons required for the MVP. File watchers are optional.
- **Append-only audit.** Every state change writes a record. Nothing is destructive except explicit `compact` commands.
- **Permissioned.** Every new tool goes through the existing `protocols/permissions.md` pre-tool-call hook.

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Enforce permissions inside direct runtime CLI paths

The proposed agentic-stack commands and cron-invoked act.py paths do not automatically traverse the existing pre-tool-call hook: the v0.19.1 adapters generally wire post-tool/session hooks, and the standalone adapter explicitly does not enforce permissions itself. Consequently, merely adding global_write and act_apply rules leaves direct federate.py promote and act.py approve invocations outside the promised gate. Require each mutating CLI to invoke the permission/approval checker internally, or specify concrete pre-tool wiring for every entry point.

Useful? React with 👍 / 👎.

**Execution model.** MVP runs each approach sequentially in its own loop run on the local machine — one bounded command at a time, matching the v0.19 scheduler guidance. v0.21: optional Modal/E2B adapter to fan out in parallel.

**Invariants.**
- Worktree lifetime and TTL are the loop supervisor's, not a second policy.

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Add loop TTL support before delegating trial pruning

The v0.19.1 loop supervisor has no worktree TTL or bulk pruning mechanism: loop cleanup --help requires run_id [target], and cmd_cleanup only removes that explicitly selected terminal/paused run. Therefore trial prune and the nightly cleanup described later cannot delegate expiration to the current supervisor, so abandoned trial worktrees will remain indefinitely. The spec needs to add TTL/bulk cleanup to the loop supervisor or assign this lifecycle policy to the trial subsystem.

Useful? React with 👍 / 👎.


**Invariants.**
- Worktree lifetime and TTL are the loop supervisor's, not a second policy.
- Losing approaches are never lost: written to `.agent/flywheel/approved-runs.jsonl` with `verdict=lost` so the flywheel learns from them too.

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Keep losing trials out of the approved-run corpus

Automatically writing every loser to approved-runs.jsonl conflicts with the existing flywheel contract, which accepts only human-approved, redacted runs. The current exporter does not interpret verdict=lost; it still emits traces and aggregates the record into context cards and metrics, while excluding it from eval/training only when its human-review fields are non-approved. This either contaminates the approved corpus or fails to provide the claimed learning signal, so losers should use a dedicated redacted failure-mode store or the flywheel schema/exporter must explicitly support them.

Useful? React with 👍 / 👎.

```

- Triggered by a git pre-commit hook on `.agent/memory/semantic/lessons.jsonl`.
- A failing eval flips the lesson's `recall_status` to `quarantined`. Quarantined lessons are excluded from `recall.py` until released — the same exclusion path v0.19.1 added for superseded lessons (`render_lessons.superseded_by_map`), so retrieval keeps exactly one filter chain.

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Preserve the two-failure quarantine threshold

This says any single failing eval immediately changes recall_status to quarantined, but the invariant below requires two consecutive failures on different commits. A runner implemented from this behavior would remove an accepted lesson from recall after the first transient failure, defeating the churn mitigation stated in the same spec. Defer the status transition until the second qualifying failure.

Useful? React with 👍 / 👎.

Comment on lines +599 to +603
{
"plans.enabled": true,
"bus.enabled": true,
"evals.enabled": true,
"retriever.enabled": false,

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Use the existing feature-toggle object shape

Existing .features.json consumers expect entries such as "mission_control": {"enabled": false} and call .get("enabled") on each entry. These boolean values—e.g. "plans.enabled": true—would instead make the shared loader call .get on a boolean, causing enabled runtime features to fail at startup. Define each subsystem as an object containing enabled (and use the established key naming convention), or update every shared feature reader as part of the migration.

Useful? React with 👍 / 👎.

joyanes97 pushed a commit to joyanes97/agentic-stack that referenced this pull request Sep 27, 2026
Retargets the draft from v0.19 to v0.20 and reconciles it with what v0.19.0
actually shipped. The original was written against a v0.18 baseline; the loop
supervisor has since landed several primitives the draft proposed to build
from scratch.

Adds section 0.1 mapping every shipped primitive (owned worktrees, budgets,
deny-path gates, approval gates, stagnation breaker, atomic checkpoints,
content-free event journal, deterministic verification, autonomy tiers) to the
subsystem that must now build on it rather than duplicate it.

Substantive changes beyond renumbering:

- Bus messages drop the inline 4KB payload for a payload_ref path. v0.19
  shipped a field whitelist for the loop event journal that excludes task,
  prompt, command, and output by construction; two append-only journals under
  .agent/ with opposite privacy guarantees means the weaker one becomes the
  leak. Both now share one whitelist helper.
- Speculative execution becomes N loop runs plus a scoreboard instead of a
  second worktree/budget/gate implementation. Renamed .agent/spec/ to
  .agent/trials/ — 'spec' already means the loop contract schema, and
  .agent/spec/ reads as 'the spec for .agent'.
- Background actions become L1 report-only loop contracts; act.py is a
  scheduler and proposal store, not a second supervisor. Carries over the
  v0.19 caveat that the supervisor is not an operating-system sandbox.
- Eval quarantine reuses the recall exclusion path v0.19.1 added for
  superseded lessons, so retrieval keeps one filter chain.
- Migration section is v0.19.x to v0.20; the Python floor is no longer
  restated, since the spec must not silently raise it.

Moves the document from the repo root to docs/specs/ so the root stays
README/CHANGELOG/LICENSE plus entry points.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant