Skip to content

docs(autonomy): ignition Phases 2-3 DONE — 7/26 demonstrable bar met#810

Merged
kyle-sexton merged 1 commit into
mainfrom
docs/ignition-phases-2-3-done
Jul 21, 2026
Merged

docs(autonomy): ignition Phases 2-3 DONE — 7/26 demonstrable bar met#810
kyle-sexton merged 1 commit into
mainfrom
docs/ignition-phases-2-3-done

Conversation

@kyle-sexton

Copy link
Copy Markdown
Contributor

Tag flips with final status notes: Phase 2 (task live, warm-up incl. failure vocabulary, budget kill proven under subscription auth) and Phase 3 (genuine scheduled fire 41a51d66, PR opened-and-stopped, durable evidence row, live checkout clean — the 7/26 demonstrable bar). Phase 4 accumulation watch now DOING.

No linked issue (tracks #778 without closing it).

Related

🤖 Generated with Claude Code

https://claude.ai/code/session_01RdeVMaDzwP26bNiDjqM59Q

@kyle-sexton
kyle-sexton enabled auto-merge (squash) July 21, 2026 08:09
@claude

claude Bot commented Jul 21, 2026

Copy link
Copy Markdown

Claude finished @kyle-sexton's task in 51s —— View job


PR Review — docs(autonomy): ignition Phases 2-3 DONE

  • Gather context and read changed files
  • Review diff against CLAUDE.md guidelines
  • Post findings

Summary

Pure documentation status update: tag flips for Phase 2 (TODO → DONE) and Phase 3 (TODO → DONE) with dated status blocks, and Phase 4 advanced from TODO → DOING. No code, config, or plugin manifests changed.

Conventions ✓

  • Branch name docs/ignition-phases-2-3-done — matches <type>/<description>
  • PR title docs(autonomy): ignition Phases 2-3 DONE — 7/26 demonstrable bar met — valid Conventional Commits form ✓
  • Single commit, one file changed ✓

Content quality

The two status blocks are well-structured: they are dated, specific about what passed and what was deferred, and name durable identifiers (session ID 41a51d66, scratch PR numbers, merge SHA 213a8b63). The accumulation-window start anchor being explicitly recorded is good for future predicate audits.

One mild inconsistency to note:

The Phase 3 sanity check still reads:

run history shows ≥1 scheduled entry; join returns the complete row; model in telemetry = Sonnet; live checkout clean

Yet the Phase 3 status block defers model-in-telemetry confirmation to Phase 4 as a residual. The sanity check criterion is therefore listed as met (phase is DONE) while explicitly not fully verified. The status prose makes the residual transparent, which is the right call, but the sanity-check line could be annotated or struck to avoid future confusion about the phase gate. This is minor and could be addressed in a follow-on touch.

Pre-existing (not introduced here):

Line 264 still says "This PLAN.md commits on feat/autonomy-ignition" — that branch name is stale relative to the current branching convention. Not a blocker.

Verdict

No issues with the changes introduced. The phase flip and status blocks accurately reflect the recorded evidence, defer residuals explicitly, and follow the plan's own precedent (Phase 1 used the same pattern). Ready to merge.

@kyle-sexton
kyle-sexton merged commit afe1cf9 into main Jul 21, 2026
21 checks passed
@kyle-sexton
kyle-sexton deleted the docs/ignition-phases-2-3-done branch July 21, 2026 08:11

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 0f95e9b314

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment on lines +168 to +172
> **Status (2026-07-21): 7/26 DEMONSTRABLE BAR MET.** Genuine scheduled fire
> `scheduled-20260721T080540Z-41a51d66` (zero manual kick) claimed scratch#12, opened PR
> #13, stopped unmerged; gate SUCCESS via check-run API; live checkout clean; evidence row
> joined on the durable path. First complete four-way row: scratch#6 (human merge
> 213a8b63). Warm-up rows excluded from the accumulation window, which starts at

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Keep the 7/26 bar open until scheduled evidence completes

For the smoke context where PR #13 is still stopped unmerged, this claims the 7/26 demonstrable bar is met while the only complete four-way row cited is scratch#6, which is then excluded as a warm-up row outside the accumulation window. The Phase 3 evidence check below still waits for the scheduled PR's merge SHA, so this records readiness before the scheduled run has the complete joined row the roadmap is meant to prove.

Useful? React with 👍 / 👎.

Comment on lines +137 to +138
> billing mode). Missed-fire detector wiring deferred to Phase 4 (native run history +
> failure tracker items are the interim signal).

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Keep Phase 2 open until the gap detector exists

In the current machine-local path where the missed-fire detector is explicitly deferred, marking Phase 2 done skips one of that phase's own work items and sanity checks (seeded gap produces the tracker item). That lets the accumulation watch proceed with only native history as a passive signal, so a reboot/sleep missed hourly fire can go unfiled and the 14-day evidence window can silently miss runs.

Useful? React with 👍 / 👎.

Comment on lines +173 to +175
> 41a51d66. Residual minor checks carried to Phase 4: model-in-telemetry confirmation,
> stray `\r` in the claimed-line `item_url` (dispatched-line key is clean; trim at writer
> on next touch).

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Verify Sonnet telemetry before closing smoke

The first scheduled-fire smoke still needs to prove the run used the pinned Sonnet model; the Phase 3 sanity check below requires model in telemetry = Sonnet. Carrying that confirmation into Phase 4 while marking Phase 3 done starts the accumulation window before proving the scheduled run used the intended execution envelope, so later evidence can be counted under the wrong model assumption.

Useful? React with 👍 / 👎.

kyle-sexton added a commit that referenced this pull request Jul 21, 2026
… recording tier, routines provenance (#845)

## Summary

Absorbs the vetted ideas from the "4 Levels of Agentic Coding" video
digest (Boris Cherny "Steps of AI Adoption" walkthrough) into the plugin
estate, per the approved plan at
`docs/topics/boris-video-absorption/PLAN.md`.

- **testing 0.3.0 (minor):** `/testing:run-e2e` gains the consumer
config surface `.claude/testing/e2e.md` (`recording: video|gif|off`
default `off`; `browser_mode: headed|headless` default `headless`) with
explicit read path, per-key merge + provenance reporting, and plumbing
to the playwright executor; keys owned by new `context/e2e-config.md`
(policy-vs-mechanics boundary; prompt > config precedence). Prerequisite
hard-fail keeps STOP and adds a structured verification-environment gap
report. Drive loop subagent-isolated; presence-gated `/verify`-first
delegation (CC >=2.1.145). Evidence contract adds optional recording
tier + session artifacts. Eval #1 updated.
- **verification 0.3.2 (patch):** `/verification:confirm` delegation
wording acknowledges the enriched run-e2e; stages/verdict machinery
untouched.
- **autonomy 0.8.0 (minor):** `reference/routines.md` gains the
instruction-provenance clause — routine instruction content lives in a
version-controlled artifact, stored prompt is a thin pointer; surface
mappings are illustrative deployment-owned bindings.
- Implementers-table row added in
`docs/conventions/consumer-config-layering/README.md`.

Coordination: zero file overlap with the ignition/enforcement lanes
(#778 comment posted; ignition Phases 2–3 merged via #810). Defaults
preserve current consumer behavior.

## Related

No linked issue: this PR closes no GitHub issue. Related, not closed:

- #778 — ignition lane coordination (comment posted; zero file overlap)
- #810 — ignition Phases 2-3 (merged; separate lane)

## Verification

- markdownlint: 0 errors on all touched files; typos gate exit 0
- `check-skill.sh`: run-e2e PASS, confirm PASS (pre-existing advisory
warnings only)
- run-e2e `evals/evals.json`: valid against bundled schema (ajv,
draft2020)
- Sanity greps per PLAN phase checks: all pass
- Stage-structure invariant on confirm SKILL.md: 16 before/after

<details>
<summary>PLAN.md</summary>

# PLAN — boris-video-absorption

## Brief

Absorb the vetted ideas from the Ray Amjad "4 Levels of Agentic Coding"
video digest (research base:

`~/.work/youtube-watch/the-4-levels-of-agentic-coding-how-to-sh-XLA-sTSJ-Wc/`)
into the plugin estate. Interview-locked contracts (this session):

- Extend existing seams, no new skill: `testing:run-e2e` (driver) +
  `verification:confirm` (judge) keep their division of labor.
- run-e2e gains: subagent-isolated surface runs; native-/verify-first
delegation (presence-gated, CC ≥2.1.145); structured missing-environment
  gap report on prerequisite failure (hard-fail stays, plus actionable
  report); video/session-artifact evidence tier.
- Recording format (video | gif | off) and browser mode (headed |
  headless) are layered config knobs per
  `docs/conventions/consumer-config-layering/` — NOT hardcodes.
- `plugins/autonomy/reference/routines.md` gains a thin-pointer clause:
  routine/scheduled-task prompts on any surface point to
  version-controlled files; pasted prose non-compliant.
- `~/.claude/scheduled-tasks/` (exists, live task) proposed for dotfiles
  tracking via the dotfiles repo's add-dotfile flow.
- Deferred-with-trigger (recorded in digest slice, NO tracker issues):
  verification self-improvement loop (trigger: accumulated recorded
  runs); bounded build→review→fix loop (trigger: post-Phase-II
  enforcement rail); Claude Tag enablement (trigger: post-8/3 cutover +
  Enterprise admin check).
- Out of scope: presenter-custom harness (Percy), chat-surface wiring,
  Chronicle adoption, lifecycle reordering.
- Boris doc capture re-verified NO-DRIFT 2026-07-21; all adopted items
  trace to the doc's step 1→2/2→3 recipes or native CC features.

Success: PRs merged with skill-quality + markdownlint green; every
adopted behavior traceable to research findings; zero collision with the
primary session's lanes (ignition Phases 2–3 merged via #810; no pending
branch touches target files — verified).

## Standards grounding

- `docs/conventions/consumer-config-layering/README.md` — three layers
  (user-global / team / local overlay), additive-preferred merge; new
  config surface must declare keys in the owning plugin's docs and point
  at the convention for layering.
- `docs/CATALOG-TAXONOMY.md` + `docs/PLUGIN-PHILOSOPHY.md` —
extend-over-
  add, seam cooperation presence-gated, fresh-eyes delegation.
- Repo commit/PR conventions (conventional commits, codex thread
  resolution before merge, `--body-file -` for PR bodies).

## Plan

### Phase 1: Lane setup [DONE]

- Create worktree off `origin/main`, branch
`feat/boris-video-absorption`
  (sibling-worktree convention).
- Materialize `docs/topics/boris-video-absorption/PLAN.md` (this file) +
`design/design-resolution.md` (Tier B early-exit: no new types; one new
  config surface whose shape is resolved here — see Decisions).
- Post coordination comment on #778: secondary session lane, files
  touched, no overlap with ignition/enforcement lanes.
- **Sanity Check:** `git worktree list` shows the new worktree on the
new
  branch; `gh issue view 778` shows the comment; PLAN.md exists in the
  worktree at `docs/topics/boris-video-absorption/PLAN.md`.

### Phase 2: run-e2e enrichment [DONE]

Files: `plugins/testing/skills/run-e2e/SKILL.md`,
`plugins/testing/skills/run-e2e/context/e2e.md`,
`plugins/testing/README.md`, new
`plugins/testing/skills/run-e2e/context/e2e-config.md` (key owner doc),
`plugins/testing/skills/run-e2e/evals/evals.json` (eval #1 asserts the
prerequisite-failure behavior this phase changes — update
unconditionally), `plugins/testing/CHANGELOG.md` (Keep a Changelog
entry;
**minor** bump — new behaviors),
`docs/conventions/consumer-config-layering/README.md`
(Implementers-table
row for the new surface — the convention tracks conformance, not
assumes it).

- Declare config surface `.claude/testing/e2e.md` (folder-form name per
  convention: surface identity = whole path relative to `.claude/`).
Keys: `recording: video | gif | off` (default `off` — evidence contract
screenshots stay the floor), `browser_mode: headed | headless` (default
  headless, operator overrides for visibility). Per-key override
  semantics. Layering resolution cites the convention doc — no
  restatement.
- **Read path (reviewer gap #3 — mandatory):** SKILL.md gains explicit
  resolution instructions — anchor at repo root, read all three layers
(user-global → team → local overlay), merge per-key, report which layer
  supplied each effective value. **Plumbing:** resolved `browser_mode`
  and `recording` are passed to the executor (`/playwright:playwright`
  invocation args / gif_creator choice) — keys without plumbing are the
  documented ai-briefing deviation; not repeating it.
- **Key-ownership framing (reviewer risk #6):** e2e-config.md states the
  boundary — run-e2e owns capture POLICY (what evidence, which format,
  visibility preference); playwright owns MECHANICS (how the browser
  runs). Keys are policy inputs run-e2e resolves and passes through.
- **Precedence (operator-locked):** config keys are DEFAULTS only; an
  explicit session prompt always wins ("run this headed" overrides
  `browser_mode: headless`). e2e-config.md states this: prompt >
  local overlay > team > user-global > bundled default.
- SKILL.md: subagent-isolated run guidance (delegate the drive loop to a
  subagent; orchestrator consumes evidence paths only); /verify-first
  delegation note (presence-gated on bundled /verify, CC ≥2.1.145, with
  existing orchestrator fallback unchanged).
- Prerequisite failure path: keep hard-fail STOP; add a structured
  verification-environment gap report (what's missing, what the operator
  must provide — keys/CLIs/environments; the "what do you need to verify
  this e2e" elicitation) written to the run's evidence output.
- context/e2e.md evidence contract: add optional recording tier (video
  via playwright CLI preferred for long flows; GIF via gif_creator for
  short demos — driven by the `recording` key) + session artifacts row
(recording path, session ID, transcript pointer) in the evidence table.
- **Sanity Check:**
`rg -c "e2e-config|\.claude/testing/e2e\.md"
plugins/testing/skills/run-e2e/SKILL.md`
  ≥ 1; `rg -c "gap report" plugins/testing/skills/run-e2e/SKILL.md` ≥ 1;
  `rg -c "recording" plugins/testing/skills/run-e2e/context/e2e.md` ≥ 2;
  **read-path present:**
  `rg -c "layer" plugins/testing/skills/run-e2e/SKILL.md` ≥ 2 AND
  `rg -c "browser_mode" plugins/testing/skills/run-e2e/SKILL.md` ≥ 2
  (resolution + plumbing sites); Implementers-table row:
`rg -c "testing/e2e"
docs/conventions/consumer-config-layering/README.md`
  ≥ 1; CHANGELOG entry present; eval #1 updated
  (`rg -c "gap" plugins/testing/skills/run-e2e/evals/evals.json` ≥ 1);
  markdownlint exit 0 on changed files; `/skill-quality:check run-e2e`
  PASS.

### Phase 3: confirm delegation touch-up [DONE]

Files: `plugins/verification/skills/confirm/SKILL.md`,
`plugins/verification/README.md` (only if wording drifts),
`plugins/verification/CHANGELOG.md` (**patch** bump — wording touch-up).

- Delegation section: acknowledge enriched run-e2e (subagent isolation +
  recording/session evidence tier + gap report); /verify stays named as
  the supplementary live-run path with unchanged presence gate. No
  structural change to stages or verdict machinery.
- **Sanity Check:**
`rg -c "gap report|recording"
plugins/verification/skills/confirm/SKILL.md`
  ≥ 1; stage structure unchanged
  (`rg -c "Stage" plugins/verification/skills/confirm/SKILL.md` equal
  before/after); markdownlint exit 0.

### Phase 4: routines.md thin-pointer clause [DONE]

Files: `plugins/autonomy/reference/routines.md`,
`plugins/autonomy/CHANGELOG.md` (**minor** bump — new normative clause).

- New subsection, phrased to honor the doc's hosting-agnosticism
  (reviewer risk #7): the NORMATIVE clause is surface-agnostic — "a
routine's instruction content lives in a version-controlled, reviewable
  artifact; the stored prompt is a thin pointer to it; pasted-prose
  prompts are non-compliant (no history, invisible drift)". The
  surface-specific mappings (cloud routines → committed
  `.claude/skills/`, documented transfer path; Desktop scheduled tasks →
`~/.claude/scheduled-tasks/<task>/SKILL.md` under dotfiles) appear only
  as ILLUSTRATIVE binding examples, marked as deployment-owned bindings
  per the doc's own Hosting stance. Cite research
  (`routines-versioning-deep-dive.md`: native prompt versioning
  documented-absent). If the clause reads better in
  `wiring-vs-advisor.md` at build time, the builder flags it back to the
  main thread rather than deciding — placement objection is the known
  review risk.
- **Sanity Check:**
`rg -c "thin pointer|version-controlled"
plugins/autonomy/reference/routines.md`
  ≥ 2; agnosticism preserved:
  `rg -c "deployment-owned" plugins/autonomy/reference/routines.md`
  count increases by ≥1; markdownlint exit 0; CHANGELOG entry present.

### Phase 5: digest-slice deferred items + through-line [DONE]

Files (home slice, not repo):

`~/.work/youtube-watch/the-4-levels-of-agentic-coding-how-to-sh-XLA-sTSJ-Wc/recommendations/menu.md`,
`README.md` (same slice).

- Update menu items to final dispositions (adopted → PR refs; deferred →
  named triggers; no-go unchanged). Add through-line note: Boris doc →
  interview locks → PLAN → PRs.
- **Sanity Check:** `rg -c "deferred" menu.md` ≥ 3; each adopted item
  carries a PR/branch reference.

### Phase 6: dotfiles tracking proposal [TODO — user gate]

- Run the dotfiles repo's `add-dotfile` flow for
  `~/.claude/scheduled-tasks/` (dir exists with live
  `autonomy-demo-hourly-drain/`). Flow owns chezmoi mechanics; never
  `chezmoi apply` from agent context.
- **Sanity Check:** dotfiles repo PR/branch exists containing the
  scheduled-tasks source state, or an explicit operator decline recorded
  in PLAN.
- **User gate:** surfaces before executing (cross-repo write).

### Phase 7: verify + ship [TODO]

- Empirical verification in worktree: markdownlint sweep, typos gate,
  `/skill-quality:check` per touched skill, eval validation.
- Version bumps via Edit on the version line (biome retab gotcha).
- Conventional commits; PR(s) with PLAN.md in `<details>`; codex threads
  resolved before merge; close-out per `/planning:plan close-out`
  (slice prune + pointer).
- **Sanity Check:** CI green on PR; `gh pr view` shows mergeable; all
  phase tags [DONE].

### Phase 8: trial run (post-merge validation) [TODO]

- After plugins update, run the enriched flow end-to-end on a real UI
  change in a consumer repo: config surface resolved + reported,
  recording produced per key, gap report exercised (remove a prereq
  deliberately), evidence table carries session artifacts.
- **Sanity Check:** recording artifact exists on disk; gap report
emitted
  on induced prereq failure; operator reviews the recording.

## Blast radius

LOW-MEDIUM. Markdown skill/contract text + one new additive config
surface; zero runtime code; three plugins touched; consumers unaffected
by default (recording defaults off, browser_mode default preserves
current behavior). Cross-repo edge: dotfiles proposal (user-gated).

## Stress-test summary

Fresh-context plan-reviewer ran against live repo. 3 CONFIRMED-GAPs
folded: (1) phantom config surface — read-path resolution + executor
plumbing now mandated in P2 with sanity greps; (2) Implementers-table
row added to P2 files; (3) CHANGELOG entries added to P2/P3/P4 with
SemVer magnitudes (testing minor, verification patch, autonomy minor).
Risks addressed: key-ownership boundary framed policy-vs-mechanics in
e2e-config.md; routines.md clause re-phrased surface-agnostic with
bindings as illustrative examples; eval #1 update made unconditional.
Collision re-verified clean (no open PR overlap). Formal
/devils-advocate skipped: blast radius LOW-MEDIUM, no
infra/CI/dependency triggers.

## Execution shape

Wave A (parallel, file-disjoint, one worktree, agents never touch git):

| Phase | Surface | Basis |
|---|---|---|
| 2 run-e2e | Opus builder subagent | mechanical doc edits, largest item
|
| 3 confirm | Opus builder subagent | small disjoint edit |
| 4 routines | Opus builder subagent | small disjoint edit |

Wave B (sequential, main thread): 5 (home slice), 6 (user-gated
cross-repo), 7 (verify + ship — main thread only, empirical checks +
git).

Phase 1 precedes all. Scope fences: each builder gets an ALLOWED list =
its phase's files; FORBIDDEN = PLAN.md, other phases' files, any git
command. Cost: 3 builders vs sequential saves modest wall-clock;
fallback
= run 2→3→4 sequentially in main thread on any fence violation.

## Open questions

None blocking — all interview-locked.

## Handoff to implementation

### User-approval gates

- Phase 6 dotfiles add-dotfile run (cross-repo write).
- Any mid-flight scope change to config-surface shape.

### Execution shape ([EXEC-SHAPE] tagged)

- [EXEC-SHAPE] PLAN drafted in `.work/` first; materialized to
  `docs/topics/` inside the worktree at Phase 1 (avoids writing to the
  canonical checkout the primary session owns).
- [EXEC-SHAPE] Wave A 3-builder parallel shape with scope fences as
  above; sequential fallback documented.
- [EXEC-SHAPE] Single PR for phases 2–4 (one review surface, shared
  rationale) unless review load argues for a split at ship time.

### Mechanical work

- Commit boundaries: Phase 1 (PLAN + design-resolution), one commit per
  phase 2–4, Phase 5 lives outside the repo (no commit), Phase 7 ship.
- Verification checkpoints: markdownlint + typos + skill-quality per
  phase before its commit; empirical greps per Sanity Checks.
- Sequential fallback: any fence violation → abandon parallel, re-run
  remaining phases in main thread.

</details>

🤖 Generated with [Claude Code](https://claude.com/claude-code)

https://claude.ai/code/session_01QdCmaHa15bwUTh9sy7Rfdz

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant