feat(review-governor): risk-routed reviewer budgets, series mode, delta reruns, sonnet-pinned subagents, explicit codex effort - #1
Merged
Conversation
…arrytan#2633) * feat: model taxonomy gains gpt-5.6-sol + per-host generation defaults Adds 'gpt-5.6-sol' to the model taxonomy with exact-match-only resolution (Terra/Luna/suffixed IDs deliberately fall back to generic gpt) and replaces the hardcoded 'claude' generation default with a validated HostConfig.defaultModel: codex renders the gpt profile when --model is absent, every other host keeps claude. Codex ship golden regenerated accordingly; ADDING_A_HOST documents the new field. * feat: gpt-5.6-sol bounded-scope overlay + scope-aware resolvers The Sol profile pins the explicit task as the lake: adjacent work is report-only, investigation is bounded, runs terminate on one clean verification pass, and the AskUserQuestion decision-brief format is never trimmed. The overlay wrapper grants scope-interpretation precedence while concrete workflow steps, gates, and skill-mandated re-verification loops still win. Sol-specific Completeness Principle and first-run intro copy. New SETUP_COMMAND resolver renders './setup --host <host>' for every non-claude host so generated upgrade skills reinstall their own host. * feat: setup reads the Codex model from config.toml New resolve-codex-generation-model.ts reads the top-level model from ${CODEX_HOME:-~/.codex}/config.toml, validates against the model allowlist, strips control characters from every config-derived string it surfaces, guards against non-absolute config locations, and warns on Sol near-misses. setup runs it on EVERY invocation (read-only TOML lookup) so a plain ./setup can never clobber a Sol user's rendered profile with the hardcoded fallback; --model <id> overrides for one run and prints the persistence hint. Kiro installs render the claude profile before copying (Kiro fronts Claude-family models), rewrite the baked setup command to --host kiro, and restore the resolved Codex profile after; the codex skills path honors CODEX_HOME. Static pins cover the resolver wiring, fail-closed exit, quoted argv, and the Kiro sandwich. * feat: hermetic Codex runner hardening + Sol scope-termination E2E The Codex E2E runner copies auth.json only (operator plugins, MCP servers, rules, and skills no longer leak into hermetic evals), pins CODEX_HOME to the temp dir, and supports per-run model, TOML overrides, and --ignore-user-config. New periodic E2E installs the FULL generated investigate skill on gpt-5.6-sol against a planted one-line bug with decoy TODOs: the fix must land inside the boundary (untracked files counted via git status --porcelain), decoys stay byte-identical, the regression oracle survives unweakened, nothing gets committed, all within 30 tool calls. The shared .agents tree is snapshotted and restored exactly in beforeAll; fixture commits disable gpg signing. Wired into the periodic CI matrix, paid-shard globs, eval scripts, touchfiles/E2E_TIERS (codex-sol-scope-termination), and diff-based selection. Real-file periodic-tier classification pins both codex E2Es out of the gate tier. Free-tier test proves an explicit --model overrides the host default through the real generation CLI. * chore: bump version and changelog (v1.67.2.0) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs: post-ship documentation sync for v1.67.2.0 - README: Codex skills path is CODEX_HOME-aware; state that --model overrides detection for one run only (persist via the Codex config.toml model key) - CONTRIBUTING: add the model-overlay axis to the per-host config table (per-host defaultModel, override precedence) - CLAUDE.md: eval results dir is ~/.gstack/projects/<slug>/evals/ (legacy fallback ~/.gstack-dev/evals/), matching eval-store.ts and the eval:* CLI headers Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs: post-ship documentation sync (v1.67.2.0) Sol exact-match and near-miss warning documented in README; CODEX_HOME-aware uninstall and troubleshooting paths; hermetic auth.json-only detail and the build-clobber gotcha in CLAUDE.md; eval-store location corrected in ARCHITECTURE.md; defaultModel row in the ADDING_A_HOST field reference; resolver test count corrected in the CHANGELOG entry. --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
… and 21 issues closed with receipts (garrytan#2632) * fix(plan-tune): reject never-ask on one-way ids at --write --check already ignored those prefs; --write still stored them and --stats counted them as a working NEVER_ASK. Refuse the write and count leftover on-disk prefs as INERT_ONE_WAY. Co-authored-by: Cursor <cursoragent@cursor.com> * Fix: gstack-config get returns "" with exit 0 for keys that have no default Skill preambles read configuration with VAR=$(gstack-config get <key> 2>/dev/null || echo "<default>") and that fallback only fires on a non-zero exit. lookup_default ended in a catch-all that echoed "" and returned 0, so for any key missing from the table VAR came back empty and the default written right there in the preamble was unreachable. The skill then branched on a value it never specified: "skip entirely if QUESTION_TUNING is false", reached with QUESTION_TUNING="". Four keys that skills actually read had no entry and took that path: question_tuning -> callers assume "false" repo_mode -> callers assume "unknown" team_mode -> callers assume "false" transcript_ingest_mode -> callers assume "off" Each default above is the value the call sites already substitute in their own `|| echo` fallback, so this only makes reachable what was already intended. The catch-all now returns non-zero. That is deliberately scoped to the unknown-key arm alone: keys whose default is intentionally empty still exit 0, because "" is their real answer and their callers depend on it -- cross_project_learnings ("unset triggers the first-time prompt"), redact_repo_visibility ("empty falls through to gh/glab detection"), salience_allowlist, user_slug_at_*. Making every empty answer an error would have broken those. test/gstack-config-defaults.test.ts pins the class rather than the four instances: it parses the case arms and asserts every `gstack-config get <key>` site in the tree is covered, so adding a read without a default fails CI. It also pins the exit-code contract in both directions. Verified failing against the pre-fix script, where it names exactly those four keys. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(redact): a typo'd subcommand no longer exits 0 having done nothing main() recognised exactly two subcommands and let everything else fall through to the stdin scan. On empty stdin that prints "(no findings)" and exits 0, so: $ gstack-redact install-prepush-hooks # plural typo gstack-redact scan — repo UNKNOWN (no findings) $ echo $? 0 No hook was installed, and the operator has every reason to believe the credential guard is armed. A guard that silently no-ops must never exit 0. Two smaller faults in the same dispatch, both of which lead people here: - There was no --help handler, so `gstack-redact --help` fell through to the scanner. Piping a credential to it scanned the secret and exited 3. - With no piped input and no --from-file, readInput() blocks on readSync(fd 0) until an EOF that an interactive terminal never sends. That prints nothing at all, so it reads as a hang rather than as "this is a filter, feed it". Now: --help/-h/help prints usage and exits 0; an unrecognised positional prints the offender and exits 1; a TTY with nothing piped in prints usage instead of blocking. "scan" stays accepted, because the human output header reads "gstack-redact scan — repo …" and that is what people type. Usage errors exit 1, deliberately not 2 or 3. Those mean MEDIUM and HIGH findings and callers gate dispatch on them, so a usage error exiting 2 would be read as "medium findings — prompt the user". A test pins that. Tests: 4 written failing first, then fixed. Full suite 7,722 pass / 0 fail. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(browse): one ambiguous ref no longer kills the whole annotated screenshot `snapshot -a` exits 1 with "Selector matched multiple elements" on most real pages, so /qa, /canary and /land-and-deploy silently produce reports whose screenshots do not exist. Plain `screenshot <path>` is unaffected. Refs are built as getByRole(role, {name}) and disambiguated with .nth() when role+name repeats. That disambiguation cannot fire for a node with NO accessible name: the locator degrades to getByRole(role) with no name filter, and the count driving .nth() is taken from the FILTERED aria snapshot while getByRole matches the unfiltered DOM. Measured on a live page: the tree surfaced 2 unnamed paragraphs, the DOM had 9. Landmarks (banner/main/contentinfo) and paragraphs are correctly unnamed per ARIA, so this is the common case rather than an edge case. boundingBox() then hits Playwright strict mode, and the catch allowlisted only timeout/closed/Target/Execution-context messages — so the strict-mode error was re-thrown and aborted every remaining annotation. Two changes: - `.first()` before boundingBox(), so an ambiguous ref draws a box on its first match instead of aborting. The heatmap path below has always tolerated this via a bare `catch {}`; annotate was the only path that could be killed outright. - the catch no longer re-throws on unrecognised messages. A box we cannot measure is a box we do not draw, never a reason to lose the rest of the page. Set BROWSE_DEBUG to see what was skipped. Also: `-o` passed without `-a`/`-H` was silently ignored (exit 0, no file), which reads as "screenshots are broken" rather than "you forgot a flag". It now warns and points at `browse screenshot <path>`. Verified by rebuilding both ways against the same page with 51 refs present: before — "Selector matched multiple elements", no file written after — exit 0, 229KB PNG Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(version-bump): missing or empty VERSION no longer repairs a fabricated 0.0.0.0 into package.json repair now fails with exit 2 when the VERSION file is absent or empty instead of folding to DEFAULT ("0.0.0.0") — which passed VERSION_RE and regressed package.json below where it started. classify gains an additive versionFileExists field so /ship can tell a real 0.0.0.0 from a fabricated one. Re-derived from PR garrytan#2612 under the generated-file screening rule. Fixes garrytan#2600 (repair half; the path-configurability half landed in v1.67 via garrytan#2531). Contributed by @Lockyer228 Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(memory-ingest): --probe counts post-attribution, through the same gate --bulk uses probeMode previously stat'd every walked file, so setup-gbrain gated its silent bulk ingest on pre-filter counts that the write path would never ingest (garrytan#2394). The attribution decision now lives in ONE shared gate (sessionIsAttributable — cheap-parse: cwd extraction + memoized resolveGitRemote, never a full page build) used by BOTH probeMode and preparePages, so the two stages' post-attribution counts are structurally identical. ProbeReport gains skipped_unattributed; the probe prints what it excluded and --include-unattributed restores raw counts. The parity is pinned at the prepare stage (probe post-attribution == transcripts reaching import), deliberately NOT == final written. Re-derived from PR garrytan#2612 under the generated-file screening rule; the shared-gate design and the remote memo are additions from the plan review. Fixes garrytan#2394. Contributed by @Lockyer228 Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat(browse): allow CPU and network throttling for performance measurement Adds Emulation.setCPUThrottlingRate and Network.emulateNetworkConditions to CDP_ALLOWLIST. Motivation: diagnosing a real "uploads take 1-2 minutes" report, the only machine available was a fast developer workstation. Client-side processing measured 1.4s where the user experienced minutes, so the conclusion had to be reached arithmetically rather than observed. Throttling would have let the measurement reproduce the reporter's conditions directly. Both fit the existing posture rather than widening it: - Emulation already allows setDeviceMetricsOverride, clearDeviceMetricsOverride and setUserAgentOverride, which are equally mutating and scoped to the tab. - Neither method reads page content. setCPUThrottlingRate affects only timing; emulateNetworkConditions constrains traffic rather than inspecting it, so no request bodies, headers or cookies are exposed. Both are output: 'trusted' because they return no page-derived data. scope 'tab' for both, matching the surrounding Emulation entries. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(session-update): lock pidfile records the live holder; hard TTL bounds every wedge (garrytan#2613) echo $$ inside the backgrounded subshell recorded the PARENT hook's PID — which exits immediately — so every subsequent session judged the lock stale and rm -rf'd a LIVE holder's lock, letting concurrent updaters run over each other. The pidfile now records ${BASHPID:-$(sh -c 'echo $PPID')} (macOS bash 3.2 has no BASHPID; the sh child's PPID is exactly this subshell). Staleness is now two independent detectors: PID liveness (as before, but against the real holder), and a 30-minute hard TTL on the heartbeat mtime — reclaimed regardless of kill -0, so a recycled PID or hung holder can't wedge the lock forever. The holder touches the pidfile after the pull and after setup, so a legitimately-slow run keeps itself alive. Empty and missing pidfiles are respected inside the TTL window (the mkdir→echo race) and reclaimed past it. Fixes garrytan#2613. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * chore(browse): explicit windowsHide on every Bun.spawn site + census tripwire (garrytan#2575 residual) Bun.spawn sites were structurally outside the windowsHide census (it swept child_process bindings only). The runtime was already safe — native Bun hides consoles by default and bun-polyfill.cjs defaults windowsHide !== false since garrytan#2523/garrytan#2539 — but implicit defaults are exactly what regress silently. Every Bun.spawn/spawnSync in browse/src now carries the explicit flag (harmless on unix-only sites like Xvfb/xattr/open), and a second SWEEP in windows-spawn-hide.test.ts fails CI on any new flagless Bun.spawn site. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(gbrain): brain worktree advances on the daily sync — no more silently stale brains (garrytan#2516) The daily pull refreshed only ~/.gstack itself, never the detached worktree at ~/.gstack-brain-worktree that gbrain actually indexes — so after setup the brain served stale pages forever unless setup-gbrain/sync-gbrain happened to run. brain-sync --once now advances the worktree once per 24h behind an ATTEMPT stamp (.brain-worktree-last-advance — a persistently-failing advance warns once a day, not at every skill boundary), inside the existing run lock and before any ingest step touches the worktree. The new gstack-gbrain-source-wireup --advance-only is built for the unattended cadence: git-only (no gbrain prereqs), pins every operation to the managed worktree (refuses paths that are not worktrees of the artifacts repo), refuses dirty worktrees, and never runs the force-remove recovery — a cron path must not be able to delete local changes. A static pin keeps the force-remove out. docs/gbrain-sync.md stops overclaiming the old cadence. Fixes garrytan#2516. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat(memory-ingest): honor the per-remote deny/read-only trust policy (garrytan#2392) Transcript ingest now respects the same trust store as code import — the gate existed only in gstack-gbrain-sync's runCodeImport, so memory-ingest happily ingested transcripts from deny-listed repos. preparePages filters prepared transcript pages through ONE batch policy lookup (new 'get --batch' verb on bin/gstack-gbrain-repo-policy — the script owns URL normalization; the client adds repoPolicyTierBatch, one spawn for all distinct remotes, so large corpora never pay a 10s-timeout subprocess per remote). Outcomes match code-import semantics: read-only → clean skip (skipped_policy_readonly), deny → counted refusal (skipped_policy_deny), corrupted/unreadable store → HARD ERROR before any write (state, staging, egress receipt, and import all untouched) with the recovery command named — policy corruption must never read as successful ingestion. Artifacts are never policy-filtered (their git_remote is a project slug, not a remote). Fixes garrytan#2392. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(config): repo_mode keeps its empty no-default semantics (garrytan#2611 follow-up) The ported defaults table synthesized repo_mode → "unknown", but EMPTY is load-bearing for that key: gstack-repo-mode treats any non-empty answer as a user override and skips its own repo classification — the synthesized default turned the classifier into dead code (REPO_MODE=unknown everywhere; caught by test/gstack-repo-mode.test.ts via the wave's cross-agent blame protocol). repo_mode joins the empty-is-real carve-outs (empty output, exit 0). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(pair-agent): consent before killing a healthy headless daemon The pair-agent headed switch spawned 'connect --force-restart' unconditionally — auto-killing a live headless daemon (open tabs, cookies, logins) in direct contradiction of the iron rule it sits beside ('only an explicit --force-restart may kill a live daemon'). The CLI now captures daemon liveness BEFORE ensureServer (which can itself boot a fresh daemon) and relaunches only when the user passed --force-restart to pair-agent; otherwise it prints the tab count and continues against the existing daemon. The /pair-agent skill gains a matching one-way-door consent question (template half rides the wave's template block). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(gbrain-status): MCP scoping is per-project, and project-local beats user scope hasRemoteOnlyGbrainMcp scanned EVERY project's mcpServers in ~/.claude.json, so one project's remote gbrain registration reclassified broken local engines as thin-client machine-wide. It now reads user scope plus only the cwd's nearest-ancestor project key. The precedence itself was verified empirically and hermetically (fake HOME + CLAUDE_CONFIG_DIR fixtures, claude 2.1.233): with both scopes defining gbrain, 'claude mcp get gbrain' reports Scope: Local config — PROJECT-LOCAL WINS. Both in-repo consumers assumed the opposite; brain-cache's endpoint resolution flips to nearest-ancestor-project-first, and the stale user-first pin in brain-cache-roundtrip now pins the verified precedence. (The user-first jq in the brain-sync preamble resolver gets the same swap in the template block.) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(slug): gstack-slug matches remote-slug's owner-repo canonical form (live misfile bug) Found live during this wave's CEO review: bin/gstack-slug emitted SLUG=garrytan for this garrytan/gstack worktree while remote-slug correctly gave garrytan-gstack — decisions, timeline, ceo-plans, and learnings were filing into the wrong project store (observed polluting Context Recovery with another repo's decisions). Root cause: a stray empty ~/.git directory made the walk-up crown $HOME as the outermost project root; the remote lookup ran only against that root, failed silently, and the basename fallback cached 'garrytan' sticky. NOT worktree-specific — any strong marker on a non-repo ancestor triggered it. Fix: the walk now finds the outermost ancestor whose .git actually resolves an origin remote and derives owner-repo with remote-slug's byte-identical parse; marker-only ancestors keep anchoring the basename fallback but can no longer shadow a real remote. A new cache self-heal recomputes the poisoned shape (cached == basename of a marker root while a remote-bearing repo exists below), preserving legit garrytan#2212 stickiness. Nested-repo walk-up, no-remote and non-git fallbacks, and the SLUG=/BRANCH= eval contract are unchanged, pinned by a 10-case parity suite. Store migration for pre-fix data is tracked in TODOS.md. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(brain-sync): per-record spool dir — the enqueue/drain race dies structurally Producers appended lines to .brain-queue.jsonl while the drain re-read and os.replace'd it; the in-code comment admitted a lockless append between the re-read and the replace was lost. Locks and rename-rotation designs were both reviewed and rejected (each retained a tail race); the shipped design is a maildir-style spool: one FILE per record in .brain-queue.d/ (tmp + atomic rename), the drain snapshots filenames, processes, and deletes exactly what it snapshotted. Writer and drainer never share an inode — nothing to race. Semantics: at-least-once (a crash between process and unlink re-drains; downstream content-hash dedup absorbs duplicates); retained (privacy-held) records keep their files; unparseable records are kept + warned, never destroyed. Legacy .brain-queue.jsonl migrates atomically on the next drain (crash-leftover .migrating files recovered too); status/drop-queue count both surfaces; discover-new writes spool records and advances its cursor per-record-written. The preamble's queue-depth line switches to spool count in this wave's template block. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(bin-context): native slug fallback walks up like bash gstack-slug slugFromEnvironment derived the slug from the INNERMOST repo's origin while bash gstack-slug walks to the outermost project root — nested/vendored repos split their stores across the bash/native boundary (win32 hits the native path constantly). The native fallback now ports _outermost_project_root faithfully (strong/weak markers, outermost-strong-wins, 64-depth cap, fixed-point termination) plus the full resolution order: env override → walk-up → sticky cache with the garrytan#1125 self-heal → remote get-url → basename. Twelve mirrored scenarios drive BOTH implementations against the same fixtures and pin identical slugs. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(next-version): git fallback queries the live remote, never mutates, and keeps 3-digit width The degraded path counted every remote-tracking ref on every remote — stale experiment branches and second remotes inflated version allocation, and a failed base read flipped 3-digit repos to 4-digit slots. Now: ls-remote --heads origin first (GIT_TERMINAL_PROMPT=0, 5s timeout, zero local ref mutation); on failure, local refs/remotes/origin ONLY with an explicit stale-refs warning; a failed base read zeroes at the LOCAL version file's width so a 3-digit repo allocates 0.0.1, not 0.0.1.0. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(setup): hooks register the global-install path and re-point stale ones Registering hooks from a dev worktree baked that worktree's absolute path into settings.json — deleting the worktree left a dead hook erroring on every session stop, and the presence-only dedup (list-sources | grep) could never re-point it. setup's hook paths now route through _hook_install_path (global install preferred, source dir fallback), and the new ensure-event verb on gstack-settings-hook compares the registered command payload against canonical: identical → no write, different → single atomic replacement (never zero or two registrations). The plan-tune hooks had the same stale pattern and get the same fix without re-triggering their consent prompt. Also hardened: bun 1.3.13 turns an uncaught sync fs error in bun -e into a SILENT exit 0 — the registrar's write path now catches, prints, and exits 1, so a failed update can never report fake-green. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(preamble): learnings capture is unconditional at completion (garrytan#2402) 43 of 44 learnings entries came from explicit /learn — the completion-status prose read 'if you discovered a durable project quirk... log it', which models treated as optional. The step now ALWAYS runs: review the session for durable learnings, log each one, and state 'No durable learnings this session' explicitly when the review comes up empty — an empty result, never a skipped step. Re-derived from PR garrytan#2612 under the generated-file screening rule. Fixes garrytan#2402. Contributed by @Lockyer228 Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat(scrape): untrusted-content warning on the page-fetching skills (garrytan#2441) /scrape and /skillify consumed page content with zero injection guidance — the CHANGELOG claimed coverage the skills didn't have. The warning now lives in ONE exported const (UNTRUSTED_CONTENT_WARNING in resolvers/browse.ts), embedded in the browse COMMAND_REFERENCE as before AND injected standalone into both skills via the new {{UNTRUSTED_CONTENT_WARNING}} token — single source, wording can never drift between surfaces. Re-derived from PR garrytan#2612 under the generated-file screening rule. (Structural isolation for skillify-generated code is tracked as its own TODO.) Fixes garrytan#2441. Contributed by @Lockyer228 Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(review): checklist paths resolve from the installed skill root (garrytan#2518) /review Step 2 read .claude/skills/review/checklist.md — a path relative to the TARGET repo, which only resolves in gstack's own checkout. Every checklist/greptile-triage/TODOS-format reference (six across five templates — two more than the issue named, same class) now uses the installed-root form ~/.claude/skills/gstack/review/... that the templates' other references already use. The install-root class itself (non-default install dirs) is garrytan#1882, deliberately its own PR. Fixes garrytan#2518. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(pair-agent): one-way-door consent question before a daemon relaunch (template half) The skill flow now checks daemon liveness before Step 4 and asks an explicit one-way-door question (tabs/cookies/logins are lost) before passing --force-restart — never proceeding on a vague reply. Pairs with the CLI-half commit that stopped pair-agent auto-killing live daemons. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs(codex): resume does not amortize the ~21K session prelude (garrytan#2387) Measured (garrytan#2387): every codex exec call pays Codex's session prelude, and a resumed call came in slightly ABOVE a fresh one — resume buys continuity, never token savings. The skill now says so where the resume flow lives: prefer one codex call per skill, batch questions into it. Fixes garrytan#2387. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(upgrade): fast-forward first; reset --hard only behind a proved-safe gate (garrytan#2517) /gstack-upgrade went straight to stash + reset --hard origin/main. Now it tries git pull --ff-only --autostash first (the same policy session-update's auto-upgrade uses). The destructive fallback runs unprompted ONLY when both git status --porcelain AND git rev-list origin/main..HEAD are empty — a clean tree with unpushed local commits is NOT safe, reset destroys them. Anything else requires an explicit one-way-door confirmation that lists every dirty file and unpushed commit being discarded. Fixes garrytan#2517. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(preamble): brain-sync block counts the spool queue and resolves MCP project-first Two resolver halves deferred from earlier wave commits: the queue-depth line counts .brain-queue.d/*.json spool records (plus legacy lines until the drain migrates them), and GBRAIN_MCP_ENTRY_JQ swaps its operands to nearest-ancestor-project-first — matching the empirically verified Claude Code precedence (project-local beats user scope) instead of the backwards user-first assumption. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * chore: regenerate SKILL.md docs + golden fixtures (single regen for the template block) Pure generator output for the six template/resolver commits above (learnings capture, untrusted-content warning, review paths, pair-agent consent, codex resume note, upgrade ff-only, brain-sync block) — bun run gen:skill-docs + --host codex + --host factory, with the three ship golden fixtures refreshed per the documented procedure. The three sidecar-path pins in gen-skill-docs.test.ts move to the new installed-root/$GSTACK_ROOT contract (garrytan#2518). Restores template freshness; full suite green from here. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * chore: TODOS.md — strike the six wave-fixed residuals, add two follow-ups The v1.67 adversarial-review residuals section shrinks to the one item the wave couldn't reach (iOS tap routing — needs real-device verification). New entries: skillify structural isolation (a prose warning is not a boundary for page-derived generated code) and the slug store migration (pre-fix sessions on stray-marker machines filed data under the degraded slug; post-fix reads go to the correct store, so history needs a merge/alias). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * test: align cross-cutting pins with the wave's contracts Three suites pinned pre-wave behavior: browse's gstack-config test asserted the old unknown-key ''/exit-0 shape (garrytan#2611 made it exit 1); the Windows-paths suite pinned O_APPEND enqueue atomicity (the spool design satisfies the same invariant via tmp + os.replace, one file per record — pinned in its new form); and nine carve-guard skeleton ceilings absorbed the garrytan#2402 unconditional-learnings prose (~450B per skill), bumped with measured values per the guard's own protocol. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * test: re-anchor the referenced-path scanner self-check to the gstack-rooted review refs The self-check pinned the review checklist as a class-1 alias-relative ref; garrytan#2518 moved those refs to the installed gstack root (class 2). The guard now proves the scanner sees them in their new class, so the class-2 assertion can't go vacuous. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * test: pin the wave's prose-tier behaviors (ship coverage-audit gap closure) The coverage audit found one regression-shaped gap: nothing pinned that the upgrade template's ff-only pull precedes the gated reset --hard (garrytan#2517) — a future template edit reverting to reset-first would fail nothing. Pinned: the ordering, the FF_OK gate, and the unpushed-commits check. Also pinned the two minor gaps: the {{UNTRUSTED_CONTENT_WARNING}} injection points in scrape/skillify (garrytan#2441) and brain-uninstall's spool-dir cleanup. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix: pre-landing review round — 8 auto-fixes + 8 accepted findings hardened The ship review army (4 specialists + red-team + checklist, 29 findings) produced 8 mechanical auto-fixes and 11 decisions; the accepted set: - win32 slug parity completed: lib/bin-context.ts gains the remote-first outermost walk + degraded-cache self-heal the bash side got this wave — the two implementations now agree on the stray-marker live-bug shape, pinned by shared fixtures (multi-specialist 9/10 finding). - probe honors the plan's bounded-read decision: 256KB prefix, extraction semantics mirrored from parseTranscriptJsonl so probe/prepare can never diverge on the same file (>1MB transcript test). - policy normalize parity: bash normalize() now matches canonicalizeRemote on .git/-trailing and uppercase-.GIT shapes (7-shape corpus pinned two ways) — a deny for those shapes could previously slip the transcript gate. - session-update reclaim is TOCTOU-safe (atomic mv-aside on both branches). - settings-hook: unparseable settings.json errors instead of being replaced with {}; ensure-event keys on (event, source) so matcher changes update in place — never zero or two registrations. - dot-only slug guard at both parse sites (hostile 'url = ..' can't escape projects/); enqueue tmp-file janitor (1h TTL, inside the drain lock); brain-sync .migrating never clobbered; drop-queue/status count .migrating; snapshot -o warning correct + surfaced in diff mode; version-bump test order-dependence removed; uninstall clears the advance stamp. Deferred with record: slug heal-probe cost sentinel (P3 TODO), FF_OK conflation (noted, misdiagnosis-only). 270 pass / 0 fail across the 10 touched suites. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix: adversarial round — the P0 finalize fail-safe and 12 hardened findings Three adversarial passes (Claude fresh-context, Codex chaos, Codex structured with P1 gate) on the full wave diff. Multi-source findings, all fixed: - P0: finalize_queue is now explicit-delete-only — a record is unlinked ONLY when classification proves it staged or dropped; a classifier crash, a missing class file, or a malformed pulled .brain-privacy-map.json (which previously nuked the whole snapshotted queue, remotely triggerable) now retains everything, warns, and re-drains next run. load_privacy_map treats corrupt maps as retain-all, never as empty. - next-version cannot silently drop a live claim: unreadable advertised refs get a targeted --depth=1 fetch + retry; still-unreadable claims surface as UNKNOWN warnings instead of duplicate-version silence. - session-update lock: ownership-checked EXIT trap (a TTL-reclaimed holder can no longer delete the new holder's lock) + a 5-min background heartbeat so a legitimately-slow pull/setup is never reclaimed while alive. - ensure-event collapses ALL same-(event,source) duplicates to one canonical entry; unique per-process tmp path; setup call sites surface (not swallow) the hardened refusals. - memory-ingest: --limit counts only policy-permitted pages (denied records no longer starve permitted ones); --probe applies the same policy filter as --bulk (skipped_policy_* fields on the report). - version-bump repair accepts a genuine literal 0.0.0.0 VERSION file. - slug heal restricted to the stray-.git shape — package.json-anchored wrapper roots keep their legit sticky identity (garrytan#2212 preserved). - brain-sync: idle fast path sees leftover .migrating records; unparseable spool records quarantine instead of warning forever; migration comment stops overclaiming the transition-window race. - CDP throttling justifications document override persistence (callers own restoration), pinned in the allowlist test. Deferred with record: deny retroactivity for already-ingested pages (P2 TODO, same semantics as the code-import gate); legacy-migration tail race (transition-window, requires pre-spool writers). 288 pass / 0 fail across the 10 touched suites. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * chore: regenerate SKILL.md docs + goldens (Windows-separator jq fix) Pure generator output for the brain-sync block's jq ancestor match now accepting backslash-formed Windows project keys — previously project-scoped brains were invisible on Windows while the TS scope resolvers saw them. Golden ship fixtures refreshed per the documented procedure. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix: codex verify-pass residuals — chunked cwd read, post-filter partial count, migrating depth The verify re-review passed the P1 gate (0 P1s) and left three residuals, all applied: transcriptCwdFromPrefix reads in chunks until one complete record (4MB cap) so a giant first prompt can't truncate mid-JSON and break probe/bulk parity; partial_pages derives from the FINAL prepared set instead of the whole scanned corpus; the preamble queue-depth line counts leftover .brain-queue.jsonl.migrating records like the status path does (regen + goldens included). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * chore: bump version and changelog (v1.68.0.0) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs: update project documentation for v1.68.0.0 BROWSER.md: fix the $B cdp example (positional JSON params, not --json; depth is the real CDP param) and add the new perf-throttling examples (Emulation.setCPUThrottlingRate, Network.emulateNetworkConditions) with their clear-override counterparts. USING_GBRAIN_WITH_GSTACK.md: the state-files table row for the sync queue now names the maildir-style spool dir .brain-queue.d/ that replaced .brain-queue.jsonl this release. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * test: align memory-pipeline probe pins with the garrytan#2394 stage-count contract The paid-tier E2E pinned the pre-fix contract (probe headline = raw discovered). Probe now counts post-attribution — the same gate --bulk uses — with an explicit unattributed-skip line. Adds the --include-unattributed companion pin so all 9 fixtures stay accounted for. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(next-version): batch missing-tip fetches — one bounded round trip, never a per-branch crawl The targeted-fetch retry for branches whose advertised tip has no local object ran ONE git fetch per branch (10s cap each). On a shallow clone against a busy remote that crawls the network for minutes — CI's shard deadline killed the free suite mid-file. Missing tips now collect into a single batched shallow fetch (15s cap); refs still missing after the batch (one unservable ref fails the whole transfer) get a capped per-branch retry, and anything past the cap warns as an UNKNOWN claim instead of fetching. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * test(next-version): pin the batched fetch + make the offline-contract tests hermetic Two new G2 pins: N unfetched claim branches resolve with exactly ONE fetch spawn (PATH-shimmed git counts invocations), and one unservable ref no longer poisons the batch — live claims resolve via the bounded retry while only the ghost warns UNKNOWN. The garrytan#2545 offline-contract tests now run the CLI in a local fixture repo instead of the repo's own checkout: the checkout path did a live ls-remote against the real origin (operator-network-dependent, and the CI shard-deadline hang). The online-contract test gains a succeeding gh stub, so fallback:null is asserted deterministically instead of only when the operator happens to be authed. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * test(redact-cli): derive the synthetic AWS-key fixture — no contiguous credential literal in source The CI quality gate scans every ADDED diff line with the redact engine, so the garrytan#2610 port's raw fixture literals failed the very gate they exist to test. The fixture is now assembled at runtime; the scanner still receives the identical bytes. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * test(next-version): pin the fixture's host via origin-URL sniff — kills the last environment dependence The hermetic offline-contract fixture had no origin remote, so detectHost() fell through to auth probes: a machine with glab authed passed via the gitlab path while a bare CI runner read host:unknown (offline stays false there) and failed. The fixture now pushes to a local bare origin at a path containing github.com — the URL sniff pins host:github identically everywhere, asserted explicitly in both tests, with every git call still local. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: y$un_ <forrest.sun527@gmail.com> Co-authored-by: Cursor <cursoragent@cursor.com> Co-authored-by: benjamin beres <benjamin.beres@bienpreter.com> Co-authored-by: Claude Opus 5 <noreply@anthropic.com> Co-authored-by: Ricky <ricky@kinokostudio.com.hk> Co-authored-by: Connex Client Access <paul@paulkortman.com> Co-authored-by: henbima <henbima@gmail.com>
…ation + self-healing settings.json (garrytan#2631) * fix(settings-hook): KNOWN_HOOKS identity healer — per-item ownership, mutation lock, fail-closed parse Claude Code strips the unknown _gstack_source key when it rewrites settings.json, so tag-based dedupe degraded to exact-command equality and every Conductor worktree's setup appended a fresh hook entry; deleted worktrees left dead hooks erroring on every AskUserQuestion fire. - KNOWN_HOOKS identity table (shared JS prelude, single source of truth): ownership is intrinsic and PER HOOK ITEM — basename + relpath suffix + event (+ matcher where defined). Tags never claim foreign items. - New `prune-stale [--repoint <root>] [--all]`: prune dead gstack items, re-point survivors at the stable install (tag restore from the table), exact-duplicate collapse, uninstall/no-team identity sweep. Explicit plan_tune_hooks:no is honored (dead pruned, live never re-pointed). - add-event / remove-source become item-aware: replace/remove only the owned item; a user's co-located hook in the same entry is never collateral. - Mutation safety: mkdir lock with owner token, ownership-checked release, atomic stale takeover; per-process-unique tmp + backup names; backup-on-change everywhere; fail-closed on parse failure (a corrupt settings.json is never overwritten — previously catch{} clobbered it); locked atomic rollback. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat(gstack-config): `has <key>` — key-presence provenance through STATE_DIR resolution `get` returns the DEFAULTS value for absent keys, so callers that need to know whether the USER decided something (vs inherited a default) had no correct primitive — setup's consent logic was about to grep a hardcoded ~/.gstack/config.yaml, which misclassifies under GSTACK_STATE_ROOT / GSTACK_HOME / GSTACK_STATE_DIR overrides. `has` exits 0 iff the key is literally present in the resolved config file, with the same C-locale key validation as get/set. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(setup): canonical-only hook registration, heal-first, PT_EXPLICIT consent provenance Three root causes of the phantom-AskUserQuestion-hooks class, all in the registration path: - Bug A: the Conductor auto-opt-in upgraded PT_DECISION "prompt" -> "yes" even when "prompt" was dev-setup's EXPLICIT --plan-tune-hooks=prompt pin, so every new Conductor workspace installed hooks. PT_EXPLICIT (flag/env/ config-key-presence via `gstack-config has`) now gates the auto-opt-in to the true silent fall-through. - Bug B: hook commands were baked from $SOURCE_GSTACK_DIR (`pwd -P` of the running tree — ephemeral for worktrees). Registration is now CANONICAL-ONLY via _hook_command_path (${CLAUDE_CONFIG_DIR:-$HOME/.claude}/skills/gstack); missing canonical hook = skip + log, never a baked tree path. SessionStart moves to schema-aware add-event under its identity source; whitespace paths are quoted. - Bug C: nothing ever pruned, and dead tagged entries blocked the "already installed" guards forever. Setup now heals FIRST on every run (prune-stale --repoint at the stable install), surfaces a one-line summary only when something changed, surfaces the plan_tune_hooks:no-vs-live-hooks contradiction, and --no-team tears down all three sources plus an identity sweep for untagged strays. dev-setup's no-mutation guarantee gains its stated repair exception (prune dead / re-point existing, never ADD). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(uninstall): run hook cleanup BEFORE install-root deletion + full identity sweep SETTINGS_HOOK resolves via $(dirname "$0") INSIDE the install root, but the cleanup ran after `rm -rf ~/.claude/skills/gstack` — a real global uninstall (running the installed copy) silently no-op'd and orphaned every hook entry. Tests masked it by running the uninstaller from the repo checkout. The relocated block also removes the auq-error-fallback source (registered by setup, previously never torn down) and finishes with a prune-stale --all identity sweep so untagged strays (Claude Code strips _gstack_source) go too. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * test: phantom-hooks heal coverage — incident facsimile, per-item safety, lock, canonical tripwires - gstack-settings-hook-schema-aware: 16 new cases — identity re-point (tag restore), foreign-basename rejection, mixed-entry per-item safety for add-event/remove-source/--all, prune-stale modes incl. bash-prefix + Windows-backslash + spaced-path idempotence, duplicate collapse preferring the tagged twin, plan_tune_hooks:no split, backup-on-change no-churn, fail-closed corrupt-JSON for every mutator, stale-lock takeover, fresh-foreign-lock skip, two-writer concurrency smoke, and an INCIDENT FACSIMILE replaying the exact 2026-08-17 production damage (6/3/2 entries, mixed tags, live-ephemeral Stop) healing to 2/1/1 canonical. - NEW setup-hook-canonical-paths: static tripwires — canonical-only resolver (no $SOURCE_GSTACK_DIR anywhere in it), heal-before-guards ordering, unsuppressed heal output, ${VAR:-0} counter idiom, shared-prelude concatenation at every bun call site, KNOWN_HOOKS completeness vs setup's registrations, uninstall cleanup-before-deletion ordering, defect-class warning present. - setup-plan-tune-hooks-noninteractive: PT_EXPLICIT pins + `gstack-config has` provenance + has-subcommand behavior (env-resolution, malformed keys). - auq-error-fallback-hook: registration + both-teardown wiring (previously untested). - uninstall: behavioral ordering test running the INSTALLED copy from inside the root it deletes. - setup-windows-fallback / gstack-config-key-locale: pins updated for the new HOOK_CMD shape and the third C-locale validator. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(test): banner-tripwire exec used JSON.stringify as shell quoting — vacuous pass + stray artifact JSON escaping is not shell escaping. Interpolating JSON.stringify(script) into `bash -c ${...}` left every JSON "\n" as a literal backslash-n inside shell double quotes, collapsing the extracted release-body tripwire block onto one line: `then\n` parsed as the command word `thenn`, and `>&2\nelse\n` parsed as the redirect `>&2nelsen` — so every full-suite run littered a `2nelsen` file (containing "bash: thenn: command not found") in the repo root, and the test's single not-contains assertion passed VACUOUSLY because all output had been redirected into that file. The "and it actually fires" functional check never verified anything. Fix: pass the script as an argv element (spawnSync array form) and assert both branches for real — ABORT case must print the leak message to stderr, clean case must print "banner tripwire clean" to stdout. Verified: `bun test test/binding-template-drift.test.ts` previously created the artifact deterministically; the full free suite now runs artifact-free. The other shell-interpolation sites (evidence, schema-aware concurrency, empty-find-fallthrough, branch-slug-hygiene) already use correct quoting. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * test: regression pin for legacy remove mixed-entry filtering + ownership negatives Coverage-audit iron rule: the rewritten legacy `remove` action filters per-item (pre-v1.67.2 it dropped the whole entry, destroying a user's co-located SessionStart hook) — modified existing behavior, previously untested. Also pins two ownership negatives: an owned basename+relpath under the WRONG matcher stays foreign, and prune-stale on an absent settings file exits 0 with removed 0. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix: pre-landing review fixes — review-army findings hardened Specialist review (testing, maintainability, security, performance, data-migration) findings, each verified against code before fixing: - legacy remove: preserve malformed/foreign entries (hooks absent, non-array, or pre-existing empty) — only entries THIS pass emptied are dropped - add-event: never tag a mixed entry (old gstack versions in sibling worktrees treat tags as entry-level ownership and would destroy the user's co-located items); tag only single-item entries; prune-stale drops tags from mixed entries for the same reason - prune-stale: within-entry twin collapse (two dead copies of one hook re-pointed to the same canonical command no longer double-fire); command quoting hardened via gsQuoteCmd (escapes \\ " $ backtick; gsStripWrap unescapes so identity round-trips); NUL bytes in the dedupe key replaced with a JSON.stringify key (bash silently dropped the NULs, degrading the separator; the file also read as binary to tooling) - gsIsAlive: only provable absence (ENOENT/ENOTDIR) counts as dead — EACCES/EIO/unmounted volumes no longer prune (one-way-ratchet guard) - gsWriteIfChanged: preserves the live settings.json mode across rewrites (a user-tightened 0600 carrying API keys was silently broadened to 0644); fresh files start 0600; backups rotate (keep 10) - remove-source: command-less items default to foreign (gstack only writes type:command items); single-item stray claim requires a command - rollback: pointer target must be a sibling settings.json.bak.* file - uninstall + setup --no-team + SessionStart registration: stderr stays attached — a lock give-up or fail-closed parse during TEARDOWN must be visible ("the next setup retries" does not apply after uninstall) - setup: team-mode banner no longer claims an auto-update hook when registration was skipped; heal log documents the rollback-pointer caveat; SESSION_UPDATE_CMD quoting mirrors gsQuoteCmd; lock constants named - list-sources: corrupt settings.json reports to stderr instead of silently printing nothing (setup guards must not misread corrupt as no-hooks) - tests: 10 new pins (malformed-entry preservation, mixed no-tag, twin collapse, 0600 mode, metachar escaping round-trip, backup rotation, rollback pointer refusal, held-lock uninstall warning, matcher-drift tripwire, ownership negatives) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix: red-team findings — verify-gate identity, single quoting authority, Windows paths Red-team pass over the hardened diff (several findings empirically verified by the reviewer before reporting): - KNOWN_HOOKS gains the sixth identity: gstack-verify-gate (README-documented opt-in Stop hook). A tag-stripped verify-gate entry previously survived prune-stale --all and errored at the end of EVERY turn after uninstall deleted the install root — the exact phantom-hook class this branch fixes. Uninstall also sweeps its tagged form. - add-event is now the single quoting authority: every registered command is normalized through the same gsQuoteCmd/gsStripWrap round-trip the healer uses. Pre-fix, only SessionStart got caller-side quoting — a spaced/metachar canonical root registered broken plan-tune/AUQ/timeline hooks that the very next heal rewrote (the codebase disagreed with its own registrations). - Windows: MSYS-form paths (/c/Users/...) are drive-translated for fs checks only (gsWinPath) — native bun resolved them drive-relative, so the heal judged every LIVE Windows hook dead and pruned it. The three AskUserQuestion hooks and the Stop hook now also get the mandatory 'bash ' prefix on Windows (previously only SessionStart did; extensionless bash shims otherwise hit the file-association dialog). - CANONICAL_GSTACK_ROOT falls back to $HOME/.claude/skills/gstack when a CLAUDE_CONFIG_DIR-derived root was never installed (the installer hardcodes the home path — split-brain left such users permanently hookless). - prune-stale preserves foreign entries that STARTED empty (they were silently deleted, uncounted, on every heal). - The timeline Stop registration and its list-sources guard join the zero-silent-mutations contract (stderr attached). Tests: verify-gate tag-stripped heal+sweep, started-empty preservation, add-event quoting-authority round-trip. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * chore: bump version and changelog (v1.68.1.0) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs: update project documentation for v1.68.1.0 README: document canonical-only hook registration + the prune-stale self-heal in the setup hooks section; expand the manual-uninstall note to cover every gstack hook identity, not just timeline-stop-hook. CONTRIBUTING: record PT_EXPLICIT provenance (Conductor auto-opt-in fires only on the true silent fall-through) and the heal-first repair exception in the dev-setup paragraph. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(settings-hook): fail-loud hardening — gsMain umbrella, lock exit 5, prototype-safe ownership bun in -e mode swallows uncaught exceptions thrown after a require() and exits 0 (verified on 1.3.13; uncaughtException handlers never fire either), so any runtime throw in a mutator was a SILENT SUCCESS. Every script body now runs inside a gsMain try/catch that prints "internal error ... refusing to mutate" and exits 4. Also: lock give-up now exits 5 instead of 0 (callers must not report a skipped mutation as registered); basename lookup uses hasOwnProperty so a foreign hook named "toString"/"constructor" can't resolve to an inherited Object.prototype member and abort the sweep; ownership-checked release also clears an empty/missing owner file; backup rotation sorts by mtime, not name; Windows-only backslash normalization (a legal Unix path containing a backslash is no longer rewritten); GSTACK_SWEEP_EXCLUDE_SOURCES lets a sweep spare named sources; lock tradeoffs documented at the lock helper. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(setup): honest hook-registration reporting + verify-gate sweep exclusion _install_plan_tune_hooks now propagates per-add-event failures (lock contention exits 5, fail-closed settings errors exit 3) and both caller sites branch on it: success logs the installed message, failure logs a visible "NOT registered — re-run ./setup" warning instead of claiming success for a mutation that never happened. --no-team's identity sweep runs with GSTACK_SWEEP_EXCLUDE_SOURCES= verify-gate: turning team mode off must not delete the user-registered verify-gate opt-in whose binary still exists (uninstall still sweeps it, correctly, because there the binary itself is being removed). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * test: adversarial regression pins — wrong-shape fail-loud, prototype basename, sweep exclusion, lock exit 5 New pins for the fail-loud hardening: a wrong-shape hooks value (object where an array belongs) exits 4 with "refusing to mutate" and leaves the file byte-identical (pre-gsMain this was a silent exit-0 no-op); a foreign hook whose basename collides with Object.prototype ("toString") survives an --all sweep that still removes gstack rows; GSTACK_SWEEP_EXCLUDE_SOURCES preserves the verify-gate row during --all; the fresh-foreign-lock test now asserts the loud exit 5 instead of a quiet skip. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * test(verify-gate): allow the --no-team sweep exclusion, keep registration banned setup now legitimately mentions verify-gate once: the --no-team identity sweep excludes it via GSTACK_SWEEP_EXCLUDE_SOURCES so team-mode teardown can't delete a user-registered gate. The opt-in pin tightens from a blanket not-contains to: every mention must be a comment or that exclusion, and no mention may sit on an add-event line. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(settings-hook): GNU-first stat in the lock stale check — Linux abort on held locks On Linux, BSD-style `stat -f %m` prints a multi-line FILESYSTEM block to stdout before exiting 1, so the BSD-first || chain captured that garbage concatenated with the real `stat -c %Y` epoch. The non-numeric mtime made `$(( now - mtime ))` a syntax error and set -e killed the binary with exit 1 whenever a lock dir already existed — every contention path (stale takeover, give-up, concurrent writers) broke on CI while staying green on macOS, where BSD stat -f succeeds cleanly. GNU `stat -c %Y` now goes first (BSD stat rejects -c with no stdout, so macOS falls through cleanly), and a numeric guard blanks any residual garbage so a future platform quirk degrades to the normal give-up path instead of an arithmetic abort. Same defect class as gstack-repo-mode's GNU-first ordering (garrytan#2195). Verified in an oven/bun Linux container: the four CI-failing lock tests now pass (62/62 across both files). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * test(uninstall): 30s budgets for the two subprocess-heavy behavioral tests Both tests spawn the copied uninstaller, which itself runs several settings-hook bun -e children (the lock-contention one also waits out a 300ms give-up per call). On a loaded box those cold starts blow bun's default 5s per-test timeout, and a timeout kill reports as a bare fail with no assertion diff — observed at 5.6-8.5s under load avg 25+. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
…ys included, verified live (garrytan#2646) * fix(browse): revokeToken deletes ALL tokens for a clientId, not the first Map hit revokeToken deleted the first Map entry matching the clientId and returned true. After a normal pairing, two entries share one clientId: the spent setup key (kept by exchangeSetupKey for idempotent re-exchange) and the session token, in that insertion order. Revoke ate the setup key, reported success, and the live session survived: DELETE /token/<id> returned a false 200 while /agents kept listing the agent. Worse, an unspent setup key created after the session survived revoke, so a "revoked" agent could POST /connect and mint a fresh session within the key's 5-minute validity window. revokeToken now deletes every matching entry and returns the delete count (truthy-compatible with the old boolean). The DELETE /token handler logs "Revoked N token(s)" and returns tokens_deleted so the multi-token class stays visible; revokeSkillToken wraps Boolean() to keep its documented contract. Regression tests pin shapes a (spent-key shadowing), b (re-grant hole), c (multiple pending keys), and bystander isolation. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat(browse): tunnel revoke/agents CLI with post-revoke verification `$B tunnel revoke <name>` was documented in the instruction block, pair-agent/SKILL.md, and REMOTE_BROWSER_ACCESS.md but implemented nowhere: the CLI forwarded it to the daemon as Unknown command 'tunnel', and nothing in the repo called DELETE /token/:clientId or GET /agents. New pre-server short-circuit (garrytan#2254 pattern: tokens are memory-only, never boot a daemon to revoke against it). `tunnel revoke <name>` DELETEs the token, prints the deleted count ("(count unknown)" for old daemons that answer {revoked} without tokens_deleted), then RE-READS GET /agents to prove the agent is gone. The still-listed branch is the version-skew net: a new CLI against a still-running old daemon with the first-match revoke bug exits 1 and says to re-run (each old-daemon call deletes the next match) or stop. An alive pid with an unreachable port reports "Could not reach daemon" (exit 1), never a false "no daemon". `tunnel agents` lists sessions plus pending (unexchanged) setup keys, which GET /agents now exposes via listTokens({includeSetup}) — without them the revocation view was blind to a paired-but-never-connected agent. Setup-key tokens never leave the server. DELETE /token/ now decodeURIComponents the clientId (400 on malformed encoding) so CLI-encoded names round-trip. Tests: subprocess CLI coverage (usage paths, no-daemon exit 0 without spawning, live pair/connect/revoke loop, pending-key listing), stub-daemon pins for the skew and unreachable branches, and e2e pins for revoke-all semantics, percent-encoded ids, and the second-DELETE-is-404 regression. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(browse): CLI always sends explicit pair scopes via shared DEFAULT_PAIR_SCOPES The effective pairing default lived in two places: the CLI omitted scopes unless --restrict was passed, and the server filled in its own literal. handlePairAgent now always sends an explicit scopes list and both sides reference one exported constant, DEFAULT_PAIR_SCOPES, so the default cannot silently drift again (pinned by a server-auth source tripwire). Three input traps closed in the same surface: - Bare --restrict (or --restrict swallowing the next flag) parsed as "no restriction" and silently granted FULL access, the opposite of the user's intent. validatePairAgentFlags rejects it pre-server, before any consent gate, so an arg error never boots a daemon. - A scopes list could smuggle the control scope past the explicit flag: --restrict "read,control" minted a control-scoped session with no --control. /pair now 400s on control in a scopes list without the control flag, and the CLI points the user at --control. - Option typos validated only at exchange time: createSetupKey stored any scope string and any rateLimit, so /pair returned 200 with a poisoned setup key whose failure surfaced to the REMOTE agent at /connect as a misleading "Invalid request body". Shared validation now runs in both creators and throws typed InvalidScopeError; /pair and /token 400 with the message, naming the bad scope or negative rateLimit. Also `opts.rateLimit || 10` became `?? 10` so the documented "0 = unlimited" survives the /pair path. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(browse): 403 hint stops recommending --admin; invariant names both scope defaults The scope-denied hint told restricted agents to "re-pair with --admin for eval/cookies/storage" — but --admin is a legacy alias for --control, so following it over-granted browser-wide destructive commands on top of the admin scope the default already carries. The hint now matches the CLI's sibling wording: re-pair without --restrict for page access, --control for browser control. Registry invariant #2 claimed "admin scope denied by default" three releases after b73f364 deliberately made /pair grant admin. It now names BOTH defaults precisely (registry API functions default read+write; the /pair ceremony grants DEFAULT_PAIR_SCOPES) so the header cannot lie one layer down. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs(pair-agent): document the full-access default, --restrict, and real revocation The pairing docs still described the pre-b73f3644 model: read+write default, --admin as the opt-in for JS/cookies/storage. Reality for three releases: /pair grants read+write+admin+meta (the pairing ceremony is the trust boundary) and --admin is a legacy alias for --control. A user following the skill believed they granted a sandboxed session and actually granted JS execution on their logged-in browser. pair-agent/SKILL.md.tmpl (SKILL.md regenerated in this commit) now states the real default, the tunnel-allowlist nuance (eval works remotely; the js/cookies/storage commands are local-only), --restrict for sandboxed sessions with an untrusted-content advisory (scope caps prompt-injection blast radius), and --control for browser-wide ops. "Revoking access" documents the now-real tunnel revoke (deletes session + pending setup keys, verifies against the agent list) and tunnel agents, and replaces the never-implemented `tunnel rotate` with `$B stop` — tokens are memory-only, so a daemon restart already rotates everything. REMOTE_BROWSER_ACCESS.md: /connect example shows the real default scopes, the scope table gains the control row, the 403 hint row matches the new server wording, and the false claim that /sidebar-chat is on the tunnel allowlist is gone (TUNNEL_PATHS is /connect + /command; /sidebar-chat no longer exists in server.ts at all). ARCHITECTURE.md drops the same phantom endpoint from the allowlist prose and endpoint table. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * v1.68.2.0: revoke-all, real tunnel revoke, truthful pairing docs Version slot allocated against the live remote via bin/gstack-next-version (clean patch bump from 1.68.1.0, no collision). CHANGELOG entry covers the revoke-all fix, the new tunnel revoke/agents CLI, the explicit-scopes wire contract, and the pairing-docs truth pass. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(browse): adversarial-review hardening — 6 findings fixed, regression-pinned Pre-push adversarial review (4 lenses, refute-style verification: 13 raw findings, 7 refuted, 6 confirmed) caught these; each fix carries a pin: 1. --restrict=read (equals form) sailed past validatePairAgentFlags — hasFlag/parseFlag are exact-token matches — so the user asked for a read-only sandbox and silently got FULL access: the exact failure mode this branch claims to close. The equals form is now a hard error before any server work. 2. handleTunnel trimmed the agent name but clientIds are stored verbatim, so a space-padded agent was unrevocable by the documented kill switch (trimmed DELETE 404'd while the grant stayed live). Names now pass through verbatim; the live-daemon test revokes ' padded'. 3. The sole pin for "CLI always sends explicit scopes" passed vacuously on a simulated revert: toContain('DEFAULT_PAIR_SCOPES') was satisfied by a comment. The tripwire now matches the code shape with a regex and bans the conditional spread formatting-insensitively. 4. The rewritten 403 scope hint was unpinned — new e2e asserts it names --restrict and --control and never --admin. 5. tunnelRevoke's verify-failure and HTTP-error branches and tunnelAgents' unreadable-list branch had no coverage — three stub-daemon pins added (an unreadable list must never render as "No paired agents"). 6. CHANGELOG claimed "40+ new test cases"; the honest count is 35. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
…e spot (garrytan#2665) * fix(pairing): reject reserved clientId 'root' at all token writers 'root' is the sentinel checkScope/checkDomain/checkRate and the server command gate use for the omnipotent caller, so a scoped token carrying it bypasses every enforcement path. Add ReservedClientIdError + a shared assertValidClientId; createToken/createSetupKey throw, restoreRegistry skips-and-logs (a corrupt state file must not brick boot). /pair and /token surface it as a named 400, and the CLI fast-fails --client root. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(pairing): release tab ownership on revoke tabOwnership cleared only on tab close, so after DELETE /token a same-name re-pair inherited the revoked agent's authenticated tabs (own-only access keys on owner === clientId). Add BrowserManager.releaseClientTabs and run it unconditionally in DELETE /token (ownership outlives the token, so an expired-token client can still own tabs); 404 only when both nothing was revoked and nothing released. Response now carries tabs_released. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * v1.68.3.0 fix(pairing): re-pair to narrow revokes the old grant on the spot POST /pair minted a new setup key but never touched the agent's live session, so re-pairing --client X --restrict read while X was connected (or whose 5-min key expired unexchanged) left the original full-access session, eval included, alive up to 24h. A reducing re-pair (fewer scopes, tighter domains, lower rate, stricter tab policy) now revokes the live session and releases its tabs before minting the new key (grantReducesAccess + revokeClientFully; superseded in the response). Non-reducing re-pairs keep the session and only drop stale PENDING setup keys, so a broaden/refresh never strands a working agent and a narrowing re-pair issued before the agent connects can't leave the old broad key exchangeable. Revoke happens before mint (revokeToken deletes all of a client's tokens). CLI prints a version-skew-safe supersede notice and warns when a re-pair-shaped call omits --client. Docs + CHANGELOG + VERSION. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(pairing): harden re-pair per adversarial review Adversarial review of the diff found four issues, now fixed: - Validate the requested grant BEFORE the supersede revoke: a reducing re-pair with a bad scope/rate no longer destroys the live session and then fails to mint a replacement (assertValidTokenOptions runs up front). - A re-pair with no live session releases tabs orphaned by an expired incarnation, closing the tab-inheritance gap /pair had (DELETE /token already released unconditionally). - Test the DELETE /token revoked=0/tabs>0 path and the /pair orphaned-tab release at the handler level (HTTP e2e can't, headless owns no tabs). - Test the CLI --client root fast-fail; fix its null-guard (parseFlag returns null when --client is absent). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: Garry Tan <garry@ycombinator.com> Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
…orbed, tracker closed with receipts (garrytan#2666) * test(wireup): make gbrain-missing PATH fixture hermetic The gbrain-missing test appended the host PATH (and a hardcoded /opt/homebrew/bin) to the fixture PATH, so on any machine with a real gbrain installed the 'missing' case saw it, exited 0 instead of 2, and could never fail where the bug exists — a false green for a whole machine class. The fixture now keeps only root-owned OS dirs on the child PATH, and a new determinism check plants a host-like gbrain to prove it is unreachable. Absorbed from PR garrytan#2615 with authorship preserved; the PR-thread liveness screenshot (docs/images/gstack-pr-liveness-2255.png) is dropped — referenced by nothing in the tree. Fixes garrytan#2255 Co-authored-by: CommandCodeBot <noreply@commandcode.ai> * fix(evidence): stop bun's dotenv autoload from reaching the spawned command `bin/gstack-evidence` has a `#!/usr/bin/env bun` shebang, and bun AUTO-LOADS `.env`, `.env.<NODE_ENV>` and `.env.local` from the cwd into `process.env`. The wrapper then spawned the command with no `env` override, so every command run through it inherited those variables — and a repo `.env.local` routinely holds production credentials. Two things go wrong, and the second is worse than the leak: 1. Secrets reach a child that would not otherwise have them. `npm test` run by hand in the same shell sees none of them; the same command through the wrapper sees all of them. 2. THE COMMAND UNDER TEST BEHAVES DIFFERENTLY, so the ledger certifies a run that is not the run CI performs. Observed in a Next.js repo on 2026-08-20: four tests failed 4/4 through the wrapper and passed 5/5 without it, because app code branched on env vars only the wrapper supplied. Nearly an hour went into chasing a "flake" that was the measuring instrument. The wrapper exists to record trustworthy evidence, so silently altering the environment defeats its purpose. The fix builds the child env from `process.env` minus the keys bun injected, and detection is exact rather than heuristic: verified on bun 1.3.11, a dotenv file does NOT override a variable the shell already exported (the shell's value wins). So a key whose live value equals the dotenv file's value was injected by bun, and dropping it restores the environment the user's own shell would have given the command. A key whose live value differs is genuinely the caller's and survives. `BUN_DOTENV_FILES()` mirrors bun's precedence, including that `.env.local` is skipped when NODE_ENV is "test" — scrubbing a key bun never loaded would strip a variable the caller legitimately provided. Escape hatch: GSTACK_EVIDENCE_KEEP_DOTENV=1 keeps the old behaviour. When keys are scrubbed the wrapper warns with the KEY NAMES ONLY, so the diagnostic cannot become the leak it prevents. Tests: 6 cases, mutation-verified — removing `env: spawnEnv` reddens exactly the two leak tests and restoring it gives 30/30. Every leak test asserts the scrub warning fired, because `bun test` sets NODE_ENV=test and the first version of these tests passed vacuously against a `.env.local` bun had never loaded. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Absorbed from PR garrytan#2652 with authorship preserved. Wave additions: a doc-comment on the ${VAR}-expansion limitation (bun expands refs, the reader compares raw text — those keys are left in the child env, failing open) and a regression pin for the unreadable-.env fail-open path with a functional DAC-override skip guard. Fixes garrytan#2624 * fix(setup): reap dangling skill dirs when the payload is gone cleanup_old_claude_symlinks derived its work list from the payload directory, so when the payload was gone — precisely when orphans exist — the glob matched nothing and the loop never ran; the -f guard also followed symlinks, hiding dangling SKILL.md links even with a payload present. The cleanup now scans the DESTINATION skills dir (-e/-L, so dangling symlinks are visible) and anchors SKILL.md provenance to path segments (gstack/*, */gstack/*, */.gstack/render/claude/*) instead of a bare *gstack* substring that would eat a user skill under ~/tools/gstack-fork/. The Windows real-file arm stays payload-gated: a real file has no provable owner. Absorbed from PR garrytan#2634 (2 commits squashed) with authorship preserved. The symmetric cleanup_prefixed_claude_symlinks hole is filed as a TODOS.md residual in this wave. Fixes garrytan#2204 * fix(redact): tolerate EEXIST from recursive mkdir in install-prepush-hook on bun/Windows (garrytan#2635) fs.mkdirSync(dir, { recursive: true }) is a no-op on an existing directory in Node, but bun on Windows throws EEXIST - crashing hook install on any repo whose .git/hooks already existed, leaving the repo unprotected. Add lib/fs-utils.ts mkdirpSync: swallow EEXIST only when statSync confirms the path is an existing directory; a regular file occupying the path, a stat failure, or any other errno still rethrows. Use it in installPrepushHook(). The regression test emulates the Windows bun fs semantics via a bun --preload fixture, so the exact crash path runs (and fails on the old code) on any platform, including CI Linux. Absorbed from PR garrytan#2641 with authorship preserved. Fixes garrytan#2635 * fix(bin): route remaining Windows-reachable mkdirSync sites through mkdirpSync Sweep follow-up to garrytan#2641's lib/fs-utils.ts helper: bun on Windows throws EEXIST from a recursive mkdir on an existing dir, so every unguarded recursive mkdirSync on a Windows-reachable path is a latent crash. Converted: bin/gstack-decision-log (unguarded, runs on every decision log — the second call on any machine hits the pre-existing projects dir), bin/gstack-evidence logsDir + ledger dir sites, and bin/gstack-redact-prepush's skip-log site (already try-wrapped, so its failure mode was a silent skip-log loss rather than a crash — the fix makes the log survive). The ~15 remaining gbrain/mac-lane sites are deliberately left alone. Regression: fs-utils.test.ts drives gstack-decision-log twice, the second run under the bun-Windows EEXIST preload fixture — the pre-sweep code exits 1 with EEXIST there; verified red against v1.68.3.0. * fix(setup-gbrain): warn about the ZeroEntropy sunset before Sept 4 ZeroEntropy was acquired by Notion and sunsets its hosted API on September 4, 2026. A gbrain configured with the zeroentropyai embedding recipe keeps importing pages after that date but embedding silently fails — pages land structurally with no semantic search, this repo's tracker P1 (TODOS.md NEXT PRIORITY). Nothing in gstack ever recommended ZeroEntropy (the dependency is gbrain-internal), so the gstack side is detection + advisory: the wireup helper warns when ~/.gbrain/config.json names the recipe (fail-open grep — a missing, unreadable, or other-provider config stays silent and never blocks a working setup), the setup-gbrain provider-default comments say never to select the legacy recipe for a new brain, and USING_GBRAIN_WITH_GSTACK.md gains a troubleshooting entry. The gbrain-side provider migration stays open upstream. Refs garrytan#2365 * fix(gbrain-source-wireup): first sync targets the registered source, not --repo The wireup registered a federated source by id, then ran 'gbrain sync --repo $WORKTREE' — which resolves against the brain's DEFAULT source and (on gbrain 0.46.x) rewrites that source's local_path anchor to our worktree. Net effect: the user's primary knowledge source silently repointed at the gstack brain worktree while the just-registered source got zero pages, and pages_synced still reported success. The sync now targets the registered id ('gbrain sync --source $id', the same form the repo's own troubleshooting documents). Because the script's stated floor is gbrain >= 0.18.0 and nothing proves --source exists there, support is probed via 'gbrain sync --help' first: an older gbrain keeps the wrong-but-working --repo call with an upgrade warning instead of converting it into a hard failure. The probe sits after the GSTACK_BRAIN_NO_SYNC early-exit and is unreachable in --probe mode. Regression tests (fail on v1.68.3.0): a no-skip sync case asserting the call log shows 'sync --source gstack-brain-<id>' and never 'sync --repo', and an old-gbrain fallback case (fake sync --help without --source) asserting --repo plus the upgrade warning. Fixes garrytan#2662 * fix(setup): --host slate exits informatively instead of silently installing nothing slate passed --host validation (added to the accept-list in v1.64.1.0) but never got a dispatch arm, and the all-INSTALL_*-zero fallback lives inside the auto branch — so './setup --host slate' configured nothing and exited 0, a silent no-op strictly worse than the original hard rejection. slate is now an informational arm (per docs/designs/SLATE_HOST.md it is blocked on the host-config refactor; Slate reads .claude/skills as a compatibility fallback, so the arm points at './setup --host claude'), and a defensive guard after the dispatch chain errors loudly (naming the host, the missing arm, and the valid targets, exit 1) if a future host is ever accepted without being wired. Regression tests (fail on v1.68.3.0): a dispatch-arm ratchet asserting every accept-listed install target has a matching dispatch branch — the exact drift class; a registry cross-check deriving both sides from hosts/index.ts and setup's case arms; a behavioral slate probe (exit 0, points at --host claude, never reaches the installer — on unfixed code it fell through into the installer); and a static pin on the guard's shape. Fixes garrytan#2361 * fix(make-pdf): resolve the sibling browse binary from execPath, not argv[0] In a bun-compiled binary process.argv[0] is the raw invocation string — often relative ('./pdf', 'pdf') — so dirname(argv[0]) yielded '.' and the sibling candidates (../browse/dist/browse etc.) resolved against the CWD instead of the install dir. Resolution was cwd-dependent: correct-by-luck when the fallbacks rescued it, wrong when a cwd-relative path matched. process.execPath is always the absolute binary path. The resolution step takes an injectable selfPath (defaulted) because under bun test the process path is the bun runtime and the compiled-binary shapes are otherwise unreachable. The issue's other half — pdf setup failing on newtab('about:blank') — was already fixed on main in v1.64.0.0 (browse/src/url-validation.ts exact-match allows about:blank; its comment names this exact smoke). This commit closes what remains. Regression tests (the sibling-via-selfPath case fails on v1.68.3.0 — pre-fix code ignores the seam and either resolves the global install or throws): sibling resolution from an install-shaped tree, and a decoy-browse-DIRECTORY case pinning that a directory never wins resolution. Fixes garrytan#2156 * fix(memory-ingest): store the normalized git_remote so unattributed pages hit the policy filter buildTranscriptPage wrote the normalized '_unattributed' sentinel into the page FRONTMATTER but stored the raw resolved remote ('' when unresolvable) on the page object. The policy filter fast-paths !p.git_remote, so under --include-unattributed an explicit '_unattributed → deny' (or read-only) policy never applied to exactly the pages it names — they ingested unpoliced. The stored value now matches the frontmatter. Regression test (fails on v1.68.3.0): seeds the REAL bin/gstack-gbrain-repo-policy store with '_unattributed → deny' through its own set verb, ingests an unresolvable-remote session with --include-unattributed, and asserts nothing reaches gbrain — pre-fix the '' remote bypassed the filter and the import ran. A fake echoing tiers would pass on both sides of the fix; the real helper prints 'none' for unknown keys, so only a genuinely applied deny distinguishes the two. Fixes garrytan#2353 * fix(land-and-deploy): MERGED recovery reconciles and reports remote-branch cleanup Step 4's merge commands carry --delete-branch, and the success path tells the user 'The branch has been cleaned up.' When gh exits non-zero AFTER GitHub already merged (routine in worktree layouts: gh's local cleanup runs git checkout <base> and fails), the §4a-postfail MERGED recovery re-established everything EXCEPT the branch deletion — and said nothing about it, so the discrepancy was invisible. The MERGED path now reconciles: git ls-remote --heads distinguishes branch-already-gone (exit 0, empty → 'already cleaned up', idempotent on re-runs) from branch-survived (offer confirm-first deletion, matching the section's worktree posture; -d not -D for any local branch) from check-itself-failed (non-zero exit → 'couldn't verify', skip the offer — never read a failed check as a clean branch). Template + regenerated SKILL.md + test extensions land in one commit (the md-sync assertion goes red otherwise). Regression assertions (fail on v1.68.3.0: no delete-branch reconciliation existed in test/ at all) pin the ls-remote check, the confirm-first delete, and the absent-vs-failed distinction. Fixes garrytan#2656 * fix(scripts): stop heredoc bodies deadlocking under Homebrew bash `./setup --help` can hang forever on macOS, printing nothing, with no way to tell it apart from a slow install. Eleven scripts carry the same latent hang, `setup` itself being the one every user hits first. bash 5.2+ delivers a heredoc body of 64KiB or less through a pipe: the forked child writes the entire body before exec, and nothing reads the other end until the command starts. Under macOS pipe-KVA pressure the kernel hands a fresh pipe a 512-byte buffer instead of the usual 16-64KiB, so any body of 512 bytes or more blocks write() permanently. The capacity check bash would need to notice (F_GETPIPE_SZ) is Linux-only, so it never fires here. It is pressure-dependent, which is why it reads as "worked on my machine" — the same script runs fine all day and then wedges. Homebrew bash is what `#!/usr/bin/env bash` resolves to on a Mac with brew on PATH, which is most of them. Apple's /bin/bash 3.2 predates the pipe path and is unaffected, so the bug is invisible to anyone testing with the system shell. The fix is `BASH_COMPAT=50` in each affected script, which restores the pre-5.2 tempfile path: $ bash -c 'probe() { [ -p /dev/stdin ] && echo PIPE || echo TEMPFILE; } probe <<EOF $(printf "x%.0s" $(seq 1 1000)) EOF' PIPE $ BASH_COMPAT=50 bash -c '...same...' TEMPFILE - Not a `#!/bin/bash` shebang swap: that pins the script to whatever bash lives at /bin (3.2 on macOS, absent on some Linux distributions) and is bypassed entirely by `bash script.sh` call sites. The variable survives both. - Not exported, so child processes keep their own compat level. - Placed below any `--help` sed range that reads $0, so usage output is unchanged (verified on all eleven). - Every guarded script is bash-3.2-clean — no associative arrays, case conversion, or mapfile — so compat level 50 costs them nothing. test/heredoc-pipe-deadlock.test.ts scans every tracked shell script for a heredoc body in the 512B-64KiB window and fails without the guard, and proves the mechanism at runtime on bash 5.2+ by asserting the body moves from PIPE to TEMPFILE. On older bash the runtime half is skipped, since the pipe path does not exist there. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Absorbed from PR garrytan#2640 with authorship preserved. Wave adaptations: the pipe-probe test skips on minimal-/dev environments without /dev/stdin (it would report OTHER for an unobservable fd), and one caveat verified during review: on bash 4.3/4.4 (e.g. Git Bash), assigning BASH_COMPAT=50 prints a non-fatal 'invalid value' warning to stderr — those bashes are already on tempfiles, so the guard is a no-op there; windows-setup-e2e exercises this empirically. * docs: TODOS.md v1.69 wave close-out Move the slate P4 entry and the ZeroEntropy P1's gstack-side half to Completed (v1.69.0.0); reframe the ZeroEntropy NEXT PRIORITY entry around the remaining gbrain-side work; file the wave's four residuals with rationale — the prefixed-cleanup symmetric conversion, the garrytan#2163 legacy-slug checkpoint heal, the invited garrytan#2657 --reconcile contribution, and the table-driven setup host dispatch behind the new cross-check ratchet. * chore: bump version and changelog (v1.69.0.0) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: Som Samantray <som.samantray@gmail.com> Co-authored-by: CommandCodeBot <noreply@commandcode.ai> Co-authored-by: Connex Client Access <paul@paulkortman.com> Co-authored-by: y$un_ <forrest.sun527@gmail.com> Co-authored-by: Lockyer <135391289+Lockyer228@users.noreply.github.com> Co-authored-by: Benjamin D. Smith <benjamin.smith@binarysword.com> Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
…ision point (tripwire + gate E2E) (garrytan#2700) * fix(ship): name the /document-release subagent at every Step 18 decision point The v1.54.0.0 carve moved Step 18 (documentation sync) into ship/sections/pr-body.md and the Claude-host skeleton stopped saying "document-release" anywhere in the workflow body — the dispatch became invisible at exactly the moments an agent decides whether to open the section. Restore visibility at three touchpoints, all subagent-framed (never bare-slash-framed, which would invite an inline Skill invocation that bypasses the fresh-context subagent + JSON contract): - manifest trigger (renders into the section-index row AND the STOP pointer): "dispatching the /document-release subagent to sync docs (Step 18) and then creating or updating the PR/MR (Step 19)" - Step 17 handoff line names Step 18's dispatch explicitly - new hoisted doc-sync invariant beside the PR-title invariant: the dispatch itself is never skipped; only a failed subagent is non-blocking Pin it in carve-guards: 'the /document-release subagent' (all three touchpoints) + 'dispatches the /document-release subagent' (invariant) must stay in the skeleton; the carved imperative 'Dispatch /document-release as a subagent' must stay carved. Skeleton cap 91,600 → 92,300 (measured 91,764; trigger renders twice). Goldens regenerated for all three hosts. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * test: pin the ship→document-release Step 18 wiring with a free tripwire Five substring/structure asserts across the carved section, the Claude skeleton's three touchpoints, the manifest trigger, and the codex/factory goldens (inlined Step 18 ordered before Step 19). Claude-golden asserts deliberately omitted: host-config.test.ts already enforces golden == generated byte-for-byte. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * test: gate-tier E2E proving /ship dispatches the document-release subagent New skill-e2e-ship-docsync: a live agent gets the sliced Step 17→19 tail of the generated ship skeleton in a bare-remote git fixture (Steps 0-16 "done"), under a fake HOME so the STOP pointer and the Step 18 subagent prompt resolve to planted copies, with a stub document-release skill that returns the empty-result JSON contract. Hard assert: an Agent/Task tool-call matching /document-release/i exists in result.toolCalls and precedes any `gh pr create`. Neutral prompt (no STOP-Read priming, no document-release mention — the prompt echoes into the transcript, so asserts read toolCalls only). Hardening from review: throw-on-marker-drift fixture slice; per-test GSTACK_HOME + .redact-prepush-prompted marker (routes Step 17's credential guard to its silent branch — the hermetic GSTACK_HOME pin defeats a HOME-only override); 480s/540s timeouts (nested subagent adds wall clock the 300s sibling never carried); 'timeout' accepted in exitReason only because the dispatch assert is independently hard; whole-file describeE2ETier('gate') composed with diff selection (keeps the file out of the periodic shard census, which sits at its ceiling, and under the hard tier-alignment invariant). Registered as 'ship-docsync' in E2E_TOUCHFILES + E2E_TIERS (gate) in the same commit — touchfiles.test.ts rejects either half landing first. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs: fix stale document-release TODOS entry + three review-deferred items The SHIPPED entry still described the deleted Step 8.5 post-PR cat-delegation design from v0.8.4; replace with the current Step 18 subagent design and its test pins. Add the three P3 items deferred from the v1.69 plan review: dispatch receipt enforcement, land-and-deploy→canary dispatch-pin pattern, and the periodic shard-census boundary. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix: pre-landing review fixes Testing-specialist findings, all mechanical: (1) pin the E2E fixture's git branch (-b main / init.defaultBranch=main) and assert every setup command's exit status so operator git config can't silently corrupt a paid run; (2) tighten the dispatch matcher to Step 18-prompt-specific markers (document-release/SKILL.md | executing the /document-release workflow) so a subagent merely quoting section text can't false-pass the regression assert (verified against recorded burn-in transcripts); (3) replace the subsumed carve-guards anchor with three non-overlapping per-touchpoint anchors (gerund/imperative/3rd-person) so each touchpoint is independently enforced. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix: red-team review fixes Five informational findings: TODOS shard-census arithmetic corrected (census is 67 with one free ungated slot; the SECOND ungated file trips the floor) and version pointer fixed (v0.18.2.0, not v0.18.1.0); the free tripwire now pins the two dispatch-matcher marker strings so a pr-body prompt reword fails the free suite instead of surfacing as a paid-tier mystery; the E2E matcher gains a section-paste exclusion (scaffold strings disqualify) — verified against all recorded runs; the E2E header documents the tierless test:evals invisibility tradeoff. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix: adversarial review fixes Pin the E2E matcher's two EXCLUSION markers in the free tripwire (an unpinned 'Parent processing:' reword would silently deaden the section-paste guard while every test stayed green); add an ordering pin (the hoisted doc-sync invariant must sit above the pr-body STOP pointer — presence-only anchors can't catch drift below it); plant a third cwd-relative pr-body copy inside the fixture repo, gitignored so the agent never tries to commit test scaffolding. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * chore: bump version and changelog (v1.70.1.0) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs: CHANGELOG accuracy fixes from the doc-release review Three factual corrections the Step 18 doc subagent caught in the fresh v1.70.1.0 entry: 5 tripwire tests (not 6), cost floor $0.63 per the cited eval store (not $0.59), and the visibility claim scoped to decision points (the re-run checklist mention survived the carve). Plus the E2E header's stale pending-burn-in note replaced with the observed numbers. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix: raise bun-polyfill subprocess budget to 60s for degraded Windows runners The 50ms-sleep test blew the 20s budget on BOTH bun retry attempts on PR garrytan#2700's windows-latest runner (run 32989821401) — sustained AV/runner pressure, not just the documented cold-start. Same flake passed-on-rerun on the prompt-token-load-reduction branch yesterday. Budget only; every assertion still checks exact output. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(ci): run the ship-docsync gate E2E in the evals matrix + silent-skip tripwire The evals.yml matrix is hand-enumerated and the Run step never exported EVALS_TIER, so the new whole-file-gated ship-docsync E2E would have self-skipped even with a row — a hollow green one layer deeper than the documented rehomed-monolith incident. Add the e2e-ship-docsync row with a row-level `tier: gate` property, exported as EVALS_TIER by the Run step (empty = unset for every existing row: all readers are `=== '<tier>'` or truthiness). New free tripwire test/evals-workflow-matrix.test.ts ratchets the class: matrix files must exist; gate-hosting files must have a row; whole-file-gated matrix files must carry a matching row tier; and the burn-down lists enforce their own cleanup. It enumerates the PRE-EXISTING holes found while wiring this (8 gate-hosting files with no row; codex/gemini rows running zero tests; the pty-plan-smoke row hollow since its files adopted describeE2ETier) — tracked in TODOS as the CI gate-lane hollow-coverage burn-down. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
…d onboarding, 20 skill carves, CLAUDE.md trim (garrytan#2691) * feat(gen): strip gen-time-only frontmatter keys from Claude renders interactive + benefits-from are read from the .tmpl by buildContext at generation time; no runtime, host, or test reader consumes them from the generated SKILL.md (e2e-harness-audit reads .tmpl; benefits-from tests assert rendered prose). gbrain: stays (bin/gstack-brain-context-load reads it from the installed render); hooks: stays (Claude Code host wires PreToolUse from it). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * chore(gen): regenerate SKILL.md — dead frontmatter keys removed Mechanical regen after hosts/claude.ts stripFields change. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat(test): context-budget ratchet — CI ceilings on always-on + eager token ledgers New free test grades the two ledgers nothing else guards: the full-frontmatter always-on catalog (aggregate) and per-skill eager tokens (SKILL.md + forced-read refs), via checkBudget from lib/context-bill.ts. Ceilings live in test/fixtures/context-budget.json with x1.05/x1.10 headroom; regenerate with bun test/helpers/capture-context-budget.ts. New skills fail until consciously budgeted; removed skills fail until the fixture is refreshed; reductions ratchet the ceilings down so wins lock in. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs(todos): file output-template carve wave + plan-ceo doctrine revisit; mark preamble-carve P3 in flight Two follow-ups deferred from the approved token-reduction program (CEO review 'NOT in scope' list), filed with full context per TODOS format. The existing P3 preamble-carve entry gets a status update pointing at the program that supersedes it. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(test): review findings — Windows path normalization, full totals rebuild, ratchet coverage Pre-landing review (5 specialists) found one critical: the ratchet test runs in the curated Windows lane, where path.relative yields backslash skill names that miss the test/ filter and mismatch every POSIX fixture key. Names are now normalized once in buildRatchetBill (toPosixName) and the fixture filter is tightened to test/fixtures/. All eight Bill.totals fields are rebuilt from the filtered list (no fixture-polluted perInvocation/totalMd numbers for future consumers). New coverage: Windows-separator normalization pins, a captureContextBudget round-trip against tree-a (headroom math exact), a stripFields regression pin (interactive/benefits-from absent from renders, hooks/gbrain preserved), and the ceilings test no longer double-reports stale-fixture entries. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(test): adversarial findings — stable root key, symlink-alias dedupe, fixture-shape guard Adversarial review (Claude subagent) verified the fixture's root-skill key was the capture machine's checkout dirname: any non-gstack-named clone (every Conductor worktree) failed the free suite, and the documented re-run-the-capture recovery baked the local dirname into the committed fixture — silent corruption through the tool's own protocol. The root skill is now pinned to ROOT_SKILL_KEY ('gstack', its frontmatter name). Symlink aliases are realpath-deduped (census precedent): connect-chrome no longer gets its own ceiling, so Windows checkouts that materialize the symlink as a plain file can't fail the stale-ceiling set-equality test. New guards: fixture-shape validation (a string alwaysOnTotal can no longer silently disable the ceiling), a mutation pin that the filter shrinks the always-on ledger vs the raw bill, an alwaysOnTotal violation test (the branch was load-bearing with only under-budget coverage), and an atomic temp+rename fixture write. Fixture regenerated: 59 ceilings, alwaysOnTotal 6344. Deferred with a TODO: anchoring transformFrontmatter's denylist strip to the frontmatter block (latent, zero live collisions, pre-existing path). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * chore: bump version and changelog (v1.69.1.0) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs: update project documentation for v1.69.1.0 CLAUDE.md: Token ceiling section documents the context-budget ratchet as the third guard (test file, fixture, new-skill budgeting, capture command). CONTRIBUTING.md: Tier 1 guard list gains a Context-budget ratchet bullet; the Adding-a-new-skill checklist gains the budget-capture step. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs: pin exact guard semantics for the context-budget ratchet in CLAUDE.md Doc-review finding: "a third enforced ceiling" undercounted the guard family (skill-size-budget floors and parity ratios also watch these ledgers, relatively). Rephrased to match the ratchet test's own header: absolute ceilings vs relative floors/ratios. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs(changelog): heaviest-skill claim matches the fixture (land-and-deploy edges review by 0.2%) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat(bin): gstack-skill-start + gstack-skill-end — the preamble runtime, consolidated Absorbs the ~13KB of bash every tier-2+ SKILL.md inlined twice over (bootstrap fence + artifacts-sync fence) and the skill-end telemetry/sync fences. Same KEY: value STATUS-line contract the prose interprets, plus SKILL_START_PROTO handshake (OV5), SESSION_ID/TEL_START echoes, GSTACK_HOME-normalized state paths (EOV7), --parent-pid session identity (EOV5: $PPID inside the script is the ephemeral tool-call shell), OV4 sanitization of passthrough output, and a receipted daily artifacts pull (_receipted_git, brain-sync class, fail-closed). Per-line || true error style throughout (F3) — a mid-script failure never drops later STATUS lines. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat(gen): preamble resolvers emit a script invocation fence instead of inline bash generate-preamble-bash: ~6.3KB fence -> 4-line gstack-skill-start invocation (quoted-tilde pitfall handled: leading ~ interpolates through $HOME; env-var hosts keep $GSTACK_BIN) + degraded-mode prose (F1/EOV8: safe defaults, consent gates deferred-never-lost; OV5: proto rule). generate-brain-sync-block: ~6.8KB bash -> interpretation prose + the privacy stop-gate (stays inline until Phase 2's gated emission). generate-completion-status: telemetry fence -> one gstack-skill-end call with SESSION_ID/TEL_START handoff. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * chore(gen): regenerate all skills + golden fixtures — inline preamble bash removed Mechanical regen after the resolver change: −12,628 lines across 52 renders (corpus 952K -> 806K render tokens; tier-2 skills −11-13KB each). Golden per-host ship fixtures refreshed from the fresh claude/codex/factory renders. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * test: skill-start contract suite + preamble A/B eval + touchfiles registration test/gstack-skill-start.test.ts (11 free tests): STATUS-key contract vs the prose (F2), per-host fence resolution shapes (E1), proto-first, OV4 marker sanitization, --parent-pid identity, headless suppression, skill-end duration math + pending cleanup. test/skill-e2e-preamble-script-ab.test.ts (gate tier, OV7): inline-bash render (pinned from 2978597) vs script render with the fence redirected at the worktree bin (EOV2 — hermetic evals otherwise resolve the operator install and silently exercise degraded mode). 21 touchfiles dep lists gain the two bin scripts (EOV9) so future script edits select the preamble evals; selection-count pin updated 23->24. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * test: repin ~70 assertions to the script contract — every literal gets a successor Assertions that pinned inline-bash internals (update-check guard, _SESSIONS reaping, telemetry start/end blocks, routing probe, repo-strip producer, first-task gating, EXPLAIN_LEVEL/QUESTION_TUNING echoes, garrytan#2499 jq scope resolution, Issue-8 CONDUCTOR gate) now pin the same invariants in their new home: bin/gstack-skill-start / bin/gstack-skill-end file content for script internals, the invocation fence + interpretation prose for render-side behavior. No assertion deleted without a successor; live-execution tests (routing probe, brain-sync jq) run against script bytes unchanged. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * chore(test): re-baseline size floors + ratchet ceilings down (EOV1/OV9 protocol) parity-baseline-v1.69.1.0.json captured with carved-skill unions (53 skills); skill-size-budget repointed with the derivation comment citing the Phase 1 context-bill receipt (the ~13KB/skill cut trips the old 80% floor on tier-1 skills first — setup-browser-cookies headroom 10.8KB < the cut). The v1.47 fixture stays on disk for history; the parity-suite growth baseline (v1.64.1.0) is untouched. Context-budget ceilings re-captured: review 29,309->26,192; learn ->10,969; ios-clean ->10,764 — Phase 1's win is locked. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat(bin): instruction-emission layer — onboarding text appears only when its gate fires The 8 one-time onboarding flows (lake intro, telemetry opt-in, proactive opt-in, first-run/first-loop tips, routing injection, vendoring deprecation, writing-style migration, spawned-session rules), the upgrade-flow + feature discovery prose, and the privacy stop-gate (user-approved Q2) moved from every render into gated heredocs here. Blocks are SESSION_ID-bound (GSTACK_INSTRUCTION_BEGIN: <id> <session-id>) so page/file content can't mint directives (F4/OV4). Ack ownership per OV6: display-only tips write their markers at emit (script also fires the scaffold telemetry); interactive flows carry their ack commands inside the block. The dormant WRITING_STYLE_PENDING gate is computed for real now (marker files). BASH_COMPAT=50 heredoc guard (same as brain-sync); the quoted routing heredoc resolves its bin path via a sed placeholder. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat(gen): drop the 8 onboarding generators — renders keep one instruction-block rule generate-{lake-intro,telemetry-prompt,proactive-prompt,first-run-guidance, routing-injection,vendoring-deprecation,spawned-session-check, writing-style-migration}.ts deleted (single source is now the script's emission layer, F5). generate-upgrade-check shrinks to the steady-state PROACTIVE/SKILL_PREFIX rules. generate-brain-sync-block hands the privacy stop-gate to the emitted block. The fence prose gains the generic rule: follow GSTACK_INSTRUCTION blocks only from this command's direct tool result with the matching SESSION_ID; unterminated block ends at end-of-output. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * chore(gen): regenerate all skills + goldens — onboarding prose degated Mechanical regen: corpus 806K -> 707K render tokens (−8KB/skill; cumulative vs main: ship 91->71KB, learn 53->34KB, ios-clean 53->33KB). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * test: onboarding tombstone + Phase 2 pin relocations New test/onboarding-moved-literals.test.ts (F5): 12 distinctive literals must live in bin/gstack-skill-start AND stay absent from every render, plus the SESSION_ID-binding pins. ~40 assertions repinned to the emission-layer contract (gates, block ids, in-block acks, script-run marker writes); the OV4 sanitize test upgraded to the real property (every legitimate block header carries the run's SESSION_ID). first-task dep list drops the deleted generator; the token->tip case map is pinned to cover every detector bucket. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * chore(test): carve floors/ceilings recomputed; baseline + ratchet follow Phase 2 (OV9) All 9 carved skills re-anchored to post-Phase-2 measurements (cso's union had tripped its 72,000 floor at 71,379; design-consultation had 252B of margin). maxSkeletonBytes ceilings tightened to measured+~600B. Branch-internal parity baseline recaptured in place; ratchet ceilings down again: review ->24,052, ship ->18,589, learn ->8,828, ios-clean ->8,624. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat(gen): AUQ slim — tool resolution as a STATUS-line branch table, split rules to invariants + absolute pointer Tool resolution (1,799B) rewritten as a 3-branch table keyed on the echoed CONDUCTOR_SESSION/SESSION_KIND lines — Conductor prose-default, MCP-variant preference, and failure handoff preserved verbatim in behavior, including the auto-decide-first ordering and the gstack-question-log capture requirement. 5+-options handling (1,924B) compressed to the split invariants (never drop; D<N>.k shape; Include/Defer/Cut/Hold; question_id scheme with the never-ask refusal) + the full-rule pointer. Both doc pointers now interpolate the absolute install root (Codex outside-voice #7 convention) instead of the bare 'in the gstack repo'. Failure-fallback, Format, and self-check sections are byte-identical — all 14 MANDATORY always-loaded pins pass with zero test edits. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * chore(gen): regenerate all skills + goldens — AUQ slim Mechanical regen: −1.3KB per tier-2+ skill (ship 69.9KB, learn 32.5KB). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * chore(test): baseline + ratchet follow Phase 3 (OV9); OV8 evaluated — shrink floor stays Branch-internal baseline recaptured; ratchet ceilings down again. OV8's floor-retirement question, evaluated as planned after Phase 3: the 80% shrink floor stays — it uniquely catches accidental body deletion in non-carved skills BETWEEN ratchet recaptures, and the capture command has amortized the fixture-refresh cost that motivated retiring it. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat(review): carve adversarial, plan-completion, and review-army into sections The three resolver macros ship already carves as siblings now load on demand for /review too: skeleton 100.2KB -> 55.0KB (-45%), union 93.4KB. Resolvers stay the single source of truth (sections wrap the macros). Step 0/1, scope drift, critical pass, confidence calibration, and fix-first stay always-loaded. Fixtures and pins follow the moved content (codex-hardening wrapped-sites, review-army E2E fixture builds skeleton+sections with an empty-fixture guard). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat(codex): carve the three mutually exclusive modes into sections Review/Challenge/Consult mode bodies (34.7KB where at most one ever runs) load on demand: skeleton 81.0KB -> 55.2KB, union 1.04x the monolith. The mode dispatch, filesystem boundary, and a new always-loaded 'Synthesis recommendation (REQUIRED) — all modes' block stay skeleton-side (the AUQ per-skill pins pass unchanged); the plan-file report + exit gate render after the last section pointer per the gateAfterStop pattern. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat(land-and-deploy): carve first-run validation, readiness gate, and merge/deploy into sections The once-per-repo dry-run validation, the pre-merge readiness gate, and the merge + deploy-strategy steps (37.8KB) load on demand: skeleton 91.1KB -> 55.7KB. Step 1.5 keeps its detection bash as the dispatch; the first-run section's fingerprint-save block gained {{SLUG_EVAL}} so it is self-contained. Zero content lost (line-coverage checked against HEAD). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat(ios): demote the four ios skills to preamble-tier 2 (Phase 5) They never consume the tier-3 sections (repo-mode ownership, search-before- building) but do fire AskUserQuestion, which tier >=2 provides — verified by grep before the plan review. -2.2KB per skill. Render assertions pin the demotion (tier-3 sections absent, AUQ format present). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * chore(guards): register wave-1 carves; monolith invariants retire; baselines + ratchet follow CARVE_GUARDS gains review/codex/land-and-deploy (12 carved skills total); their MONOLITH_INVARIANTS entries retire (invariants now generate from the registry, cso precedent). Touchfiles: carve-section-loading covers the three new carves; the codex + land-and-deploy LLM-judge dep lists widen to their sections. Regen + goldens + branch-internal baseline + ratchet ceilings recaptured (review 24,052 -> skeleton-based ceiling; union floors hold). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * test(gen-skill-docs): review render pins read the carved union The review carve's readSkillUnion conversions (same pattern its neighbor carved-skill pins already use). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat(autoplan): carve the four review phases + tasks aggregator into sections Phase bodies (CEO/Design/Eng/DX consensus flows) and the Implementation Tasks aggregator load on demand; Design and DX stay separate sections because each is independently conditional on scope. Skeleton 83.7KB -> 58.7KB (-30% always-loaded); the 6 decision principles, classification, sequencing, and explicit skip-condition dispatch stay always-loaded. The chain E2E's phase-complete markers now live only in sections, so its assertions double as section-read proof (behavioral: external). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat(spec): carve the post-confirmation gate-and-file tail into one section Phases 1-4 are the turn-1 conversational spine — carving them would force the Read on the first user message for zero real savings. The mechanical tail (4.5/4.5a/4.5b redaction gates + Phase 5 filing + TTHW telemetry) fires only after draft confirmation: a genuine lazy boundary, kept as ONE section so the gh-issue-create bash can never load without the fail-closed redaction gate that precedes it. Skeleton 65.4KB -> 50.7KB; all ~85 phase-structure invariants migrated location-aware plus a new carve-shape suite (56 tests). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat(setup-gbrain): carve the branch-exclusive install paths into sections Brain-init (Paths 1/2/3/4 bodies), engine remediation, transcript gate, and CLAUDE.md persist load on demand — at most one install route ever runs. Skeleton 75.3KB -> 57.0KB; the Step 1 detect and Step 2 path dispatch stay always-loaded. New buildSetupGbrainFixture helper gives the periodic E2Es extract-don't-copy fixtures with a non-empty guard; the voyage-code-3 gate counts scan the tmpl union (the third init site lives in engine-remediation). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * chore(guards): register wave-2 carves (15 carved skills); autoplan monolith retires; baselines follow CARVE_GUARDS gains autoplan (behavioral: external via the chain eval), spec, and setup-gbrain; autoplan's MONOLITH_INVARIANTS entry retires. Touchfiles: setup-gbrain periodic dep lists gain the section tmpls + fixture helper; the stale-brain-refs scan covers setup-gbrain/sections. Regen + goldens + branch baseline + ratchet recaptured. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat(qa): carve QA patterns + health rubric into on-demand sections (68→48KB skeleton) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat(browse): carve full command list + snapshot flags into sections/command-list.md (39→27KB skeleton) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat(retro): absorb inline git/awk metrics into bin/gstack-retro-metrics + carve report format RETRO_METRICS_PROTO: 1 contract, local git reads only (fetch stays in the skill prose), degraded path documented in the skeleton. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * test: register wave-3 carves (qa, browse, retro) — guards, touchfiles, pins, baselines CARVE_GUARDS gains the three entries; qa's monolith invariant retires. auq-format carve-safety now keys on the skeleton+sections union shipping the AUQ block (first tier-1 carve: browse never renders it by design). Baselines: parity v1.69.1.0 at 18 sectioned skills; ratchet recaptured. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(test): drop stale generate-lake-intro import (generator deleted in the emission-layer move) Sol scope discipline stays pinned via the model overlay + completeness section; the lake intro is now a single script-emitted blurb. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat(office-hours): carve Phase 2A/2B into mode-exclusive sections (81→67KB skeleton) A session runs exactly one mode, so a builder session never loads the 13KB startup diagnostic. Mode mapping and the vibe-shift upgrade rule stay in the skeleton. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat(design): carve UX doctrine + Pretext patterns into read-on-demand sections design-html 57→49KB, design-shotgun 53→50KB. Sections wrap {{UX_PRINCIPLES}} so scripts/resolvers/design.ts stays the source of truth; the pretext-patterns STOP sits at the top of Step 3 so the read provably precedes the Write. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * test: register wave-4 carves (office-hours ext, design-html, design-shotgun) — 20 carved skills Both design entries carry requiredReads + loading-eval scenarios (D3A condition). office-hours phase sections are mode-exclusive, so only the always-reached design/handoff section is a deterministic requiredRead. Baselines and ratchet recaptured. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs: trim CLAUDE.md 66.4→44.9KB — verbatim moves to docs/, pointers stay inline Moved: browser/sidebar/server internals, CHANGELOG release-summary format spec, project tree, hermetic-E2E detail, slop-scan reference, OpenClaw publishing. Kept inline: every hard behavioral rule (dist/ ban, redaction scan-at-sink, egress receipts, bisect commits, eval detach, CHANGELOG entry rules), the machine-managed GBrain block (byte-identical), and the '## Deploying to the active skill' header with gbrain-refresh in range (pinned by test/gbrain-refresh-install-render.test.ts). No voice rewrites. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(test): seed onboarding markers into the hermetic child GSTACK_HOME EOV7 made bin/gstack-skill-start honor GSTACK_HOME, so the operator-HOME seeding in e2e-helpers.ts no longer reaches hermetic children — the emission layer fired lake-intro/telemetry prompts that burned turns and stalled PTY tests waiting on an answer (observed: plan-mode-no-op derailed by the telemetry question). Onboarding-specific tests pin their own GSTACK_HOME per-test, which merges over this seed. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * test: raise carve-section-loading wall clock to 480s SDK / 540s bun The heavy full-workflow scenarios satisfy their required section reads inside 60s but need 300-450s to finish the report on slower sandboxes; the 300s default read as a loading failure when the carve invariant held (traces: plan-eng-review read its section at 8s, office-hours all three at 24s, design-html both at 50s — all timed out mid-report). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(security): harden the skill-start trust boundary — review-army findings Session ID gains a urandom suffix (block binding unforgeable by reflected content); _sanitize also neutralizes spoofed SESSION_ID: lines; branch names are charset-clamped before JSON embedding (skill-start + skill-end); .brain-last-push reads first line only with a charset clamp; the artifacts URL echo routes through _sanitize; the privacy consent gate fires in interactive sessions only (spawned auto-choose could accept consent no human gave — emission order is not a safety property); the daily pull gets non-interactive + slow-network git guards and stamps only when the receipted path ran; ~/.claude.json gets a grep pre-filter before the jq parse. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(resolvers): question-log session_id becomes a substitution placeholder + stale-comment sweep The question-log block bound $_SESSION_ID, a shell variable the consolidated fence never sets — hook-less hosts logged empty session_id, breaking /plan-tune per-session grouping. It now uses the same substitute-from-the-skill-start-echoes contract as the telemetry block. Also: retired the pre-Phase-2 stop-gate docstring, repointed the gbrain-local-status cross-reference at the script's inline jq, dropped an orphaned section comment, documented retro-metrics' suffix-only census. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * chore: regenerate renders for the question-log placeholder; goldens + baselines follow Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * test: hermetic update-check, onboarding gate sequencing, seeding parity The contract test's child did a live git ls-remote + curl to github.com on every bun run test (update_check config now gates it off); the headless test gets a fresh GSTACK_HOME so the suppression is actually exercised; a new OV6 test drives the script three times to pin ack-at-emit and gate sequencing; hermetic seeding covers the config-keyed privacy gate; the EVALS_HERMETIC=0 debug seeding reaches marker parity. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * test(ci): demote the preamble A/B to periodic (OV7) and add it to the periodic matrix Post-Phase-3 demotion per the plan; the eval needs fetch-depth 0 (it git shows a pre-Phase-1 sha), which only the periodic workflow provides — and a static matrix entry so it can't silently never run. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * chore: bump version and changelog (v1.70.0.0) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs: update project documentation for v1.70.0.0 ARCHITECTURE.md: the preamble section now describes the v1.70 runtime — the rendered {{PREAMBLE}} block invokes bin/gstack-skill-start and reads STATUS lines, gstack-skill-end logs telemetry, and one-time onboarding text arrives as gated GSTACK_INSTRUCTION blocks instead of riding in every render. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs: doc-review fixes — repair moved-file links, drop unbacked session-count claim docs/BROWSER_INTERNALS.md: the two ARCHITECTURE.md anchor links broke when the section moved from repo-root CLAUDE.md into docs/ — now ../ARCHITECTURE.md. ARCHITECTURE.md: the preamble's session-tracking item claimed an active-session count and an "ELI16 mode" that no shipped code implements (the count computation was deleted with the inline preamble); describe the real touch-and-prune behavior instead. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs(changelog): correct numeric claims against measured counts 50 of 62 installed skills dropped (fixture/alias entries have no preamble); 11 new carves + a deeper office-hours carve = 9→20; test counts match the files (13 / 11 / 3 / 7). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs: repoint the preamble-runtime version reference after the queue rebump (v1.71.0.0) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * test(e2e-design): widen the Aesthetic synonym set — vocabulary variance, not a regression Both attempts in run 33090283032 produced judge-praised DESIGN.md files phrased as 'design principles'/'design language' without any of the four original literals; inputs were identical to the prior passing run 32899975845 (design-consultation untouched by the intervening merge). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(test): stage design-consultation's sections/ into the E2E fixture The skill has been carved since v1.57.0.0 — the DESIGN.md structure prescription (the AESTHETIC proposal template) lives in sections/proposal-and-preview.md behind a STOP-read. The fixture only copied SKILL.md, so the agent improvised structure from the skeleton and the section-synonym check has been a coin flip since the carve (CI run 33090283032 trace shows 'no sections dir'; the local eval store has the same failure on 2026-08-25 while that day's CI run passed on lucky vocabulary). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
…arrytan#2710) * fix(browse): never chmod shared, symlinked, or foreign-owned dirs to 0700 restrictDirectoryPermissions unconditionally chmodded its target. On hosts where the process holds CAP_FOWNER (Docker as root, CI sandboxes) that chmod SUCCEEDS on root-owned /tmp whenever a state file is configured there (BROWSE_STATE_FILE=/tmp/x.json derives stateDir=/tmp), and a 0700 /tmp breaks access(2)-based checks machine-wide for every other process. The POSIX branch now refuses shared sticky dirs, world-writable mounts under root, foreign-owned dirs, and symlinked state dirs; refusals warn once per process instead of failing silent; owned-but-unreadable dirs keep their chmod self-repair; and the check-then-act race is closed with fd-anchored O_NOFOLLOW + fstat/fchmod on a single inode. Regression tests cover the sticky-dir, foreign-uid, mkdirSecure-reapply, and symlinked-dir shapes. * fix: hash with sha256sum before shasum on Linux (config slugs + setup verify) shasum is perl/macOS; coreutils-only Linux ships sha256sum. Two call sites hard-coded shasum: gstack-config's sha8_of/sha16 (so resolve-user-slug exited 127 for any Linux user with a git email, the Layer-3 fallback) and the generated bun-installer checksum snippet in the browse/qa NEEDS_SETUP flow (spurious "checksum mismatch" on the same distros). Both now resolve sha256sum first and fall back to shasum -a 256. New shim-PATH tests pin BOTH hasher branches of sha8_of to a known vector and cover the sha8->sha16 collision escalation end to end. * feat(contract): Aside is the recommended driver for third-party web actions The Third-Party Web Actions contract (ship, spec, office-hours, land-and-deploy, setup-deploy) now names the Aside AI browser as the recommended driver: it acts across the user's real logged-in sessions, which is what vendor-dashboard moments need. Supersedes the v1.65.0.0 de-Aside stance by explicit user directive (2026-08-27). Detection is a runtime probe (command -v + aside --version under a portable gtimeout/timeout/bare guard; nonzero exit = not detected). Consent options render per detection state with Aside recommended and the first-party stack ($B headed + handoff, GStack Browser) as the universal fallback. Absent on macOS, the contract mentions the aside.com download (macOS 15+) once per task; gstack never runs an installer and binary presence is never consent. Drive discipline: step-wise over whole-task delegation, vendor confirm mode on, vendor skill/--help text scoped to operational syntax only, secrets minimized (autofill / human-used copy buttons), Apple credential creation never a drive target in any skill, failure path quotes redacted errors and falls back only with fresh consent. test/third-party-actions.test.ts pins every load-bearing sentence (21 tests) plus repo-wide tripwires: an aside command allowlist (--version/--help only, code spans AND prose) and a ban on Aside installer invocations across all generated docs. Budget ratchet fixture and carve skeleton ceilings refreshed in this commit per the ratchet protocol. * chore: regenerate remaining browse-setup snippet consumers The sha256sum-first checksum fallback in the generated NEEDS_SETUP snippet renders into every browse-consuming skill, not just browse/qa. Mechanical regen of the other ten consumers; no template changes here. * test: consent-gate E2E suite + functional fs-capability probes Five hermetic gate-tier E2E cases (tpa-present / absent-linux / broken / absent-darwin / apple-ban) drive the real contract section through claude -p with PATH shims for aside and uname; the absent cases filter any REAL aside binary out of the child PATH and assert absence with Bun.which before spawning, so dev machines cannot leak into detection. Registered per-case in E2E_TOUCHFILES/E2E_TIERS with template-level deps (ship/SKILL.md.tmpl, gen-skill-docs.ts) and added to the evals.yml matrix with tier: gate. eval:bg:periodic's detach timeout rises to 36000s for the grown periodic shard census (floor-enforced by test/eval-detach-timeout-floor.test.ts); CLAUDE.md doc updated to match. test/helpers/fs-caps.ts adds canRevokeWrites/canRevokeReads functional probes; 13 chmod-based tests swap their uid-0-only guards for the probes so suites skip honestly on CAP_DAC_OVERRIDE containers (this sandbox: uid 1000 with full caps) instead of asserting revocations the kernel ignores. path-validation's symlink test targets /etc/passwd (exists everywhere; /etc/crontab is absent on Amazon Linux). * docs: file the Aside follow-ups in TODOS Phase-2 QA logged-in-evidence path (P3), a hostile-vendor-skill E2E for the contract's override sentence (P2), and fd-anchoring the file-level permission writes to match the directory hardening (P3). * chore: bump version and changelog (v1.72.0.0) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs: update project documentation for v1.72.0.0 docs/skills.md: Third-Party Web Actions subsection under /ship (Aside recommended driver, consent rules, credential boundaries). BROWSER.md: "Aside and third-party drives" subsection under Real-browser mode + ToC entry, including the no-gstack-side-audit-trail caveat (ship adversarial finding 12). TODOS.md: mark the finding-12 doc note done. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs: apply cross-model doc review fixes for v1.72.0.0 docs/skills.md: restore the /ship closing line above the new subsection. BROWSER.md: ToC label matches the heading; BROWSE_STATE_FILE env row documents the new dir-hardening refusal + one-time warning. CHANGELOG: correct the hasher precedence wording (sha256sum first, shasum fallback) and the fs-caps count (14 test files, verified against the diff). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs: close cross-model doc-review gaps for v1.72.0.0 setup's manual bun-verify instruction gets the same sha256sum-first fallback the automated snippet got (coreutils-only Linux); BROWSER.md's BROWSE_STATE_FILE row now lists the under-root world-writable refusal; test-cost ceilings in CLAUDE.md/CONTRIBUTING.md updated for the five new gate E2E cases (~$4.20 E2E / ~$4.35 evals). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(test): gate the symlink-refusal test to POSIX and drop the umask assumption The symlink regression test exercised the POSIX O_NOFOLLOW branch but ran on Windows, where restrictDirectoryPermissions takes the icacls branch and stat has no POSIX modes (0o666 always) — windows-free-tests failed on mode 493 vs 438. Early-return on win32 like every sibling test in the file, and assert the target's mode is UNCHANGED (captured post-mkdir) instead of hardcoding 0o755, which a strict umask would also break. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
…r speed (garrytan#2721) * fix(ci): free-tests lane actually runs the make-pdf e2e gates The 9 make-pdf/test/e2e gate tests probe make-pdf/dist/pdf, browse/dist/browse, and the diagram-render bundle, then self-skip when absent. The required free-tests lane never built any of them, so the gates silently skipped on Linux for their entire life (verified: 9 of 14 skip, exit 0). make-pdf-gate.yml's justification for deleting its Linux leg claimed the free lane covered this — it didn't. - new build:gates script: exactly the three artifacts the gates probe (full bun run build compiles five binaries; ~60-90s tax on the only required check is not warranted) - free-tests.yml: build:gates step + poppler-utils + fonts-noto-color-emoji (fonts must precede the first browse daemon launch — Chromium snapshots fontconfig at startup; verified live: a warm daemon renders tofu, a fresh one embeds NotoColorEmoji) - make-pdf/test/e2e/ci-prereqs.test.ts: GSTACK_EXPECT_BINARIES=1 (set by the workflow) inverts the skip polarity in CI — dropping the build step or poppler fails the lane instead of re-opening the silent-skip hole Pre-flight: all 9 gates green on Linux locally. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(ci): kill the three zero-test eval jobs (hollow green) - delete the vestigial e2e-codex / e2e-gemini matrix rows: both files are whole-file periodic-tier, so with no row tier: they ran ZERO tests and reported green on every PR (~2 min of runner each, pure false confidence; the periodic lane owns those suites) - e2e-pty-plan-smoke gains tier: gate — its two files are whole-file describeE2ETier('gate'), so the job burned ~7 min of container setup then skipped every describe - KNOWN_TIER_UNSET burned down to empty; the ratchet stays armed so a future row/file tier mismatch fails the suite instead of shipping hollow green Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(ci): least-privilege permissions + fork-safe concurrency keys - evals.yml / evals-periodic.yml evals jobs: explicit contents:read + packages:read (container-image pull) and persist-credentials:false — the jobs that execute PR-authored code with three provider API keys ran on the repo-default token grant with the token written into .git/config - permissions blocks for the 4 workflows that had none (skill-docs, make-pdf-gate, windows-free-tests, windows-setup-e2e) - fork-safe concurrency keys: actionlint, skill-docs, make-pdf-gate, windows-setup-e2e switch from head_ref to PR-number keying — a bare branch name carries no fork prefix, so same-name branches from two forks shared one group and cancelled each other's runs Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(ci): one bun version everywhere + drift tripwire Lanes disagreed four ways: 1.3.13 (free-tests, windows, Dockerfile.ci), latest (quality-gate, make-pdf-gate), unpinned (skill-docs, version-gate — setup-bun installs latest), 1.3.10 (.gitlab-ci.yml). Different Bun versions change the runner output shapes the strict classifiers regex-match, spawn semantics, and shell parsing — a lane on a different Bun tests a different product; Dockerfile.ci's own comment records this class biting once already (silent 1.3.13/1.3.14 drift). All surfaces pinned to 1.3.13; test/bun-version-drift.test.ts scans every workflow setup-bun stanza + Dockerfile.ci + .gitlab-ci.yml and fails on any mismatch or unpinned stanza. skill-docs also gains --frozen-lockfile (was bare bun install). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * test(ci): bind the three-way image-tag hashFiles() expressions evals.yml, evals-periodic.yml, and ci-image.yml each compute the CI image tag from hashFiles('.github/docker/Dockerfile.ci', 'bun.lock', 'patches/**') — synced by comment only (TODOS.md 'CI three-way image-tag drift'). If one input list drifts, that workflow computes a different tag for the same content: eval lanes silently rebuild the image every run, or ci-image prebuilds a tag nobody looks up. The test extracts each tag-computation site and fails on any mismatch. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(ci): ci-image stops rebuilding the identical image every ship - package.json out of the trigger paths: the tag hash deliberately excludes it (version bumps every ship), so every merge rebuilt and re-pushed the IDENTICAL tag (~2m26s for zero content change); patches/** added (it IS a tag input) - manifest existence check (mirrors evals.yml): tag already exists → skip the build - concurrency group: two rapid main pushes raced pushing the same :latest/:buildcache tags - cron staggered 06:00→04:00 Monday: it shared the exact minute with evals-periodic, which could race a half-pushed tag or duplicate the build - timeout-minutes: 30 (was unbounded → 360-min default for a hung docker build) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(ci): quality-gate drops the 74s full-history checkout fetch-depth:0 cost 74 of the job's 92 seconds; the three gates it feeds take ~12s combined. Shallow checkout + exact-SHA fetches for the diff's base/head (an exact-SHA fetch, not a guessed depth — long-lived branches and merge queues still resolve), with a --deepen fallback for push events whose 'before' is unusable. timeout right-sized 20→10 min. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(ci): small-lane batch — timeouts, right-sizing, windows cache warm-start - timeout-minutes on the 6 remaining unbounded jobs (actionlint 5, skill-docs 10, version-gate 10, make-pdf-gate 15, pr-title-sync 5, evals build-image 15) — a hung step sat on GitHub's 360-min default - right-size measured-over-long timeouts: dependency-review 10→5, windows-setup-e2e 15→10 - dependency-review: 2-core runner (28s API call on an 8-core box) and drop .github/workflows/** from its trigger paths (workflow edits have no dependencies to review) - windows caches gain restore-keys: a lockfile bump paid the 26s/43s restore for a guaranteed cold miss Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(test): scope GSTACK_HOME to each file's execution window Five files assigned process.env.GSTACK_HOME at module scope. Shard processes evaluate sibling modules before running their tests, so the assignment leaked into every other file in the shard — the damage was already visible in defensive workarounds (relink.test.ts:28 'fresh install test saw a neighbor's skill_prefix'; cdp-e2e's own comment documents a sibling's temp dir baked into artifacts). Pattern: save original, assign in beforeAll, restore in afterAll (cdp-e2e already restored but still assigned at load — its window now matches the others). GSTACK_TELEMETRY_OFF and GSTACK_PROJECT_SLUG get the same treatment where they rode along. Victim files' defenses stay in place (cheap insurance). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * test: tripwire against module-scope GSTACK_HOME assignments Column-0 assignment of GSTACK_HOME / GSTACK_STATE_ROOT in any tracked *.test.ts fails with the file:line and the fix (beforeAll + afterAll restore). Kills the cross-file env-leak class the previous commit swept. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(test): e2e-harness-audit derives its skill census from disk The hand-maintained 39-name SKILL_GLOBS list had drifted to 39 of 54 SKILL.md.tmpl on disk. No live gap today (none of the 15 unlisted skills is interactive), but the next interactive skill would have landed unguarded with zero signal. The audit now walks top-level dirs for SKILL.md.tmpl (statSync so symlinked dirs like connect-chrome count), so new skills are in scope the commit they appear. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(evals): judges honor the eval-model resolution chain + real 429 backoff callJudge inlined GSTACK_EVAL_MODEL_JUDGE || sonnet, silently ignoring the global GSTACK_EVAL_MODEL override every other eval call site honors via lib/eval-model.ts. New 'judge' kind in DEFAULTS (sonnet — the D1a pin-on-regressors calibration stands; model CHOICE unchanged) and callJudge resolves through it: explicit arg > GSTACK_EVAL_MODEL_JUDGE > GSTACK_EVAL_MODEL > default. 429 handling upgraded from one fixed 1s retry (reliably lost races at CI concurrency) to three jittered exponential retries (~1s/4s/16s), honoring the server's retry-after when present. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(test): the two expect(true) paid stubs become test.todo skill-e2e-spec-execute (600s budget) and skill-llm-eval-spec (300s) reported PASS on every periodic run while asserting nothing. Deleting them would remove the periodic-tier selector surface they exist to register (diff-based selection for spec/ changes), so they become test.todo — reported as todo/skip, never pass — with the v1.1 implementation specs kept in-file. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(test): reactivate 5 quarantined browse tests (2 security) extension-sender-auth's two privileged-message denial tests (content script + missing sender.url — the extension's security boundary) and snapshot's three skips were quarantined 'pre-existing' failures. Root cause: machine-local state on the quarantining dev machines — the test and gate code are byte-identical between the quarantining commit (410b492) and HEAD, and all five pass deterministically on a clean checkout (68/68 across both files, multiple runs). No assertions weakened, no product changes. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(evals): activate the 4 paid test files that could never run anywhere carve-section-loading, codex-e2e-plan-format, codex-e2e-recommendation-substance, and llm-judge-recommendation gated on EVALS/tier (free suite loads them as describe.skip) but their names fell outside PAID_TEST_GLOBS, so no paid lane ever selected them — net execution zero, forever. The existing matrix tripwire filtered on isPaidTestFile() first, so it was blind to exactly this class (the same bug that hid the pre-split monolith's gate tests for ~8 releases). - PAID_TEST_GLOBS: codex-e2e* + skill-llm-eval* wildcards (replacing exact names) + llm-judge-recommendation + carve-section-loading; package.json's six test-script glob lists mirrored - codex-e2e-plan-format gains the explicit periodic tier gate its siblings carry (external-service rule) — without it the sharded runner's no-guard default would spawn Codex in the gate tier per PR - eval:bg:periodic --timeout 32400→37800: the census growth pushed the periodic worst case to 35910s; the old value had 270s of headroom BEFORE this change and would now kill healthy runs mid-flight - new test/paid-orphan-tripwire.test.ts: any EVALS/tier-gated test file outside the globs fails the free suite (reasoned SCANNER_EXEMPT for the gate helpers + meta-tests) — the class-killer - paid-shards pins updated: the four orphans now assert INSIDE the census Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(browse): restrictDirectoryPermissions warns and skips symlinked dirs Closes the Windows Free Tests red: recent lane failures showed a platform-unguarded POSIX mode-bit assertion ('Expected: 493' — a symlink-skip test) from PR-branch variants; the KNOWN_WINDOWS_SAFE force-include reason ('mode-bitmask hits are POSIX-branch only') did not hold for that shape, and main had neither the guard nor the behavior. - product: lstat first; a symlinked dir gets a warning and a skip on both platforms — chmod AND icacls dereference the link, so restricting through a symlink hardens an unvetted target (and /inheritance:r could lock out its real owner). All callers already treat hardening as best-effort (try/catch). - test: the symlink regression test, platform-aware — symlinkSync in the house try/catch skip pattern (Windows runners without Developer Mode can't create symlinks), mode-bit assertion guarded off win32, behavior assertions (no throw, warning text, target readable) everywhere; POSIX still proves the skip (0o755 unchanged, not 0o700) - KNOWN_WINDOWS_SAFE reason updated to the now-true premise 20/20 pass on Linux. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(test): unique tmp dirs for plan artifacts + audited live-repo cwd sites Six paid PTY tests wrote their expected plan artifact to a FIXED shared /tmp path ('/tmp/gstack-test-plan-<mode>.md') and rmSync'd it in finally — under --retry 1, EVALS_JOBS>1, or two concurrent worktrees, a sibling's cleanup deletes this run's artifact and the D19 'agent did not produce expected plan file' assertion fires spuriously. Each test now mkdtemps its own dir, interpolates the unique path into the agent prompt (fixture-sourced prompts get a replaceAll + drift guard that throws if the fixture's literal ever moves), and cleans up its own dir. The 18 cwd:-into-the-live-repo sites were audited: all deliberate (skill registry + hermetic pre-trusted dir, in-repo gen renders, git history reads, slug resolution) — each now carries a '// LIVE-REPO CWD: <reason>' comment so the next audit can tell deliberate from accidental. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(test): trim the seven over-wall 1700s timeouts to the 1500s physical ceiling 1,700,000ms (28.3 min) exceeded every wall these tests run inside: the 25-min CI job timeout and the 1800s sharded-runner wall (which also leaves --retry 1 zero room for a second attempt). Budget above the wall is fiction, not headroom — a test that actually used it produced a job-level kill (no bun summary, no artifact) instead of a clean per-test timeout. No recorded p95 exists for this family (they are being retiered to periodic in the re-platform wave); the trim stops at the physical ceiling rather than guessing lower. Final policy lands in the Wave-2 eval-budgets constants module. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * refactor(gen): main() guard — importing gen-skill-docs no longer regenerates the tree The generator's whole body executed at module load, so any import of it (test/gen-skill-docs.test.ts pulls assertSinglePreamble via require(); test/catalog-trim.test.ts imports helpers) regenerated all 71 SKILL.md in place — the root cause of half the TREE_MUTATING serial-shard entries (hazard class garrytan#2532). The body now lives in an exported main(): number behind if (import.meta.main). Semantics preserved exactly: failure exits are immediate (matching the old top-level process.exit), success leaves the event loop to drain so the llms.txt fire-and-forget IIFE finishes its write, and the module stays synchronous/require()-able. Proofs: byte-identical --host all output (git status clean), --dry-run stale-tree still exits 1 (the skill-docs freshness lane depends on it), and the new test/gen-skill-docs-import-purity.test.ts pins load-time purity via a subprocess probe (mtime-based, so a dirty worktree can't false-fail). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat(gen): --out-dir renders every host, outputs-only --out-dir was Claude-host-only (gen-skill-docs.ts:842), which forced the codex/factory-regenerating tests (gen-skill-docs, skill-validation, host-config) to mutate the live tree — the reason they sit in the TREE_MUTATING serial shard. The flag now mirrors ALL outputs into the out-dir: external-host trees (.agents/.factory/... via processExternalHost), external section files, openclaw docs, and gstack/llms.txt (a catalog-mode render must never rewrite the tracked index). OUTPUTS ONLY — inputs (templates, sections/, host configs) are always read from ROOT, so an empty out-dir can never feed the render. rewriteSectionBase stays Claude-only (external hosts have their own path grammar). Proofs: in-place --host all is byte-identical (tree clean); --host all --out-dir <mkdtemp> renders the full multi-host tree with ROOT untouched; gen-skill-docs-out-dir tests + 415/415 gen-skill-docs.test.ts green (bin/dev-setup's claude rendering byte-compat). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(evals): every E2E key's dep list names its own declaring test file 129-of-177 keys omitted their own test file, so editing only a test's prompt or assertions selected NOTHING — the changed test never ran on the change that changed it. 135 keys self-registered (110 E2E + 25 LLM-judge), resolved by strict declaration evidence (testName:/ testIfSelected/judge call sites), with skill-name false positives excluded. e2e-tier-alignment's warn-only branch for unregistered files is now a hard failure with a 4-entry KNOWN_UNREGISTERED ratchet (template- literal testNames, fail-open-safe) + a burn-down test so the set only shrinks. Selection sanity: a one-file diff on skill-e2e-qa-workflow now selects its 4 tests (was 0); skill-llm-eval 0 → 25. Known follow-ups (filed): 15 E2E + 2 judge PHANTOM keys select tests that exist nowhere; codex-e2e-plan-format's testIfSelected names have no map keys (run-all only). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * test(evals): ratchet the 8 newly-visible gate-matrix gaps The self-registration sweep made these eight files' gate-tier keys visible to the census for the first time — their gate tests run in NO CI lane today (pre-existing hole, newly measurable). Ratcheted into KNOWN_MATRIX_GAPS with the burn-down note: the paid-lane re-platform runs every gate file by construction and retires this ratchet class. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat(test): duration-aware LPT shard packing for the free suite Hash sharding balances file COUNTS (1.15x spread) but not cost — the Playwright-launching files landed 4/3/4/1/2/1 across 6 shards, giving a measured 28s–97s shard spread and ~40s of idle tail on every run. Full-suite mode now packs by recorded per-file durations (longest-processing-time-first) when the committed seed scripts/free-test-durations.json exists. - ONE store, no overlay: the seed is refreshed occasionally via the new --record-durations mode (each file timed in its own child — exact, and immune to bun's stream buffering, where silent passers print no header to timestamp); GSTACK_FREE_TEST_DURATIONS overrides the path for experiments; CI never records - seed is a hint: missing → silent hash-shard fallback; corrupt (bad merge) → one warning + fallback; unknown files → 75th-percentile pessimism so a surprise long-runner can't recreate the tail - packed shards get duration-aware walls (max(base, predicted x 3)) — LPT decouples count from cost BY DESIGN, so the 5s/file heuristic would undersize a shard holding few expensive files - one log line per shard (files + predicted seconds) so packing regressions are diagnosable from any run log - the --shard CI-matrix path is untouched: stable hash indices are its contract - successor note in-code: bun >=1.3.14 ships native --timings/--shard LPT — swap this packer when the repo unpins 1.3.13 Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(test): decouple slop:diff from bun run test; quality-gate runs it per PR 'bun run test' silently appended up to two 120s npx slop-scan runs plus a git worktree add/remove after the suite (2>/dev/null || true) — invisible in the documented '~90-100s' timing and pure friction in the pre-commit loop. Decoupling is not coverage removal: quality-gate.yml now runs slop:diff on every PR (advisory, matching its in-repo 'never blocking' contract), and /review already invokes it explicitly. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat(test): eval-budgets timeout tiers + fit/ceiling policy test Five named tiers (JUDGE 120s / CAPTURE 300s / CAPTURE_LONG 600s / PTY 900s / PTY_LONG 1200s) replace hand-ratcheted sprawl (46x300s, 46x120s, 44x360s, 44x180s, 27x240s, 19x150s, 13x420s, 12x600s...), much of it inflated to paper over the old 40-way in-shard concurrency that the sharded runner's 1-file-per-shard model kills. Policy test pins: every tier fits the shard wall minus 120s overhead (the structural fix for budgets-above-the-wall fiction), tiers stay ordered, and no paid literal exceeds PTY_LONG x1.25 — oversized tests get split, not budgeted past the wall. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat(test): shared runBin helper for bin-script unit tests ~36 free test files each carry a near-identical local run() (spawnSync + utf-8 + {status, stdout, stderr}) differing only in env composition, cwd, and timeout. runBin absorbs the invariant core; options carry the variance (gstackHome sets BOTH GSTACK_HOME and GSTACK_STATE_DIR — the config-precedence trap several locals rediscovered independently; home for $HOME-anchored bins; input/trim/timeout/maxBuffer). Free-test-only by design so it never becomes a de facto global touchfile. Migration of the 36 call sites lands separately (mechanical batches). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(test): runBin trim assertion — trim shapes stream ends, not interior Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * refactor(test): mechanical sweep — 298 paid-test timeouts onto eval-budget tiers 69 files, both shapes (trailing bun-test budgets and runner timeout/timeoutMs options), ROUND-UP ONLY so nothing that passed can start failing: 75 → JUDGE_MS, 137 → CAPTURE_MS, 74 → CAPTURE_LONG_MS, 9 → PTY_MS, 3 → PTY_LONG_MS. Raw >=60s literal count in the paid scope: 395 → 97, of which 51 are non-timeout noise (fixture dates, run IDs) and 46 are enumerated justified holds (comment-carrying calibrated budgets, poll-loop constants, utility spawn waits, and the seven physical-ceiling 1_500_000 sites). The eval-budgets policy ratchet keeps the residue from regrowing. Known collapse: where an inner runner budget and its enclosing test budget now share a tier, the old stagger is gone — an overrun surfaces as a bun test timeout instead of a graceful runner timeout (diagnosability trade, not a correctness one). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * test: coverage fill — 95 tests for six zero-coverage surfaces - eval CLI family (eval-list/compare/summary + eval-select smoke): the primary interface to eval results had no tests; isolation via a fake gstack-slug under a mkdtemp HOME (the scripts' real resolution path — they do NOT honor GSTACK_EVAL_DIR; only EvalCollector does). Pinned current behavior: eval-list does NOT exclude _partial runs (documented improvement candidate) - slop-diff (runs on every /review + quality-gate): fixture git repo + first-on-PATH npx stub (never downloads real slop-scan); no-diff early exit, missing-scanner fallback, fingerprint line-insensitivity, merge-base worktree scan - bin/gstack-code-intelligence CLI arg surface (lib was covered, the 284-line CLI wasn't): select/consent/suggest/index/search gating; pinned: --help routes to usage failure exit 1 (no handler) - browse media-extract: the page.evaluate callback exercised in-process against a mock DOM (no exports added) — lazy-src fallback chain, HLS/DASH detection, bg-image url() parsing, 500-element cap - browse session-cookie-store: factory contract (cookieName/ttlMs/ maxSessions eviction, cross-store isolation, mint→validate round-trip); store is in-memory — no fs cases exist - lib/version-source direct unit tests (gstack-version-bump.test.ts spawns the bin, never imports the lib): parse/format/cmp/bump coercion, npm 4→3 translation, garrytan#2501 mangled-JSON regression class All hermetic (mkdtemp homes, runBin child isolation); windows curation correctly partitions the six. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * refactor(test): first runBin migration batch (3 of ~36 run() duplicates) explain-level-config, benchmark-cli, evidence move onto the shared helper; each file's remaining special-case spawnSync sites (raw-buffer probes, env-scrub probes) stay put deliberately. 55/55 green. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * refactor(evals): paid shards spool to disk + shared runShardChild lifecycle - runPaidShard no longer buffers whole 30-min stream-json streams in RAM (x concurrent jobs): every byte tees to a per-shard log file (slug-named, path printed at START for mid-run inspection and on the FAILED terminal line); failures print a 64KiB tail read back from disk; passing shards stay quiet (the file is the record) — the free runner's proven contract. Classification unchanged: the strict classifier still sees every byte first. - the ~35 duplicated spawn/group-kill/wall-timer/finally-reap lines move into runShardChild in test-strict-output.ts (detached-per- platform spawn, signal forwarding, SIGKILL group kill at the wall, drain-before-verdict); designed so the free runner can migrate later - expectedFiles drift fixed toward ENFORCEMENT: the injected-command exemption is gone — a fake command exiting 0 without bun's terminal summary now reads FAILED (pinned: silent-pass → failed) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat(evals): parent-computed selection propagates to shard children The sharded runner computed diff selection once, then each of its 48-73 children recomputed it at module load — including, on touchfiles-diff branches, a per-child bun subprocess evaluating the old data file (20s timeout each). The parent now serializes {version, selected, reason} as EVALS_SELECTION_JSON into the shard env; e2e-helpers adopts it at load. Fail-open preserved: any parse/shape violation → ONE stderr warning + local recompute; absent env → silent local compute (non-sharded entrypoints unchanged). Drift test pins parent→child round-trip to identical selection decisions plus the malformed/absent cases. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(test): kill the four worst fixed sleeps (300s/30s/30s/20s) - watchdog.test: the 20s blind wait for one production parent-watchdog tick becomes BROWSE_PARENT_WATCHDOG_INTERVAL_MS=250 (new env knob in server.ts, NaN-safe, production default unchanged) + polls for the boot line and the tick's stay-alive log — strictly stronger (the old form never proved a tick observed the parent death). 24s → 3.6s. - stop-dead-daemon / terminal-agent-owner-watchdog: the 300s/30s stand-in child lifetimes become stdin-EOF-bound — the child can never self-exit mid-test on a slow runner (spurious-failure class) and self-reaps instantly if the test dies (no 300s orphans). Node-compat stdin APIs (owner-watchdog runs on the Windows lane). - browser-skill-commands: the sleeper fixture's 30s self-time becomes 8s (no stdin pipe exists in runToFiles) — far above the 1s product timeout it must outlive, below the test ceiling, so a timeout-kill regression fails on clean assertions instead of an opaque bun timeout; added: stdout must NOT contain 'done'. 45/45 green across the four files + server tripwires. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(test): gen-skill-docs + catalog-trim leave the serial mutator shard gen-skill-docs.test.ts's 15 in-place generator spawns now render into mkdtemp out-dirs (gitignored-artifact reads repointed; the handshake scan's silent console.warn degrade became a hard assertion); its tracked-tree reads (freshness dry-run, SKILL.md content pins) stay reads. catalog-trim needed no change beyond the earlier main() guard — its import is now side-effect-free (pinned by the import-purity test). Both TREE_MUTATING entries deleted in this commit, per the transition rule: an entry leaves in the same commit as the file's last in-place write. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(test): skill-validation renders codex host into an out-dir Its 3 in-place --host codex regeneration sites collapse into one module-level --out-dir render; assertions untouched. TREE_MUTATING entry deleted. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(test): host-config self-provisions goldens (ordering dependency severed) Its goldens were 'produced by gen-skill-docs.test.ts' with a when-missing beforeAll fallback that wrote the live tree — an inter-test ordering dependency the serial shard hid. It now renders codex+factory UNCONDITIONALLY into its own out-dir and reads goldens only from there (the Claude golden deliberately keeps reading tracked ship/SKILL.md — a read; out-dir claude renders repoint section-base paths by design). TREE_MUTATING entry deleted. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(test): gbrain-detection-override drops mutate-then-git-restore regenAndSnapshot renders --host claude --out-dir <mkdtemp> (+ --respect-detection) and snapshots probes from the out-dir. The git-restore machinery is deleted outright — it restored only PROBE_FILES of the 71 files each call wrote, so a stale tree kept the other 68 dirty (the partial-restore bug), and its 'no output-path arg' comment had been false since --out-dir landed. TREE_MUTATING entry deleted. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(test): catalog-mode-full renders to out-dir; restore machinery deleted The full-catalog smoke no longer rewrites all 71 SKILL.md then regenerates to restore (with its 'CRITICAL: failed to restore' prayer path) — it renders into a mkdtemp and additionally asserts tracked ship/SKILL.md is byte-unchanged. TREE_MUTATING entry deleted. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(test): idempotency proof strengthens to two-out-dir recursive diff Two renders into two separate out-dirs, EVERY file diffed byte-for-byte (claude-only and --host all; normalization only for each dir's own sanctioned section-base repoint; presence-sanity lists guard against a vacuous empty-dir pass) — strictly stronger than the old in-place double-regen that sampled 5 files. TREE_MUTATING entry deleted. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(test): spec-template-sync compares an out-dir render, not an in-place one TREE_MUTATING entry deleted. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat(test): the serial tree-mutating shard dissolves — TREE_MUTATING is empty Zero mutators remain (all eight render into out-dirs now), so the four ratchet READERS (parity caps, size budgets, carve parity/ordering) get a quiet tree by construction in any shard and rejoin the parallel phase. The ~35-40s serial tail on every full-suite run is gone. The mechanism stays: a future test that genuinely must write shared artifacts in place earns an entry with a reason and is serialized again; the census pin still fails on renamed keys. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * test(gen): out-dir byte-identity + tree-clean pins for external hosts codex render: porcelain unchanged AND out-dir gstack-ship/SKILL.md byte-identical to a fresh in-place render (+openai.yaml presence); --host all render: exit 0, porcelain unchanged, claude + .agents + .factory + llms.txt + openclaw docs all present in the out-dir. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat(test): commit the initial free-test durations seed (496 files) Recorded via --record-durations on a quiescent tree: 479s serial total, p50 92ms / p90 1.8s / max 31.4s — the top-heavy cost shape LPT packing exists for. A hint, not a contract: refresh opportunistically with bun run test:free --record-durations. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat(evals): planner/executor/report modes — the CI re-platform surface One PLANNER computes diff selection + the slice plan ONCE and writes a manifest (--emit-plan <path> --slices K); K executors consume it (--plan <path> --slice i), never self-selecting, and write slice-result artifacts; a REPORT reconciles results against the manifest (--report <dir>) fail-closed: a slice whose artifact never landed is a FAILURE, a planned shard nobody reported fails, wrong-slice/duplicate/cross-tier results fail. Kills per-slice selector divergence and hollow-lane aggregation at the root. - hollow-shard guard: under EVALS_ALL, exit 0 with ZERO executed tests (bun's 'Ran N tests' now captured by the classifier — additive) is 'passed-empty' and fails the run; selective runs keep it 'passed' with one warning (in-file diff/tier self-skips are legitimate there); unknown counts are never guessed hollow - retry parity: --retry 1 default + RETRY_OVERRIDES literals for the three files whose old matrix rows earned retries: 2 (stale entries pinned against disk) - live smoke: gate plan = 48 shards across 6 slices; report mode exits 1 on a fabricated missing slice, 0 when complete Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat(ci): sliced paid lane (planner -> 6 executors -> fail-closed report) The parity-phase re-platform: evals.yml gains a second, sliced lane driven by scripts/test-paid-shards.ts — the SAME engine local eval:bg:gate uses, so CI and local share one selection engine. - plan-slices: ONE planner (fetch-depth 0 — the only job needing history) emits the manifest; selection fails open to run-all, never per-slice (the divergence class is structurally dead) - eval-slices: 6-way matrix consuming the manifest; PTY seed + skill-registration steps run unconditionally (idempotent — a sliced lane cannot key them on suite names); aggregate spawn budget 6 x EVALS_JOBS=2 x EVALS_CONCURRENCY=2 = 24 lane-wide (the matrix's 40-way per row queued session startup behind 39 siblings — the timeout-flake family root); slice results + spooled shard logs uploaded as artifacts - slices-report: reconciles slice artifacts against the manifest FAIL-CLOSED via --report — a slice whose artifact never landed, or a planned shard nobody reported, is a failure, not an absence - sequenced needs: evals so provider concurrency never doubles while both lanes coexist; the matrix + its ratchets are deleted after demonstrated parity (intersection + expected-additions comparison) - workflow_dispatch gains evals_all (default true) for parity runs and post-merge smokes — a dispatch can never silently select zero Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat(ci): weekly periodic lane runs EVERY periodic test + gate census backstop evals-periodic.yml re-platforms onto the sharded runner: planner manifest → 6 executor slices → FAIL-CLOSED report. This IS the coverage contract: all ~70 periodic-tier files weekly (EVALS_ALL=1), killing the silent-rot class where a hard-coded 9-file matrix left ~57 files running NOWHERE (the autoplan E2E rotted invisibly for months). - test/helpers/periodic-exclude-data.ts: reasoned exclusions in their OWN literals file (deliberately not touchfiles-data — map-diff evaluates old versions of that file standalone). Every entry carries reason + tracking with a re-entry condition; the runner surfaces each exclusion per run; policy test pins real-file + non-empty fields. Initial: ship-idempotency + brain-privacy-gate (documented-red, never green) and skill-e2e-ios (manual hardware). The TODOS 'sidebar E2E trio' turned out already deleted — only tombstone tests remain. - gate-census job: weekly EVALS_ALL gate-tier run — PR lanes are diff-billed, so without this the full gate census might never execute anywhere; with the hollow-shard guard it is a census-health check (exit 0 + zero executed tests fails), not just a test run. - failure notification is a concrete gh issue UPSERT (one tracking issue, commented per red week — never issue-per-week spam), with issues:write scoped to the report job. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs: TESTING_INTERNALS covers the 2026-08 runner overhaul LPT-packed free suite + --record-durations, the emptied TREE_MUTATING mechanism, the sharded paid runner as the single selection engine, CI planner/executor/report with the fail-closed report and hollow-shard guard, the weekly coverage contract + exclusions policy, and the eval-budgets timeout tiers. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs(CLAUDE.md): testing prose matches the overhauled runners - bun run test: duration-packed shards + --record-durations; the trailing serial tree-mutating shard no longer exists - two-tier system: the sliced CI lanes (one engine local+CI), the weekly all-periodic coverage contract + exclusions, the gate census - periodic detach timeout 32400 → 37800 Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs(TODOS): close the absorbed test-infra items, file the overhaul follow-ups Closed with receipts: the periodic coverage contract (implemented as full weekly coverage + exclusions), the eval-harness observability P1 (verified already landed: heartbeat, incremental _partial persistence, live stderr + eval-watch), and the sidebar trio (already deleted — tombstones remain). Filed: matrix deletion after parity, the required-check maintainer decision, browse /tmp-namespace hardening, PTY boot-readiness waits, the single typed test registry, bun-native LPT swap, runBin/free-runner migrations, eval-list partial exclusion, phantom key cleanup, duration-weighted slicing. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * v1.73.0.0: test/CI overhaul — green means green, suites restructured for speed Version + release notes for the audit-and-overhaul branch: every silently-skipping or never-running test class fixed and tripwired, the free suite duration-packed with the serial mutator shard dissolved, the paid lane re-platformed onto the sharded runner (planner/slices/ fail-closed report, parity phase), the weekly all-periodic coverage contract, eval-budget timeout tiers, and 95 new coverage tests. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(ci): first-live-run fixes — executor history + two environment-blind assertions The sliced lane's first run (PR garrytan#2721) did its job: the planner and report worked, the manifest governed, and every failure had a name. Three were fixable on the spot: - executor + gate-census checkouts get fetch-depth: 0 — files with SELF-derived selection (the LLM-judge map, routing) walk git at module load, and selection is deliberately fail-closed on git errors, so the shallow checkout crashed those shards ('ambiguous argument main...HEAD'). The manifest still governs WHICH shards run. - landscape --toc gate: the exact toBe(3) landscape-page count was font-metric-dependent (3 on Amazon Linux, 2 on ubuntu CI — the same disease the file's own page-index comment warns about). Now a comparative invariant: --toc must not CHANGE the landscape count vs a baseline render. - paid-run-manifest parse test builds its manifest under EVALS_ALL so it never walks git (proven with GIT_DIR=/nonexistent). Remaining first-run failures are newly-exposed rot in gate files that had never executed in CI (skillify D1 refusal, session-intelligence context-restore, one tpa-apple-ban retry flake) — being probed separately; they are the lane WORKING, not the lane failing. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs(TODOS): file the three first-execution findings from the sliced lane's live run Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * v1.74.0.0: queue-advance — garrytan#2722 claims the v1.73.0.0 slot The version gate caught a live queue collision (its whole job); same MINOR bump level, next free slot per bin/gstack-next-version. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(test): per-shard CHROMIUM_PROFILE — the collision class duration packing exposed Nine test files launch in-process persistent contexts or daemons that default to the SHARED ~/.gstack/chromium-profile. Two concurrent shard processes on one profile dir kill each other's browser — observed live on CI once duration packing recomposed shards: handoff's launchPersistentContext died 'Target page, context or browser has been closed' (--user-data-dir=~/.gstack/chromium-profile in the call log) while a sibling shard's daemon logged 'Chromium process crashed'. Hash sharding had masked the collision by chance placement; handoff passes standalone everywhere. Fix at the runner, not per file: each shard child gets CHROMIUM_PROFILE=<shard-state>/chromium-profile (the documented env knob, same isolation idea as the existing per-shard TMPDIR). Files within a shard run serially, so sharing the per-shard profile is safe; config.test's resolution-order tests save/restore the env around their assertions. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(test): landscape --toc gate asserts promotion PRESENCE, not counts Two rounds of CI receipts: the exact toBe(3) was font-metric-coupled (3 on Amazon Linux, 2 on ubuntu), and the baseline-comparison repair then failed 2-vs-3 across renders SECONDS apart in one CI job while the sibling no-toc test saw 3 — per-render image-promotion timing makes any count assertion here a coin flip. The sibling test owns exact promotion counts; this test's actual invariant is that --toc does not break the promotion machinery: >=1 landscape page + the TOC rendered. Also drops the second render (halves the test's runtime). Flaky per-render image promotion itself is worth its own look — noted in TODOS with these receipts. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs(TODOS): file the per-render image-promotion nondeterminism (receipts from PR garrytan#2721) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(test): per-FILE Chromium profiles for the nine in-process launcher files Completes the profile-isolation work: the per-shard CHROMIUM_PROFILE stopped cross-shard kills; these nine files launch in-process persistent contexts and could still collide with a lingering daemon a sibling file spawned on the SAME shard profile. Each now scopes a mkdtemp profile via beforeAll/afterAll (the module-scope-tripwire-safe pattern), cleaned up per file. All nine green solo and in combined runs, except the pre-existing commands+snapshot pairing — proven identical WITH and WITHOUT these edits (baseline receipts) — which is the daemon-lifecycle follow-up now extended in TODOS with this session's receipts. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(browse): Chromium-crash exit is daemon-only — embedded launches never kill their host handleChromiumDisconnect unconditionally process.exit()ed. Correct for the standalone daemon (its supervisor/user must notice); suicidal when a TEST launches BrowserManager in-process: a mid-suite Chromium death exited the whole bun shard with no terminal summary — the exact truncation class the strict runner flags (observed live: CI shard 1 on eb23329 died at cache-concurrent-refresh right after a daemon-spawning gate test; with this fix the same pairing runs to completion and REPORTS instead of dying). The standalone entrypoint opts in via markDaemonProcess() under server.ts's import.meta.main gate — the same embedder contract its signal handlers already use (gbrowser phoenix keeps its own handlers). Embedded contexts now get the disconnect log line and continue. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(test): context-restore assertion is evidence-based, not prose-matching The test failed twice per run in TWO CI cycles while passing locally 4/4: the prompt said 'present the content' and the check grepped the FINAL message for exact phrases — local runs quoted the file, CI runs paraphrased ('the most recent context is from branch-b...') and the substring check lost the coin flip. - prompt now demands machine-checkable output: the newest file's '## Working on:' heading VERBATIM + a literal 'RESTORED: <filename>' marker (the mtime-scramble and cross-branch subject matter untouched) - assertion ordered strongest-first: RESTORED marker → legacy content phrases → tool-call corroboration (Read/Bash input naming the newer file, credited ONLY when the older file was never read — a both-files run must still present the right one) - the older-file negative got STRONGER: an explicit RESTORED marker naming the older file fails even if wintermute words appear elsewhere - sibling scan: context-recovery-artifacts got the additive prompt-side treatment only (quote the matched literals verbatim); its lenient 1-of-6 assertion deliberately unchanged 3/3 consecutive local green with all evidence classes firing (marker=true, content=true, toolNewer=true, toolOlder=false). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(test): skillify family — HOME==cwd broke project-skill registration Root cause (forensically pinned from stream-json init events + a kill-after-init probe): with HOME set EQUAL to the child's cwd, claude resolves <cwd>/.claude/skills as the PERSONAL skills directory and the seeded project-tier skills never register — the Skill tool returned 'Unknown skill'. The provenance-refusal test then improvised a refusal whose wording missed the regex (the deterministic CI+local red); the happy-path and approval-reject siblings passed only because their agents self-recovered by Reading SKILL.md manually — silently not exercising the Skill-tool path at all. All three tests now use HOME=<workDir>/home (a fresh subdir keeps the override's intent: child ~/.gstack writes land in the assertable sandbox, without the cwd collision). Refusal test additionally: a 'not registered/unknown skill' tripwire (a not-loaded skill can never pass as a refusal) and the refusal regex now matches assistant text only — the skill BODY echoed into the transcript contains the exact refusal message, so the old full-surface match could pass vacuously once the skill loaded. Sibling disk assertions sweep both $HOME/.gstack and cwd .gstack roots (positives and negatives). Verified paid: refusal 2x consecutive green with the skill's EXACT message rendered ('Launching skill: skillify' in-transcript), then the full file 5/5 green (~$1.35) with both siblings driving real Skill calls (25-27 turns each). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs(TODOS): two of three first-execution findings fixed (skillify family, context-restore) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(test): context-restore gets a private home — the REAL root cause was fixture sharing The evidence-based assertion fix was treating a symptom. The slice artifact's embedded transcript showed the CI agent restoring 20260829-context-save-skill-test.md — the checkpoint the SIBLING context-save test wrote into the SHARED gstackHome checkpoints dir, which by filename-prefix ordering genuinely IS the newest. The agent behaved CORRECTLY; the test's fixture set was open to concurrent sibling writes, and bun --concurrent ordering differs between CI (save finished first) and local (restore listed first) — the entire local-green/CI-red split explained. The restore test now uses its own .gstack-restore-home (the whole home moves, not just the handed path — an agent deriving the dir from GSTACK_HOME/projects/<slug> must land in the closed set too). Full file 4/4 paid green with all evidence flags firing. Also: the on-failure shard-log artifact glob uploaded nothing — the Fix-bun-temp step points TMPDIR at /home/runner/.cache, so the spool lands there, not /tmp. Both eval workflows now glob both locations (this gap is why diagnosing THIS failure required digging transcripts out of the slice-results artifact). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(evidence): carry the real index mtime onto gstack-wtree's temp copy The stat-cache seed (cp of the real index) stamped the temp index "now", which defeats git's racy-git protection: an entry is only re-hashed when its cached mtime is not older than the index file itself, so a same-size rewrite landing in the same second as the last real index write looked non-racy, kept its stale stat-cache entry, and vanished from the fingerprint — evidence stayed FRESH after a source change. This is the CI flake in test/evidence.test.ts "allow-paths carve-out" (sub-second alignment on fast runners: expected STALE exit 1, got FRESH exit 0). touch -r restores the original index timestamp, reinstating the exact racy window git itself uses. Deterministic regression pin in test/review-log.test.ts reproduces the miss with pinned zero-nsec timestamps (fails on the old script, passes now); receipts: manual probe shows the fresh-stamped copy returning the clean tree for a same-size 'hello'→'howdy' rewrite while the mtime-carried copy detects it. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(test): landscape gate bounds the promotion count instead of pinning 3 The alt-hinted image promotion rides the per-render measurement race already filed in TODOS (2-vs-3 landscape pages on renders seconds apart — CI receipts from PR garrytan#2721, now reproduced locally). Pin the two deterministic promotions as the floor and the three promotable blocks as the ceiling (anything above 3 means the veto leaked); the veto/portrait assertions remain exact. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com> Co-authored-by: Test <test@test.com>
…m benchmark, reuse ladder, instruction-tier digest (garrytan#2722) * feat(autoplan): eng review always runs last — the gate reviews the final amended plan Reorder the pipeline to CEO -> Design (if UI scope) -> DX (if developer-facing scope) -> Eng. The old order (CEO -> Design -> Eng -> DX) let DX findings land AFTER the required gate signed off, so eng validated a stale plan. Accept-all semantics made explicit: every AskUserQuestion resolves to the recommended option; premises no longer pause the pipeline mid-run (clearly-wrong ones queue as User-Challenge items at the single Final Approval Gate). Eng's Codex voice now sees the DX consensus summary. New free static test pins the order; the chain E2E gains DX-between and Eng-terminal assertions. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat(review): simplification specialist — advisory over-engineering lens with ponytail's tag vocabulary New 8th Review Army specialist (DIFF_LINES > 100, --simplification force flag) hunting unrequested STRUCTURE only: delete/stdlib/native/speculative/shrink closed tags, one-line findings, lines_removable field. speculative: replaces ponytail's yagni: tag — we import the lens, not the posture; coverage stays sacred (Completeness Gaps owns it, suppressions inlined, shrink needs >=5 lines). Advisory carve-out in the merge step: advisory findings are excluded from quality_score and the findings-count header, render with an [ADVISORY] label, and are ASK-only in Fix-First. Zero-findings case prints the lens-scoped 'Simplification: lean already — nothing to cut.' from the PARENT (the specialist keeps the exact NO FINDINGS contract); with findings, the parent prints 'net: -N lines possible' summed from lines_removable. Tests: static pins for the carve-out + early-out contract (gen-skill-docs), two periodic e2e cases with planted fixtures — activation (over-build traps: hand-rolled Intl, one-impl abstract, dead config) and false-flag precision (a lean ETHOS 'choose A' diff must yield NO FINDINGS). Inspired by dietrichgebert/ponytail's /ponytail-review. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat(preamble): reuse ladder in Search Before Building — rungs 2-5 of ponytail's ladder, completeness kept Tier-3+ skills gain a per-edit reflex the section only stated as research discipline: before writing new code, stop at the first rung that holds — repo helper, stdlib, native platform feature, installed dependency — then build the COMPLETE version of what remains. The closing clause is the explicit reconciliation with Boil the Ocean: the ladder governs structure, never coverage. Rungs 1/6/7 (YAGNI / one line / minimum that works) are deliberately NOT imported. Also ports ponytail's root-cause rule: one guard in the shared function beats a guard in every caller. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat(preamble): bounded-closer output rule for tier-2+ skills After completing work, skills report in a few short lines — what changed, what was skipped, what to watch — and cut any explanation that outgrows the change. Explicit exemptions protect every mandated output: decision briefs, completion-status blocks, user-requested explanations, and report-shaped skills' report formats (the report IS the work in /qa-only, /plan-*-review, /retro, /document-generate). Rationale is signal-to-noise, not tokens: ponytail's own benchmark shows terse prose alone doesn't cut cost (caveman arm: -20% LOC, +7% tokens), and independent replications found its 'skipped on purpose' essays ate the code savings. Includes a good/bad closer example pair per the model-overlay guidance that a positive example beats a 'don't be verbose' instruction. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs(resolvers): terse-mode savings claim matches measurement — 2.6KB, not 3-5KB Measured on the v1.71 render: --explain-level=terse saves exactly 2,611 bytes per tier-2+ skill. The old ~3-5KB claim predated the preamble restructuring. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat(retro,preamble): gstack-shortcut debt ledger — accepted shortcuts leave a joined trail When the user accepts an option that is BOTH Completeness <= 7 AND a durable-scope call, the decision ledger entry (gstack-decision-log, ceiling + upgrade trigger in the rationale) is the source of truth, and the agent marks each cut corner in code with gstack-shortcut(dec-<id>): <ceiling>, upgrade when <trigger> — same edit, no follow-up question, never agent-initiated. /retro Step 11.5 harvests markers into a debt ledger (grep || true — zero matches is the healthy case; skill installs and docs excluded), joins on the decision id so nothing double-counts, tags unlinked and no-trigger rot risks, and closes with 'N markers, M with no trigger.' /review suppressions: a marker with ceiling+trigger downgrades a would-be Completeness Gaps finding to acknowledged debt. Redaction test pins that the marker ships untouched (the ledger is the point) — it does not match the TODO(owner) hygiene shape. Format from dietrichgebert/ponytail's ponytail-debt; store inverted to gstack's existing decision ledger. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * test: refresh golden ship baselines after preamble additions (reuse ladder + bounded closer) The golden-file regression test pins the rendered ship skill byte-for-byte; the WS3/WS7 preamble sections are deliberate changes, so the baselines re-capture per the goldens' own update protocol. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat(hosts): instruction-only tier — a 2KB committed rules digest any agent host can read New agents-digest/gstack-AGENTS.md (1,765 bytes, hard 2,048-byte budget): gstack's ethos one-liners, the reuse ladder, and voice rules for hosts with no install arm — Zed, Amp, Jules, or any AGENTS.md-reading agent. Generated by scripts/gen-agents-digest.ts, auto-refreshed by gen:skill-docs, committed like llms.txt so setup's explainer arms can point at it before any toolchain exists. First line carries the gstack version as its own staleness nudge. Delivery is print-path + user-performed copy ONLY: setup never writes or overwrites a user's AGENTS.md (a test pins this — no cp/ln/mv/redirect into AGENTS.md anywhere in setup). openclaw and hermes explainer arms print the path; slate keeps routing to the full Claude install and gbrain ships from its own repo. HostConfig gains the optional install.instructionTier slot, declared by both instruction-tier hosts. README host table now matches what setup actually does. Inspired by dietrichgebert/ponytail's instruction-tier AGENTS.md fallback — one generated source, never per-host hand copies. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat(preamble): AskUserQuestion repetition cut — gated, passed NOT-WORSE A/B Removes the duplicate statements v1.71's compaction left in the AskUserQuestion Format section: the completeness rule restated in the prose triad, the auto-decide marker syntax stated twice, the Conductor-flakiness explanation stated twice, and the self-check's full triad restatement. Every verbosity floor and all 14 format pins stay (Layer 0 green). The gate this decision rested on ran before landing (new periodic skill-e2e-auq-repetition-cut-ab.test.ts, pre-cut ref 3263fff vs this render, same harness as auq-verbose-vs-carved-ab): POST 7/7 format elements, substance 5 — identical to PRE. No degradation; the load-bearing-repetition hypothesis did not hold for these duplicates. Net: -236 bytes per tier-2+ skill (~9.7KB corpus). Golden ship baselines re-captured for the deliberate change. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat(evals): with-skill vs without-skill arm benchmark — measures whether gstack's behavioral layer earns its tokens Ponytail's honest-benchmark method pointed at gstack itself: 3 build-shaped tasks (native-platform over-build trap, CRUD endpoint, bug fix with planted decoys) x 2 arms, real claude -p sessions, scored on the git diff left behind. A research instrument, not a release gate — no assertion compares arm scores. Arms use the PROVEN project-scope pattern: the with-arm installs a build-discipline skill (extracted reuse-ladder + bounded-closer content, not whole-file copies) into the fixture's .claude/skills/ with a CLAUDE.md routing line and an explicit invocation; a live spike confirmed claude -p discovers and invokes project-scope skills via the Skill tool (3 turns, exact-output probe). Fixtures are git init + local bare origin; diff capture is three lines of git, no worktree machinery. Failure taxonomy: zero-diff arms are VALID scored cells (deterministic 0/none, no API call), harvest failures record harvest:null, judge_error cells are excluded from aggregates but named in the report — nothing drops silently. armJudge: fixed sonnet judge, 0-3 unrequested-structure rubric, must name the construct or say none, bounded retry-on-malformed; callJudge gains optional temperature/max_tokens (defaults unchanged). recordE2E now populates tokens_used for every E2E. Eval schema v2: harvest gains {insertions, deletions, net}, tolerant reads keep v1 runs comparable. Registered periodic in E2E_TIERS + touchfiles (with the auq-repetition-cut A/B); periodic detach timeout raised to the new shard-census floor. Free selftest (8 tests, zero API) pins fixtures, extraction, arm asymmetry, diff capture, judge plumbing, and the retry bound. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * test: absorb the ponytail-import wave into the guard fixtures — ceilings, schema pin, triad phrasing Skeleton ceilings re-captured for the 17 carved skills the wave deliberately grew (reuse ladder + bounded closer + shortcut trail, net of the gated -236B AUQ cut), each with its measured size in the comment per the carve-guards protocol. eval-store schema pin updated to v2 (harvest gains insertions/deletions/net). The AUQ prose-triad keeps its pinned per-choice phrasing ('explicit on EACH choice') while still deferring the score scale to the canonical Format rule — the shipped cut is strictly closer to the pre-cut text than the render that already passed the NOT-WORSE gate. Autoplan carve anchors follow the Phase 2.5 renumbering. Golden ship baselines re-captured. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * test: observability partial-file pin follows eval-store schema v2 Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat(test-runner): GSTACK_FREE_JOBS + opt-in flaky-retry pass for syscall-supervised sandboxes GSTACK_FREE_JOBS overrides the computed shard count (the free runner's analogue of the paid runner's EVALS_JOBS). On Vercel sandboxes, PID 1 installs a seccomp filter whose supervisor spuriously fails access(2) for busy processes — measured: 200/200 git-init probes fail 'Cannot access work tree: Permission denied' while the suite runs at 6 shards, 0/200 idle; statx succeeds while access fails on the same path in the same process. One serial mega-shard maximizes per-process pressure and fails too; 2 shards is the measured sweet spot. GSTACK_FREE_RETRY_FLAKY=1 (default OFF — dev boxes should see flakes) re-runs attributed failures once, serially, capped at 5 files; a clean retry downgrades to a loud FLAKY-PASS naming the offenders, a repeat failure stays red, timeouts and unattributed failures never retry. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(browse): portable temp paths — TEMP_DIRS allowlist, tmpdir()-based test files Local path validation now accepts os.tmpdir() alongside the classic /tmp (new TEMP_DIRS in platform.ts): on macOS os.tmpdir() is /var/folders/..., and TMPDIR-honoring CI/sandbox environments point it elsewhere entirely — both are legitimate scratch space. Remote file serving (TEMP_ONLY) stays pinned to TEMP_DIR alone; no change to the exfil boundary. commands.test.ts drops 41 hardcoded /tmp literals for a tmpp() helper on os.tmpdir() (two message assertions now reference the same variable), and path-validation's symlink-escape test targets /etc/hosts instead of /etc/crontab — the target must EXIST for realpath to resolve the link (a dangling target falls back to the link's own path and passes vacuously), and /etc/crontab is absent on Amazon Linux. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(config): portable sha256 — Linux ships sha256sum, not shasum resolve-user-slug and endpoint hashing exited 127 on Amazon Linux (shasum is a macOS/perl tool). New _sha256_hex helper prefers sha256sum and falls back to shasum, matching gstack-verify-gate's existing pattern; both call sites converted. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(next-version): only trust ls-remote when origin is actually configured Without the guard, git DWIMs the literal 'origin' as an ssh host/path; on hosts whose transport launders exit codes the probe 'succeeds' with zero branches and the allocator silently sees an empty queue — the exact duplicate-allocation failure (garrytan#2545) fetchGitClaimed exists to prevent. git remote get-url origin gates the probe; absence falls through to the existing local-refs path with its staleness warning. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs(testing): sandbox-doctor — one command makes a cloud sandbox run the suite green Measured failure taxonomy for Vercel/Conductor sandboxes (missing /dev/fd, 64M /dev/shm, seccomp-supervisor access(2) EACCES under load, uid-1000 processes with FULL capabilities defeating chmod-denial tests, no X server, no git identity, Conductor git-shim exit-code laundering) plus the idempotent script that treats all of it and seeds the run recipe. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(config): converge on main's self-contained sha8_of — its tests extract the function standalone The merge kept a branch-local _sha256_hex helper; main's v1.72 landed the same portability fix inline WITH tests that extract sha8_of()'s text and run it under a shim-only PATH — a helper call can't satisfy that shape. Adopt the landed implementation at both hash sites. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * test: coverage for GSTACK_FREE_JOBS override and failingFiles attribution Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * test: coverage for TEMP_DIRS widening and remote-serving TEMP_ONLY asymmetry Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * test: coverage for gstack-shortcut marker grammar and retro harvest joint Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * test: coverage for sandbox-doctor shell syntax and idempotency guards Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(test-runner): empty-shard outcome carries failingFiles; harden flaky-retry list The empty-shard early return omitted the (required) failingFiles field — tsc TS2741 — feeding undefined into the flaky-retry flatMap. Also drop the dead 'else if (worst !== 0)' guard (the enclosing if already pins it). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(release): version-bump write regenerates the version-stamped agents digest agents-digest/gstack-AGENTS.md embeds VERSION in its first line and is byte-freshness-gated (test/agents-digest.test.ts + Skill Docs Freshness CI), but nothing in the release path regenerated it — every version-bumping ship of this repo would land red. write now spawns the repo's own generator when present (agentsDigest true/false/null in the output JSON), and ship's evidence gate allow-lists the digest alongside VERSION/package.json. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(setup): instruction-tier explainer prints the script-anchored digest path $(pwd) printed a nonexistent path when setup ran from any other directory; both arms now share one print_instruction_tier() using SOURCE_GSTACK_DIR. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * test(digest): broaden AGENTS.md writer tripwire; pin digest-resolver ladder lockstep The print-path-only guard now catches tee/install/rsync/dd/truncate, >> appends, and laundered variable-destination writes. New test ties the digest's hand-rendered reuse-ladder text to the preamble resolver so an edit to either fails CI instead of shipping drift. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(retro): shortcut harvest drops placeholder markers and convention docs The Step 11.5 grep matched documentation mentions (dec-<id>, dec-*) in checklists, resolver sources, and convention tests, reporting phantom debt rows on gstack itself. A trailing filter kills placeholder forms; prose tells the agent to discard convention-quoting hits. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(review): advisory findings count in per-specialist stats Without this, simplification (all-advisory by construction) would log findings:0 every run and auto-gate itself into permanent silence after 10 dispatches. The advisory carve-out governs score and header only. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(evals): arm-benchmark harvest and judge hardening - Harvest diffs against the recorded seed SHA (origin/main is movable by an agent that commits AND pushes; a recorded SHA is not). - Fixtures get a node_modules .gitignore and the git wrapper a 64MB maxBuffer, so a vendored-dependency arm is scored instead of killing the cell. - The judge diff cap is a named constant with loud truncation (log + judge_reasoning suffix). - Judge prompt block markers carry a per-call random sentinel, so a diff containing a faked closing marker cannot escape the data block. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(evals): AUQ A/B vendored pre-cut arm + judge-error inconclusive taxonomy - The PRE arm read a branch-local SHA (3263fff) that becomes unreachable on fresh clones after the squash-merge; the pre-cut render is now a vendored fixture. - A judge failure on one side no longer coerces substance to 0 (which fabricated DEGRADATION on POST-side failures and masked regressions on PRE-side failures): null substance = inconclusive, format still gates. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * test: regression pin for the originConfigured guard vs laundering git shims On healthy hosts the guarded and unguarded paths behave identically, so a revert passes the suite; only a shim that makes 'git ls-remote' exit 0 with empty output (the Conductor wrapper's observed behavior) exposes it. Pins that the empty 'successful' probe is never trusted as an empty queue. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(sandbox-doctor): missing /dev/shm no longer aborts the doctor under set -eu Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * chore(touchfiles): close dep-list gaps for the new evals - arm-benchmark entries gain ship/SKILL.md (buildBehavioralSkill extracts sections from the rendered ship skill) - review-army-simplification entries gain their planted fixtures + test file - auq-repetition-cut-ab gains llm-judge.ts and the vendored PRE fixture Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * chore: re-capture context-budget fixture — lock the WS6-3 reduction and Step 9 deltas Per the ratchet protocol: the AUQ repetition cut shrank per-skill eager tokens but the fixture was never re-captured, leaving the win unlocked. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(release): digest regen is an explicit --regen-digest opt-in, not presence-sniffed code exec Review (security) caught the cycle-1 fix executing any repo's scripts/gen-agents-digest.ts on plain 'write' — arbitrary code exec from a hostile clone on a routine bump, contradicting the binary's own containment posture. The regen still runs the TARGET repo's generator (a 'trusted' copy beside the binary would false-red the freshness gate on version drift), but only under the flag: /ship passes it deliberately, in a repo whose code the operator already executes (its test suite). Plain write is side-effect-free again. Also: uniform output shape (agentsDigest: null on the JSON-manifest branch), a REAL generator round-trip test replacing the misnamed lockstep check, and land-and-deploy's evidence gate gets the same digest allow-path as ship so the two grading surfaces agree. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(test-runner): flaky-retry vetoes on ANY unattributable failure evidence The gate equated 'some failure attributed' with 'all failures attributed': a shard with one attributed failure plus a headerless failure, an unhandled error between tests, or a truncated run (no terminal summary) qualified for retry — re-running only failingFiles and masking the rest as FLAKY-PASS, re-opening the silent-truncation hole the strict classifier closes. FreeShardOutcome now carries unattributedFailures; nonzero vetoes the retry. Pins: mixed shard, truncated-with-attributed shard, empty-shard field values. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(next-version): a configured origin advertising zero heads is never trusted The originConfigured guard covered only the no-origin laundering case. With origin configured (the normal Conductor worktree state), the laundering shim makes a failed ls-remote exit 0 with empty stdout — read as 'the queue is empty', the exact duplicate-allocation bug (garrytan#2545) one layer up. A reachable remote always advertises at least its default branch, so an exit-0 zero-head probe now falls back to local refs/remotes/origin with a laundering-specific warning. Regression test shims git for both configurations. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(sandbox-doctor): loud on git-shim patch drift; document the retry-contract override - The /conductor/bin/git patch was a silent no-op if the shim's bytes drift from the exact pattern — now warns that laundering is NOT fixed. - The bashrc block documents why GSTACK_FREE_RETRY_FLAKY=1 deliberately overrides the runner's default-OFF contract on this sandbox, and how to undo it. - Test pins the guarded shm form (missing /dev/shm must not abort set -eu). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * test(digest): pin the script-anchored explainer path; catch declaration-prefixed writers - Asserts $SOURCE_GSTACK_DIR/agents-digest path and forbids $(pwd)/agents-digest (the cycle-1 fix was revertible without failing anything). - The laundered-assignment arm now matches local/export/declare/readonly/typeset prefixed assignments — the likeliest in-function writer shape in setup. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * refactor(evals): arm-benchmark selftest runs FREE on every PR The selftest lived inside the paid skill-e2e-* file, so fixture-integrity and plumbing pins executed weekly at best — a broken fixture would ship past every gating check and be discovered when the periodic run burned money on a dead instrument. Harness extracted to test/helpers/arm-benchmark-harness.ts, selftest to test/arm-benchmark-selftest.test.ts (free suite). Touchfiles: harness added to the three benchmark dep lists; the auq-repetition-cut-ab tier comment now states the MANUAL re-run obligation honestly (periodic runs force EVALS_ALL, so dep lists cannot auto-trigger it). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * chore: re-capture context-budget fixture after cycle-2 template deltas Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(sandbox-doctor): keep both heredoc bodies under the 512B pipe-deadlock window The cycle-2 additions pushed the python-patch and bashrc heredocs into the 512-65536B window test/heredoc-pipe-deadlock.test.ts guards (sh scripts get no BASH_COMPAT escape hatch). Same content, tighter prose; the drift warning now reuses the patch pattern variable instead of a second literal. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(review): a gstack-shortcut marker only suppresses findings when its decision id resolves in the ledger Cross-model catch (Claude adversarial + Codex agreed): any diff author could fabricate a marker and silence Completeness review of that gap. Reviewers now resolve the dec-id via gstack-decision-search; an orphan marker is reported as a forged suppression, not honored as debt. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(autoplan): define the B2 gate path — accepted premise challenges amend the plan and re-run Eng The final gate offered B2 (respond to User Challenges) but the option handler table omitted it, leaving accepted challenges with no amendment or Eng re-review path. B2 now walks challenges one at a time; an accepted one amends the plan and re-runs Eng (the gate always reviews the final plan), sharing D's 3-cycle cap. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat(evals): arm benchmark runs each fixture's functional oracle — correctness before LOC The plan's metric order is diff-quality FIRST, but cells never ran the fixtures' own run-tests.js, so a refusal, a broken implementation, and working code were indistinguishable in aggregates (Codex adversarial catch). Tasks with an oracle declare checkCmd; every cell records checks=pass|fail|none in the report line and eval store. Selftest pins the oracle declarations and that the planted bug fails its own check pre-fix. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(ship): check the bump's agentsDigest result; state the --regen-digest trust envelope honestly A failed digest regen warned and moved on — ship now instructs re-running the generator and staging the digest with the bump (the freshness check stays red otherwise). The 'no-op everywhere else' phrasing oversold safety: the step now names what executes and why that is inside the envelope Step 5 already opened (the repo's own test suite). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(test-runner): GSTACK_FREE_JOBS accepts digits only — parseInt truncation defeated the loud-failure contract '2abc' silently became 2 and '3.7' became 3 despite the error text claiming a positive-integer requirement. Strict /^\d+$/ pre-check; both shapes pinned. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(sandbox-doctor): atomic git-shim patch, :99-socket Xvfb check, dnf gate, non-interactive sudo - The /conductor/bin/git patch writes tmp-then-rename with a .orig backup — a concurrently spawned git can never exec a truncated shim. - Xvfb running-check looks for the :99 socket, not any-display pgrep. - Xvfb install is dnf-gated so non-dnf distros degrade to a warning instead of aborting the remaining fixes under set -eu. - The bashrc /dev/fd restore uses sudo -n || true — no password prompt at every shell start on non-passwordless machines. - BASH_COMPAT=50 keeps heredoc bodies off the bash pipe window. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(build): a failed agents-digest regen fails gen-skill-docs instead of deferring the red to CI Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(browse): an untrustable TMPDIR (/, $HOME, a cwd ancestor) never widens the local allowlist TEMP_DIRS honors os.tmpdir() at daemon start; a daemon launched with TMPDIR=/ would have trusted the whole filesystem for local path validation for its lifetime. Subprocess pins cover /, $HOME, cwd-ancestor rejection and that a benign distinct TMPDIR (the sandbox recipe's $HOME/tmp) stays honored. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * chore: zero-heads warning names the benign cause too; digest path declaration made load-bearing; ratchet re-capture - The ls-remote zero-heads warning no longer accuses an empty remote of running a laundering shim. - instructionTier.rulesFile now must equal the generator's DIGEST_RELPATH (and setup must print it) — the declaration fails with the real path instead of lying silently. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * chore: file ship-time follow-ups in TODOS skillify HOME-override gate red (pre-existing, proven on main), the auq-verbose-vs-carved-ab branch-local ref, eval-store harvest union, evidence digest allow-path scoping, and the WS6-2 dead-frontmatter live-host verification deferral. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * v1.73.0.0 chore: version bump + CHANGELOG — ponytail import wave Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * chore: raise ship skeleton parity ceiling — measured 75,592 after the v1.73 release-step prose The --regen-digest trust-envelope paragraph (Step 12) and the evidence-gate digest note (Step 16) grew the ship skeleton past the previous 75,420 ceiling. Re-measured per the deliberate-change protocol. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs: update project documentation for v1.73.0.0 - README.md, docs/skills.md, AGENTS.md: /autoplan phase order corrected to CEO → design → DX → eng (eng always last); /review rows note the advisory simplification lens - docs/PROJECT_STRUCTURE.md: add agents-digest/, gen-agents-digest.ts, sandbox-doctor.sh, test-free-shards.ts to the annotated tree - CONTRIBUTING.md: document GSTACK_FREE_JOBS, GSTACK_FREE_RETRY_FLAKY, and the sandbox-doctor one-command fixer in the Tier 1 test section Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs: apply cross-model doc-review fixes for v1.73.0.0 - README.md: host table gains the OpenClaw explainer arm row (setup has the arm; the table claimed to match setup) - docs/skills.md: /review completeness-gaps section documents the gstack-shortcut(dec-<id>) acknowledged-debt suppression and orphan-marker flagging; /autoplan deep-dive states the recommended-option default with the 6 principles as tie-breakers - CONTRIBUTING.md: host count 8 -> 10 (Hermes, GBrain), supported-hosts list completed - docs/TESTING_INTERNALS.md: sandbox recipe says to source ~/.bashrc after the doctor seeds it; GSTACK_FREE_JOBS wording fixed from "caps" to "overrides in either direction" (matches the un-clamped runner) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(test): temp-dirs asymmetry pins are topology-aware; TMPDIR probes are POSIX-only CI exposed two wrong assumptions in the new temp-dirs tests, neither a product bug: - The remote-serving asymmetry test assumed a distinct os.tmpdir() lies OUTSIDE TEMP_DIR, but the free-shard runner nests each child's TMPDIR inside /tmp on CI — a file there is under TEMP_DIR, so serving it remotely is legitimate. The test now pins the actual exfil boundary on every topology (a cwd project file is locally readable, never remotely servable) and branches the os.tmpdir() case on nested-vs-outside. Reproduced locally with TMPDIR=/tmp/nested-tmp before fixing. - The untrustable-TMPDIR subprocess probes set TMPDIR, which Windows os.tmpdir() ignores (reads TEMP/TMP) — and on Windows TEMP_DIR is DEFINED as os.tmpdir(), so the fixed+movable two-dir topology the guard filters does not exist there. Probes now skip on Windows with that rationale; the benign-TMPDIR assertion compares realpaths. Verified under all three POSIX topologies: TMPDIR=$HOME/tmp (outside), TMPDIR=/tmp/nested-tmp (CI shard shape), TMPDIR unset (identical). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(build): DIGEST_RELPATH is a forward-slash literal on every platform path.join built it with backslashes on Windows, so the wiring test's string comparisons against setup and hosts/*.ts (which carry the forward-slash literal) could never match there — windows-free-tests red. path.join(root, DIGEST_RELPATH) at the write site normalizes fine. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(sandbox-doctor): bashrc block re-heals the /dev/shm remount on sandbox restart The 4G remount does not survive restarts; a reverted 64M shm made the multi-tab browse handoff test fail consistently under suite concurrency (observed live: two consecutive full-run failures, green in isolation, green again after remounting). Same guarded arithmetic as the doctor body. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(test): close the cross-shard porcelain race that failed Windows CI Two-part fix for the gen-skill-docs-out-dir isolation-pin failure: - cookie-import-browser built its scratch cookie DBs inside the TRACKED browse/test/fixtures/ dir (created in beforeAll, deleted in afterAll), so they flash as untracked files mid-run — a concurrent shard's porcelain snapshot caught the window on Windows. The DBs now live in a per-run tmpdir; zero source-tree writes. - gen-skill-docs-out-dir is the free suite's only LIVE porcelain-snapshot test, so it joins TREE_MUTATING (the serial quiet window): any concurrent transient tree-write can race it, and its own spawned render rewrites llms.txt/agents-digest in place (idempotent on a fresh tree). The race is pre-existing; this branch's +5 test files reshuffled shard composition and exposed it. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * v1.75.0.0 chore: queue-advance rebump — perth-v2 landed v1.74.0.0 on main The v1.73.0.0 slot this branch claimed was superseded when garrytan#2721 merged; same MINOR level relative to main per the versioning invariant. CHANGELOG entry renumbered (1.73.0.0 was branch-internal and never landed on main), digest restamped via --regen-digest. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(test-runner): duration-packed walls keep the per-file floor — predictions don't transfer across machines The committed duration seed is recorded on fast CI; a syscall-supervised sandbox replays the same files 2-4x slower. Observed post-merge: a 253-file shard predicted ~242s was wall-killed at its predicted-x3 725s wall while genuinely progressing (the old count heuristic guaranteed 1265s). Packed walls may be looser than the count floor, never tighter. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
…+ multiSelect join in question-log-hook Local patch carried on the phase0-pr-gating branch (previously working-tree dirt that /gstack-upgrade reverted twice). Verified via ~/.gstack/projects/dohma/question-log-hook-shapes.sh. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01LB2a2HTKsqKBd3vK3oRXqh
…ase 0 PR-gating instrumentation, commit 1/6) New bin/gstack-gate-log writes <branch>-gates.jsonl, a SIBLING of <branch>-reviews.jsonl: review-read cats the whole reviews file into model context per dashboard render, and specialist-stats globs *-reviews.jsonl only, so per-invocation gate rows stay invisible to every existing reader by construction. Validate-and-stamp block deliberately duplicated from gstack-review-log (load-bearing; must be safe to evolve independently). Writer-enforced contract: record_type:"gate" + gate + run_id required; effort_source:"user-override" REQUIRES a non-empty effort_reason (the "no implicit xhigh" rule made structurally uncheatable). Everything else passes through schema-free so shadow-classifier fields ride along. artifacts-init: sync allowlist + privacy map gain projects/*/*-gates.jsonl. Tests pin the writer contract AND the reader-isolation claim (gate rows invisible to specialist-stats and review-read), not just round-trips. Baseline note: 2 diff-scope migration tests + 1 gbrain-sync symlink test fail on this machine at pristine origin/main (verified in a clean worktree) — environment-specific, unrelated, not introduced here. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01LB2a2HTKsqKBd3vK3oRXqh
…itignore silently drops fixtures test/diff-scope.test.ts ran fixture git calls under the user's real git config. A ~/.gitignore_global carrying *.sql makes `git add .` silently skip every migration fixture, turning the prisma and root-migrations cases red on such machines while the script under test is correct (measured on a real machine; CI never sees it). All spawn sites now run with GIT_CONFIG_GLOBAL/GIT_CONFIG_SYSTEM=/dev/null. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01LB2a2HTKsqKBd3vK3oRXqh
…nifest (Phase 0, commit 2/6) bash wrapper + adjacent .ts under bun (embedded `bun -e` rejected: the glob→regex code is backslash-dense and bash double-quote escaping of it is a standing bug class). Emits manifests/<wtree12>.json content-addressed by gstack-wtree, plus shell-safe RUN_ID/MANIFEST_PATH/DIFF_LINES/DOC_FP/ SHADOW_TIER/SCOPE_* stdout for source <(...). The manifest is an INDEX for reviewer subagents — never a replacement for the raw diff. Shadow verdicts (A/B/C/D tier fail-upward, auth glob-vs-policy disagreement, doc-impact, load-bearing docs) are computed from an optional repo-root .gstack-policy.json and are LOGGED ONLY: nothing routes on them in Phase 0. diff-scope's exit-2 SCOPE_ERROR is carried in-band (scope.error), never swallowed. run_id: minted on first call, passed back verbatim on later calls in the same skill run so one run keeps one id across fix cycles. 12 deterministic tests (throwaway-repo pattern, hermetic git config). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01LB2a2HTKsqKBd3vK3oRXqh
…am on the >200 path (Phase 0, commit 3/6)
Scheduling: when DIFF_LINES > 200 the Red Team trigger is knowable BEFORE the
specialists run (measured: launched strictly after the last specialist in
12/14 runs, +p50 7.5 min critical path), so it now joins the SAME parallel
dispatch message. The specialist-CRITICAL late path is preserved exactly,
merged findings included. The activation sentence is verbatim-unchanged —
this moves WHEN Red Team runs, never WHETHER. review/specialists/red-team.md's
scope line carried a NARROWER trigger ("security specialist") than the
resolver ("any specialist"); aligned to the WIDER form.
Telemetry: Step 9.1 generates the shared diff/risk manifest (an INDEX in
every specialist prompt — raw diff command retained verbatim) and the merge
step appends one gstack-gate-log record per dispatched gate (trigger,
start/end, verdict, severity counts, fix_cycle, rerun_cause, manifest_wtree).
Best-effort and additive; the reviews.jsonl specialists object is untouched.
test/red-team-scheduling.test.ts pins all of it across every host that
renders Review Army (claude section, /review, factory golden; codex strips
the section by design) in both directions, including that the narrower
trigger wording can never return. Factory golden re-blessed in this commit.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LB2a2HTKsqKBd3vK3oRXqh
…cision, xhigh recorded-override-only (Phase 0, commit 4/6) Demotions (audit: doc voice output is 95% wording-level over 166 runs; plan voices on routine plans do not need high): codex plan outside voice (review.ts plan-review site) high→medium, codex doc voice (document-release) high→medium, plan-design-review voice high→medium (design.ts ternary now keys on isDesignReview — the live audit alone keeps high). Kept at high, deliberately: the codex adversarial challenge and the codex structured review (P1 gate) — adversarial coverage is untouched. Recording: adversarial-review / codex-plan-review / codex-doc-review / codex-review reviews.jsonl rows all gain effort + effort_source; the adversarial step appends per-gate gate-log records (adversarial-claude, codex-adversarial, codex-structured) with tokens from the codex stderr "tokens used" line. /codex --xhigh stays the ONLY path to xhigh and must be recorded as effort_source:"user-override" with a non-empty effort_reason — gstack-gate-log refuses the record without one. test/effort-routing.test.ts pins all of it line-scoped (tempfile markers disambiguate call sites, so demoting the WRONG site fails); the gen-skill-docs plan-design-review pin flips high→medium — the one legitimate existing-pin change. Factory golden re-blessed in this commit. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01LB2a2HTKsqKBd3vK3oRXqh
…elemetry + fix-cycle index (Phase 0, commit 5/6) Step 18 now computes the doc fingerprint (list-based, from the shared manifest — stable across fix cycles that edit already-listed files) and skips a REDISPATCH only when all three hold: same run_id, same doc_fingerprint, completed non-error record. Cross-run fingerprint matches STILL dispatch and log shadow.redispatch_would_skip:true — the Phase 1 evidence for widening. Backgrounding was evaluated and NOT adopted: the only race-free overlap window (post-push, pre-splice) is ~1-2 min against doc-release's p50 8.8, while this guard alone removes the 12/36 re-dispatches. Every dispatch (including failures) writes a gate:"doc-release" record with files_updated / doc_commit / documentation_section (stored verbatim so a same-run skip can reuse it) and the doc-impact shadow block, making doc-release's success observable downstream for the first time. The non-blocking failure contract and the S17→S18→S19 position are verbatim untouched and now pinned. Fix-First loop gains a 0-based fix_cycle index + rerun_cause on re-dispatched gates — telemetry only, loop contract unchanged. All three ship goldens re-blessed in this commit. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01LB2a2HTKsqKBd3vK3oRXqh
… computed alongside the deciding wtree rule (Phase 0, commit 6/6) The content-first wtree rule still decides exactly as before; Step 3.5a now ALSO computes the 4-commits/>5-files heuristic even when wtree already graded CURRENT, and logs the pair as a gate:"staleness-check" record with a disagreement flag (critical_path:false). The audit found the heuristic biased toward STALE by construction in a WIP-commit workflow — this record measures that claim instead of arguing it. Logged, never routed: no readiness decision, threshold, or inline-review offer changes. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01LB2a2HTKsqKBd3vK3oRXqh
…p, was 1.231x) The parity ratchet caught the instrumentation prose pushing the ship union past its v1.64.1.0 size baseline. Compressed rationale sentences and JSON placeholder descriptions across review-army/adversarial/pr-body — every pinned contract string survives (the whitespace-tolerant pin variant covers one re-wrapped phrase). All three goldens re-blessed. Ship now sits at exactly 1.2200x — the NEXT addition to ship prose must bring its own compression or a deliberate baseline rebase. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01LB2a2HTKsqKBd3vK3oRXqh
…te-log fields, compress Phase 0 prose Upstream v1.68.0-.3 grew ship to 1.1805x, leaving 7415 bytes of headroom against the 1.22 cap; Phase 0 instrumentation adds 8365. Two changes close the 950-byte gap without losing an instruction: - gstack-gate-log now DERIVES commit (from the authoritative commit_full it already stamps) and DEFAULTS diff_scope/critical_path, only when the caller omits the key. All five payloads stop hand-substituting three fields that were pure redundancy. manifest_wtree is deliberately NOT derived: it records the manifest's worktree while log-time is already stamped as wtree, so deriving it would make the two always agree and destroy the only signal that a manifest came from a different tree than the gate it logs. - 30 wording-only compressions across the Phase 0 blocks. Ship is now 1.2194x (115 bytes of margin). Two pins in ship-doc-release-fingerprint.test.ts broke on reflow, not on meaning, and are re-pointed whitespace-tolerantly; both mutation-verified. Five new gate-log tests cover derive-when-absent, explicit-value-wins, and the manifest_wtree asymmetry; mutation-verified by removing the undefined guards. Codex/Factory ship goldens re-blessed (carved claude SKILL.md skeleton is unchanged, so its golden was already correct).
v1.75 renders review/codex sections behind STOP pointers instead of inlining them, so the rendered-prose pins retarget review/sections/ review-army.md, review/sections/adversarial.md and codex/sections/ review-mode.md; the doc-release skip counter stops counting the new negated 'Never skip the dispatch' invariant; the codex xhigh paragraph loses 40 bytes to refit the 57800-byte skeleton parity cap. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01FzUsqrfLRmeuVv5nmqq3oh
…D, auth/migrations → D), outcome metadata, reviewer budgets with distinct slots and fail-closed completion, delta rerun-check on the recorded SHA, review packet, outcome report, context guard Phase 0.5 follow-up (P0 cost control, 2026-09-03). New: bin/gstack-outcome, gstack-review-budget, gstack-review-packet, gstack-outcome-report, gstack-context-guard. gstack-diff-manifest gains a routing block (shadow untouched); gstack-gate-log stamps outcome_id/slice/risk_tier. Hardened after one Sol-high adversarial review: --no-renames canonical paths, scope.migrations → D, duplicate-slot refusal, one escalation per run, verify-of bound to a recorded finding, per-cycle budgets with a tier floor, rerun-check bound to the plan-time HEAD (fail closed on git failure), BLOCKING_CATEGORIES vocabulary, manifest reuse refreshes run_id. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01UWhbwHBzWYCEbWiJSVVXd1
…h the governor; every subagent pinned to sonnet; every codex call pins effort Specialist selection, adaptive gating, LOC-triggered Red Team, the Claude adversarial subagent and the free-form codex challenge are replaced by the plan from gstack-review-budget: A/B one codex-structured@medium, C two, D three (codex-structured@high + security + data-migration|red-team). Every dispatch is gated by `dispatch --cycle`, records a `verdict`, and `complete` fails closed. Advisory (informational) findings are never fixed inside /ship (max 5 advisories in the PR body); blocking = CRITICAL/P1/P2 or a BLOCKING_CATEGORIES category; repair cycles bounded by REPAIR_CYCLES_MAX; after a fix only rerun-check-scoped verification unless FULL_RERUN. Coverage audit, plan completion and doc-release are slice-gated by outcome metadata (missing → final). The UNDER_CODEX preflight stays on the routed codex review. Autoplan phase voices and greptile pinned (medium / sonnet). Old-routing test pins updated; goldens regenerated; size corpus ratio 0.788. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01UWhbwHBzWYCEbWiJSVVXd1
…override; no template may spell ultra; every rendered codex call pins its effort dohma ruling 2026-09-03: Codex 5.6 Sol MEDIUM is the project default (repo-level .codex/config.toml), HIGH is explicit and reserved for C/D-risk implementation or the final adversarial review, ultra is never selected automatically. gstack-gate-log refuses an effort above high whose effort_source is not user-override (which already requires effort_reason); the codex skill documents the rule; the effort-routing suite walks every rendered SKILL.md and section and refuses ultra anywhere and any codex exec/review line without model_reasoning_effort. Codex skill cap lifted 100B for the clause (measured 57,809). Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01UWhbwHBzWYCEbWiJSVVXd1
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Review governor (dohma P0 cost control, 2026-09-03)
Makes the Phase 0 shadow classifier the router for
/shipand/review, with hard reviewer budgets, outcome/series slice mode, delta-only reruns, one shared review packet, per-outcome telemetry, a context-size guard hook, explicitmodel: "sonnet"on every subagent dispatch, and an explicit effort on every codex call. Companion: dohma PR v639dragoon/dohma#728 (policyroutingsection,.codex/config.tomlproject default, fixtures for docs-only / garrytan#726 / feature / garrytan#727).a214804deterministic bins:gstack-outcome,gstack-review-budget(plan / dispatch / verdict / complete / rerun-check / finding / resolve / report),gstack-review-packet,gstack-outcome-report,gstack-context-guard;gstack-diff-manifestrouting block (fail-upward to D; auth surfaces and migrations → D);gstack-gate-logstamps outcome fields. Hardened after one Sol-high adversarial review (renames, distinct slots, fail-closed completion, verify-of bound to a finding, per-cycle budgets, recorded pre-fix SHA, blocking-category vocabulary).3a212afprose: specialist selection, LOC-triggered Red Team, the Claude adversarial subagent and the free-form codex challenge replaced by the plan (A/B 1, C 2, D 3 reviewers); informational findings never fixed inside/ship; repair cycles bounded; coverage/plan-completion/doc-release slice-gated; UNDER_CODEX preflight kept on the routed codex review; autoplan voices and greptile pinned.gstack-gate-logunless they are a reasoneduser-override; no template may spell ultra.Tests: focused governor + prose suites 1,151 pass;
bun run test:freegreen except two pre-existing failures (gbrain-sync symlink pin, benchmark-runner shard timeout) verified on a clean checkout. Rebase onto upstream v1.79 is a separate follow-up.🤖 Generated with Claude Code
https://claude.ai/code/session_01UWhbwHBzWYCEbWiJSVVXd1