Skip to content

Design sketch: planned search, experiment ladder, multi-agent scaling - #25

Merged
renmengye merged 2 commits into
mainfrom
docs/scaling-sketch
Aug 6, 2026
Merged

renmengye merged 2 commits into
mainfrom
docs/scaling-sketch

Conversation

@renmengye

Copy link
Copy Markdown
Member

v1 of docs/design/scaling.md — iterating with Mengye, explicitly not build-gating phase 4. Three parts: (1) a planning agent proposing hypothesis-motivated single-knob searches around a per-benchmark leader config, plans published as vetoable GitHub issues; (2) resource tiers with pre-registered small-scale proof-of-idea gating large runs, plus periodic consolidation runs with leave-one-out ablations; (3) multi-agent via repo-level assignment — persistent agent ids, one tick chain each, per-agent credentials (API key or subscription-seat setup-token), frequency caps. Ends with four open questions for the next iteration turn.

🤖 Generated with Claude Code

Captures Mengye's direction (2026-08-06) plus a first iteration: leader
configs, twice-motivated plans as vetoable GitHub issues, pre-registered
proof-of-idea before large runs, consolidation runs with leave-one-out
ablation, per-agent identity/chains/credentials, repo-level assignment
with frequency caps, seats via setup-token.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@github-actions

github-actions Bot commented Aug 6, 2026 •

Copy link
Copy Markdown

Advisory review — not an approval. Automated findings from autoresearch; a human code owner still owns this PR. Reply to any finding you disagree with, or add the opt-out label to silence future runs on this PR.

  • Veto window on bot-authored plan issues is fail-open with no fallback if no human is watching (low confidence, docs/design/scaling.md:33)
    The plan-issue flow is described as: bot opens an issue, waits a provisional one-day veto window, then runs under the self-initiated budget. Nothing in the sketch describes what happens if no human sees the issue (vacation, notification failure, issue opened on a repo nobody watches) — silence is consent, so a mis-scoped or expensive search program executes by default. If that is intended, it would be worth stating the compensating control explicitly (e.g. the frequency cap and small-tier gate are the only backstops); otherwise a fail-closed variant for the first plan on a newly onboarded repo may be safer. Design-level observation only, not a code defect.
  • Plan issues are a bidirectional channel with untrusted input; no trust boundary is stated (low confidence, docs/design/scaling.md:30)
    Plans live as GitHub issues on target repos and a human is expected to "re-scope the plan with a comment". That implies the planner (or a later session) reads issue comments, which on a public or org-wide-visible repo can be authored by anyone, not just the intended reviewer. The sketch does not say who may author an effective re-scope (author association / write permission check) nor how comment text is treated relative to the bot's own instructions. Worth pinning down in the next iteration alongside the requested-lane gate discussion, since the whole point of that gate is to prevent the pipeline from feeding itself privileged tasks.
  • Partition validation is per-target, but the leader ledger and consolidation runs are per-benchmark (low confidence, docs/design/scaling.md:76)
    Isolation is claimed "by construction" because the agent→targets mapping is validated as a partition of repos. However the leader ledger is described as per benchmark, and consolidation runs update the leader; if two targets served by different agents contribute to the same benchmark, the repo-level partition does not by itself prevent two agents from driving the same leader/consolidation line. The doc does mention lease/CAS on the ledger, but ownership of consolidation runs (which agent launches one, under whose budget and frequency cap) is unspecified. Flagging as a gap to resolve rather than a bug in code.

Docs-only PR (a design sketch plus a pointer paragraph in architecture.md); there is no executable code, config, or test to verify, so all findings are design-gap observations rather than defects. The relative link [design/scaling.md](scaling.md) in architecture.md resolves correctly since both files sit in docs/design/. Claims about claude setup-token behaviour, seat/ToS permissibility, and the existing "env-injection seam" cannot be verified from the provided context; the doc already marks the first two as needing a spike/ToS check.

…ceiling, ToS flag

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant