Skip to content

docs: phased build-out plan for the research loop (substrate, depth, parallel) - #127

Merged
renmengye merged 2 commits into
mainfrom
docs/research-loop-buildout
Aug 23, 2026
Merged

renmengye merged 2 commits into
mainfrom
docs/research-loop-buildout

Conversation

@renmengye

Copy link
Copy Markdown
Member

What

A plan design doc (docs/design/research-loop-buildout.md) that turns the
research-loop.md north-star into a concrete, phased build-out — the three
efforts we converged on, sequenced with their dependencies.

The shape

One substrate, two axes, one governor:

  • Substrate — consolidate the role-runner (climb/steward/follow-up onto the
    shared run_role, as the judges already are).
  • Depth (k axis) — the agent iterating on its own results, k a contract
    dial, generalizing the panel-on-wake slices (k = 1 today).
  • Parallel dispatch (M-within-a-climb) — a portfolio on one objective,
    deadline-salvaged.
  • Governor — the benchmark-integrity gate (suite / no-regression), threaded
    through so scaling the engine doesn't scale benchmark-gaming.

Sequencing (and why)

  1. Phase 1 — role-runner consolidation, thin/feature-driven (no dispatcher
    dep). Enables the rest; avoids each axis growing its own per-lane
    orchestration (the Tick: drive the codex author across the climb-authoring lanes #123 lesson).
  2. Light up the dispatcher on a real long-eval benchmark — the entry gate to
    cross-session depth and parallel (it's proven but dark today).
  3. Phase 2 — depth (2a inline, then 2b cross-session), with the metric
    taxonomy + suite gate landing alongside.
  4. Phase 3 — parallel dispatch, on the consolidated runner + live dispatcher.

Depth before parallel: depth extends something already built and is cheaper to
make honest; parallel is the larger new capability.

Status

Status: plan — reviewable before any code. Non-goals and open questions are
called out in-doc (esp. for parallel: workspace sharing, member interaction,
budget split, salvage ranking, depth×parallel composition). Cross-linked from
research-loop.md.

No secrets / no large files

Docs only.

🤖 Generated with Claude Code

…parallel)

A reviewable plan that sequences the three efforts toward research-loop.md's
north-star: (1) consolidate the role-runner (thin, feature-driven), (2) the depth
axis (k as a contract dial, generalizing panel-on-wake), (3) parallel dispatch
(portfolio + deadline-salvage), with the benchmark-integrity gate threaded
through as the governor and the dispatcher-must-go-live dependency called out.
Cross-linked from research-loop.md.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

@github-actions github-actions Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Round 1 — reviewed head 55937afd — reviewer hermes/gpt-5.6-terra.

second opinion — terra
Advisory findings from autoresearch — the code owner decides. Reply to disagree; the autoresearch:no-review label opts this PR out.

Verdict: nothing blocking — 1 advisory note.

1 finding attached to the lines below.

Comment thread docs/design/research-loop-buildout.md Outdated

@github-actions github-actions Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Round 1 — reviewed head 55937afd — reviewer claude/claude-opus-5.

Advisory findings from autoresearch — the code owner decides. Reply to disagree; the autoresearch:no-review label opts this PR out.

Verdict: nothing blocking — 3 advisory notes.

Advisory (non-blocking):

  • Phase 1 premise is stale: the authoring roles already run through run_role (docs/design/research-loop-buildout.md:30; high)
  • Phase 1 acceptance criteria are already met by current code (docs/design/research-loop-buildout.md:46; high)
  • parallel-climb reference does not resolve to anything in the repo (docs/design/research-loop-buildout.md:95; low)

Verified doc cross-references that do resolve: consolidation.md, agent-substrate.md ("The role-runner: one loop replaces five drivers", line 266), dispatcher.md, scaling.md, orchestrator-verify.md, judge-placement.md, plus the code hooks panel_reads/panel_revisions (orchestrator.py) and supports_resume (harness.py, climb.py:777). Claims I could not check from the repo: the dispatcher being "proven but dark" on real Slurm, PR numbers #123/#125, and the climb-lessons notes (the notebook is a separate private repo per architecture.md).

… run_role

Review caught that the five-drivers→run_role consolidation is already DONE
(consolidation.md; climb/steward/followup all call run_role). Reframe Phase 1 as
the layer ABOVE run_role — the shared run→await→decide composition seam that depth
and parallel both need — rather than a from-scratch role-runner consolidation.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
@renmengye
renmengye merged commit c535648 into main Aug 23, 2026
1 check passed
@renmengye
renmengye deleted the docs/research-loop-buildout branch August 23, 2026 13:15
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant