Skip to content

feat: implement bounded parallel story fan-out with isolated worktrees #229

Description

@rasanderoland

Context

scm.max_parallel already exists, but upstream main clamps every value greater than 1 to 1 because Phase 5 parallel fan-out is not implemented. The engine remains sequential: it dispatches one story, runs its development/review/verification flow, integrates it, and only then selects the next story. Worktree isolation, per-unit branches, and merge-back already exist. This is an upstream bmad-loop Phase 5 design request, not an OpenEngine-specific fork.

Design request

Implement bounded parallel story fan-out with these invariants:

  • A configured scm.max_parallel greater than 1 must hard-fail unless scm.isolation is worktree and scm.branch_per is story. Keep scm.max_parallel = 1 behavior unchanged.
  • Admit work using declared file globs plus named exclusive resources, such as database, lockfile, service:name, generated artifacts, or shared verification infrastructure. File/resource conflicts serialize. Unknown or invalid footprints are exclusive, and the journal must record why. Shared verification infrastructure should be able to reserve migrations, test databases, or services so parallel source work cannot race them.
  • Allow up to N workers in flight, but keep exactly one coordinator/integration lane. Workers must never write the target branch concurrently.
  • Persist deterministic dispatch and integration order in run state. Completion timing must not determine run history.
  • Each worker must record the target base SHA from which it was started.
  • Integrate each worker against the latest target through a candidate branch/worktree. Run deterministic verification commands against the combined candidate, then advance the target only if its SHA is still the candidate's recorded target SHA; never force-advance a target that moved.
  • Define and test stale-base behavior for all merge strategies: ff may advance only when the candidate is a valid fast-forward from the current target; merge must compose against the latest target and preserve/escalate conflicts; squash must create its deterministic squash candidate against the latest target. A stale target must trigger deterministic recomposition/retry or an explicit escalation, with the candidate and evidence preserved.
  • Define one worker pause/escalation behavior: the coordinator stops dispatching, lets already-running siblings reach a safe terminal boundary, persists their outcomes, and pauses without silently dropping results. If maintainers prefer cancel-on-first-failure, make that an explicit design choice rather than an assumption.
  • Persist durable state for every in-flight worker: task, worktree, branch, phase, session, and base SHA. Support crash-safe resume, stop, orphan cleanup, and useful TUI/status/attach behavior for these workers and candidate integrations.
  • Preserve the existing semantics of max_stories, epic boundaries, selectors, plan/spec/done checkpoints, notifications, keep_failed, delete_branch, seed_adapter_defaults, and worktree_seed.
  • rollback_on_failure remains an in-place-only knob and is irrelevant in parallel worktree mode.
  • Defer sweep-bundle parallelism unless maintainers explicitly opt it into this issue.

Acceptance tests

The implementation should cover:

  • N-in-flight cap.
  • Overlap of independent stories, proven with deterministic barriers.
  • Serialization of file/resource conflicts.
  • Unknown footprint treated as exclusive, with the journal reason.
  • Isolation and branch_per guards.
  • Deterministic merge order when completion order is inverted.
  • ff, merge, and squash stale-base paths.
  • Post-composition verification failure.
  • Target SHA movement during integration.
  • Merge-conflict preservation and escalation.
  • A worker pause while siblings continue to a safe terminal boundary.
  • Crash/resume, stop, and orphan cleanup.
  • Regression equivalence for max_parallel = 1.

A manual two-story timing run is supplementary only; deterministic barrier-based tests are the primary concurrency proof.

Implementation Notes

Relevant current-main seams include policy parsing and validation; the engine dispatch loop; per-unit worktree creation, integration, and cleanup; persisted run state and journal; adapter sessions; verification; orphan/reconciliation cleanup; sweep; and diagnostics/TUI/status/attach surfaces. Please use the current symbols and contracts in those areas without depending on exact line numbers, which are not stable.

No activity

Activity on this issue will appear here.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    P4Parked - needs design, blocked upstream/PR, or speculativearea:engineOrchestrator engine and run lifecyclearea:parallelParallel story executionenhancementNew feature or requestneeds-designAwaiting a maintainer design decision before code

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions