Skip to content

fix(server): isolate provider command lanes by thread - #7071

Open
clintebbesen wants to merge 2 commits into
pingdotgg:mainfrom
clintebbesen:fix/provider-command-cross-thread-lockout
Open

fix(server): isolate provider command lanes by thread#7071
clintebbesen wants to merge 2 commits into
pingdotgg:mainfrom
clintebbesen:fix/provider-command-cross-thread-lockout

Conversation

@clintebbesen

@clintebbesen clintebbesen commented Aug 15, 2026

Copy link
Copy Markdown

Fixes #6517.

ProviderCommandReactor used one global DrainableWorker for lifecycle, turn, response, and stop commands. A provider start/restart that never resolves therefore prevented unrelated threads from starting and left their statuses blank or stuck on Connecting.

This change extends the existing DrainableWorker owner with keyed FIFO lanes. Commands for different threads now process independently; commands for the same thread retain order. Idle lanes remove themselves after draining so the reactor does not retain a queue and fiber for every historical thread.

Focused regression coverage proves independent progress for different keys and FIFO ordering for the same key. The existing ProviderCommandReactor test file still passes (47 tests); the shared worker tests pass (2 tests). Server/shared typechecks, targeted lint, formatting, and done-check --base upstream/main --no-tests passed. Tests were run in the isolated worktree; no live T3 state was used.

Scope disposition: this PR fixes the cross-thread scheduling lockout only. It does not duplicate the still-open interrupt responsiveness work in #6531, nor fold in the separately scoped lifecycle timeout/stale-session/restart-recovery work tracked by #6560, #4944, and #4584. The Codex framing defect remains separate in #5389 and fork coordination issue clintebbesen#1.

Verification environment: Codex harness, gpt-5.6-sol.


Note

Medium Risk
Changes orchestration concurrency for provider lifecycle/turn commands; same-thread ordering is preserved but cross-thread behavior is no longer globally serialized.

Overview
Fixes cross-thread lockout where a single global provider command queue let one thread’s stuck start/restart block every other thread’s lifecycle and turn handling.

Adds makeKeyedDrainableWorker in shared DrainableWorker: one FIFO drainable worker per key, concurrent across keys, FIFO within a key, with idle lanes torn down after drain so historical threads don’t leak queues/fibers. ProviderCommandReactor now keys lanes on event.payload.threadId instead of makeDrainableWorker.

Regression test covers independent progress for different keys and FIFO for the same key.

Reviewed by Cursor Bugbot for commit 4822cc3. Bugbot is set up for automated code reviews on this repo. Configure here.

Note

Isolate provider command processing into per-thread FIFO lanes

  • Adds makeKeyedDrainableWorker to DrainableWorker.ts, which maintains a per-key makeDrainableWorker lane, creating and cleaning up lanes dynamically as work arrives and drains.
  • Updates ProviderCommandReactor.ts to key the worker by event.payload.threadId, so events on different threads are processed concurrently while preserving FIFO order within each thread.
  • Behavioral Change: previously all provider command events shared a single FIFO queue; now concurrency is possible across threads.
📊 Macroscope summarized 4822cc3. 1 file reviewed, 0 issues evaluated, 0 issues filtered, 0 comments posted

🗂️ Filtered Issues

No issues evaluated.

@coderabbitai

coderabbitai Bot commented Aug 15, 2026

Copy link
Copy Markdown

Important

Review skipped

Auto reviews are disabled on this repository. Please check the settings in the CodeRabbit UI or the .coderabbit.yaml file in this repository. To trigger a single review, invoke the @coderabbitai review command.

⚙️ Run configuration

Configuration used: Repository UI

Review profile: CHILL

Plan: Pro Plus

Run ID: c6db509b-8e09-4829-b3fc-91dfc1be273c

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@github-actions github-actions Bot added vouch:unvouched PR author is not yet trusted in the VOUCHED list. size:M 30-99 changed lines (additions + deletions). labels Aug 15, 2026

@cursor cursor Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Cursor Bugbot has reviewed your changes using high effort and found 1 potential issue.

Fix All in Cursor

❌ Bugbot Autofix is OFF. To automatically fix reported issues with cloud agents, enable autofix in the Cursor dashboard.

Want fixes drafted automatically? Bugbot Autofix can create code changes for findings. A team admin can enable Autofix in the Cursor dashboard.

Reviewed by Cursor Bugbot for commit bf77599. Configure here.

Comment thread packages/shared/src/DrainableWorker.ts
@macroscopeapp

macroscopeapp Bot commented Aug 15, 2026

Copy link
Copy Markdown
Contributor

Approvability

Verdict: Needs human review

1 blocking correctness issue found. This PR introduces significant concurrency changes by switching from a single worker queue to per-thread isolated lanes. There is also an unresolved High severity finding about a race condition in lane creation that could break the intended per-key FIFO ordering.

You can customize Macroscope's approvability policy. Learn more.

const enqueue = (item: A): Effect.Effect<void> =>
Effect.gen(function* () {
const key = keyOf(item);
const existing = entries.get(key);

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟠 High src/DrainableWorker.ts:65

Concurrent first enqueues for the same key create separate workers, so their items are processed concurrently instead of in per-key FIFO order; the later entries.set also hides the first lane from drain and idle cleanup. Serialize the keyed lookup, lane creation, and initial enqueue (or reserve the key before yielding) so only one lane can be created per key.

🤖 Copy this AI Prompt to have your agent fix this:
In file @packages/shared/src/DrainableWorker.ts around line 65:

Concurrent first enqueues for the same key create separate workers, so their items are processed concurrently instead of in per-key FIFO order; the later `entries.set` also hides the first lane from `drain` and idle cleanup. Serialize the keyed lookup, lane creation, and initial enqueue (or reserve the key before yielding) so only one lane can be created per key.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size:M 30-99 changed lines (additions + deletions). vouch:unvouched PR author is not yet trusted in the VOUCHED list.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Bug]: Changing runtime mode mid-turn hangs ProviderCommandReactor; all new sessions stay Connecting or show no status

1 participant