fix: InstanceState reset-caused deadlock/hang during subagent return - #50667
RainbowLazer wants to merge 2 commits into
Conversation
- OpenCode has a bug identified by Claude/Copilot in which instance state can reload while a subagent is working, which prevents the original agent from ever getting a response—it checks for stale state that no longer remains. - Add Claude-generated defenses and tests for this bug - Claude summary of changes: Three-layer defense against the instance lifecycle deadlock: 1. task.ts — Release handler cleanup forked to a daemon fiber, breaking the re-entrant ScopedCache deadlock chain. Safe because SessionRunState.cancel already calls cancelBackgroundJobs before interrupting runners. 2. instance-registry.ts — 10-second disposal timeout added as a safety net. Even if an unforeseen deadlock occurs in any disposer, disposal completes and the new instance can boot. 3. core/background-job.ts — Scope-close finalizer resolves all pending done Deferreds with cancelled status. Prevents zombie waiters from blocking forever on a Deferred that will never be resolved. Tests added: - test/project/instance-lifecycle.test.ts — 4 tests covering disposal/reload with active background jobs, fiber blocked on background.wait, and disposal timeout behavior. - test/background/job.test.ts — 1 new test verifying scope-close resolves pending done Deferreds.
Claude summary:
- task.ts: Effect.forkDaemon → Effect.forkDetach
- job.test.ts: Effect.forkDaemon → Effect.forkDetach
- instance-lifecycle.test.ts: Effect.forkDaemon → Effect.forkDetach,
and rebuilt the layer to use
Layer.mergeAll(InstanceStore.defaultLayer,
BackgroundJob.defaultLayer,
CrossSpawnSpawner.defaultLayer
).pipe(Layer.provide(noopBootstrap))
so InstanceStore.Service is a provider
(not just a requirement-satisfier)
|
Hey! Your PR title Please update it to start with one of:
Where See CONTRIBUTING.md for details. |
|
The following comment was made by an LLM, it may be inaccurate: Based on my search results, the only PR that appears is the current PR #50667 itself. While I found some related PRs in the cleanup fiber and subagent context (#45482 and #45822), they are not duplicates of this PR—they address different aspects of the system. No duplicate PRs found |
|
Thanks for your contribution! This PR doesn't have a linked issue. All PRs must reference an existing issue. Please:
See CONTRIBUTING.md for details. |
|
Thanks for updating your PR! It now meets our contributing guidelines. 👍 |
Issue for this PR
Closes #50671
Type of change
What does this PR do?
Please provide a description of the issue, the changes you made to fix it, and why they work. It is expected that you understand why your changes work and if you do not understand why at least say as much so a maintainer knows how much to value the PR.
If you paste a large clearly AI generated description here your PR may be IGNORED or CLOSED!
This PR fixes a bug where InstanceState can be reset while a subagent is working, causing it to never return back to the main agent—it checks for the old state instead of the new state and deadlocks with the cleanup logic. New code is Claude-generated and I only vaguely understand it. Adds three levels of defense for the bug: forking of the cleanup fiber so it doesn't cause a deadlock with the cache access otherwise hitting the old state, 10 second timeout to override any deadlocks that occur anyways, and automatic resolution of deferred tasks as cancelled when exiting a scope so they cannot wait forever.
How did you verify your code works?
Added and ran (locally) Claude-generated tests that exercise the bug, e.g., by forcing InstanceState reload with a background task running and ensuring the client does not hang.
Screenshots / recordings
If this is a UI change, please include a screenshot or recording.
Checklist
If you do not follow this template your PR will be automatically rejected.