fix(core): recover partial provider failures - #44002
kitlangton wants to merge 6 commits into
Conversation
4dbc250 to
1b0435e
Compare
— automated review (ox-alpha, round2) |
|
Automated PR Cleanup Thank you for contributing to opencode. Due to the high volume of PRs from users and AI agents, we periodically close older PRs using automated criteria so maintainers can focus review time on the most active and community-supported contributions. This PR was closed because it matched the following cleanup criteria:
PRs created within the last month are not affected by this cleanup. If you believe this PR was closed incorrectly, or if you are still actively working on it, please leave a comment explaining why it should be reopened. A maintainer can review and reopen it if appropriate. Thanks again for taking the time to contribute. |
What
Automatically recover retryable provider-internal and rate-limit failures that arrive after partial model output. Recovery can cross eagerly executed local tools once their outcomes are durable, but stops at provider-hosted activity that cannot be replayed uniformly.
Before / After
Before: A provider could emit reasoning, text, or a local tool call and then fail with a retryable provider error. Because output had already started, the ordinary transparent retry path no longer applied, so the assistant response ended in error and the user had to prompt again manually.
After: The runner joins eager local tool executions, durably settles every observed call, and then schedules its existing bounded retry for incomplete streams, transport-read failures, retryable rate limits, and retryable provider-internal failures. It records the failed attempt, appends the existing continuation prompt, reloads projected history, and starts a fresh model request. Durable local tool results are included, so side effects are not repeated. Provider-hosted activity remains terminal because hosted results are not uniformly replayable across protocols and storage modes.
How
packages/core/src/session/runner/llm.tsclassifies eligible post-output failures separately from transparent pre-output retries and requires every observed call to be settled and locally executed.packages/ai/src/protocols/openai-responses.tssurfaces hosted activity atresponse.output_item.added, ensuring a stream failure cannot hide a provider-side tool boundary beforeoutput_item.done.Scope
Testing
bun run test test/session-runner.test.tsfrompackages/core: 165 passedbun run testfrompackages/core: 1950 passedbun run test test/provider/openai-responses.test.tsfrompackages/ai: 94 passedbun typecheckfrompackages/core: passedbun typecheckfrompackages/ai: passedgit diff --check origin/v2: passedpackages/aisuite: 542 passed, 6 unrelated existing OpenRouter fixture/reasoning failuresFlow
flowchart TD A[Retryable failure after output starts] --> B[Join eager local tools] B --> C[Durably close every observed call] C --> D{Failure eligible?} D -->|No| E[Keep failure terminal] D -->|Yes| F{Any hosted or unsettled call?} F -->|Yes| E F -->|No| G[Schedule bounded retry] G --> H[Append continuation prompt] H --> I[Reload projected history and local tool results] I --> J[Start fresh model request]