Skip to content

feat: make the session retry policy configurable - #44517

Open
JoaoBerne wants to merge 4 commits into
anomalyco:devfrom
JoaoBerne:configurable-retry-policy
Open

JoaoBerne wants to merge 4 commits into
anomalyco:devfrom
JoaoBerne:configurable-retry-policy

Conversation

@JoaoBerne

@JoaoBerne JoaoBerne commented Aug 23, 2026 •

Copy link
Copy Markdown

Issue for this PR

Closes #43596

Type of change

  • Bug fix
  • New feature
  • Refactor / code improvement
  • Documentation

What does this PR do?

The retry constants in session/retry.ts are compiled in, so a turn is dropped after 5 attempts and there is no way to change that. This adds an optional experimental.retry block and passes it into SessionRetry.policy. Each field falls back to the constant it replaces, so nothing moves unless you set one.

I exposed backoffFactor as well as maxRetries, because the attempt count alone barely helps. With no response headers the delay is already clamped at 30s from attempt 5 on, so more attempts is mostly more waiting. With headers but no retry-after it isn't clamped at all (#33728), so attempt 10 waits ~17 min. retry-after handling is untouched.

maxRetries: -1 retries while the error stays retryable, which is the long quota window case in the issue. That does hand back the unbounded retry the cap was added to stop (#41848). It's opt-in and off by default, but say the word and I'll drop it and keep only the finite knobs.

maxRetries: 0 is the other end: a failed request is never replayed. The issue has since asked for exactly that, for turns with side effects.

Tuned values can't produce a delay the scheduler mishandles: cap() keeps every delay finite, non-negative and under the setTimeout limit, since Duration.millis turns NaN and negatives into an immediate retry. Details in the review thread below.

How did you verify your code works?

Rebased on dev on 2026-09-23. bun test test/session/ (424 pass), bun test test/config (230 pass), bun turbo typecheck on core and opencode.

Eleven new tests in retry.test.ts. The one that matters: absent and empty tuning both still produce [2000, 4000, 8000, 16000, 30000, 30000], so the default schedule is unchanged.

I also ran script/schema.ts to look at the generated schema. That's why the two float fields use Schema.Finite and not Schema.Number, which was emitting "NaN" and "Infinity" string enums into the public schema.

Screenshots / recordings

Not a UI change.

Checklist

  • I have tested my changes locally
  • I have not included unrelated changes in this PR

@github-actions github-actions Bot added needs:compliance This means the issue will auto-close after 2 hours. and removed needs:compliance This means the issue will auto-close after 2 hours. labels Aug 23, 2026
@github-actions

Copy link
Copy Markdown
Contributor

Thanks for updating your PR! It now meets our contributing guidelines. 👍

@github-actions github-actions Bot added needs:compliance This means the issue will auto-close after 2 hours. and removed needs:compliance This means the issue will auto-close after 2 hours. labels Aug 23, 2026
@github-actions

Copy link
Copy Markdown
Contributor

Thanks for updating your PR! It now meets our contributing guidelines. 👍

@Enough1122

Copy link
Copy Markdown

AI code review — automated review for reference, author can ignore or act on any point.

Overall: clean, well-scoped configurability — resolve() keeps every default in one place, the schedule math is faithfully parameterized (headers vs no-headers caps preserved), and the tests pin both exact default schedules and the new -1 = retry-forever semantic. Two issues and a nit:

  1. packages/core/src/v1/config/config.ts:~206 — maxDelayMs is PositiveInt with no upper bound, but its own default exists precisely because delays feed setTimeout, which breaks above 2^31−1 ms (the old constant's comment in packages/opencode/src/session/retry.ts:31 says exactly that). A user setting maxDelayMs: 4294967296 passes validation, and combined with a large initialDelayMs/backoffFactor (where exponential can reach Infinity) you get an immediate-fire or NaN-ish sleep instead of a capped wait. Add .check(Schema.isLessThanOrEqualTo(2_147_483_647)) (and/or clamp in cap()).

  2. packages/opencode/src/session/retry.ts:policy,~213 — limits is resolved from opts.tuning once, but delay() re-resolves internally on every attempt. Harmless today, but two sources of truth invites drift if one later gains normalization (e.g. clamping). Either pass the resolved Limits into delay or have delay remain the only resolver.

  3. Minor API shape: delay(attempt, error, random, tuning?) now takes tuning as the fourth positional arg — every future caller must remember Math.random() placement (the policy call site already threads it explicitly). An options-object overload (delay(attempt, {error, random, tuning})) would be harder to misuse; not blocking.

Test gap worth one small addition: an upper-bound case with random = 1 so the jitter ceiling (base * (1 + jitterFactor)) is pinned alongside the random = 0 floor cases. Otherwise this is ready — the schema descriptions documenting each default are a nice touch.

@JoaoBerne

Copy link
Copy Markdown
Author

Fixed. Bounded in the schema, and clamped in resolve() as well, since the plugin config hook mutates the loaded config without revalidating it, so the schema can't be the only guard.

Looking for the same shape turned up two more, so cap() now just pins every delay finite and non-negative:

  • a big backoffFactor or jitterFactor overflows the exponential to Infinity, and Infinity * 0 jitter is NaN
  • a malformed retry-after gave a negative delay. The HTTP-date branch guards for that, the two numeric branches didn't. Predates this PR, but 5 attempts kept it cheap and -1 doesn't.

Duration.millis treats NaN and negatives the same as 0, so all three fold into the one guard.

Your 2 goes away with the clamp in resolve(), both sides read identical limits now. Skipped 3, didn't seem worth churning an exported signature the tests already call. Added the random = 1 ceiling case.

One I haven't fixed: -1 plus any zero delay still loops unbounded, and a server sending retry-after: 0 is enough to get there. That wants a delay floor, or no -1 at all. Maintainers' call, same as the offer in the description.

RETRY_MAX_RETRIES and the backoff constants in session/retry.ts are
compiled in, so a turn is abandoned after 5 attempts (~68s) with no way to
change it. That default is right for the infinite-loop reports it was added
for, and wrong for providers whose transient errors outlive it.

Add experimental.retry to the config schema and thread it into
SessionRetry.policy. Every field falls back to the existing constant, so
the schedule is byte-for-byte unchanged when the key is absent. maxRetries
of -1 retries for as long as the error stays retryable.

Expose backoffFactor alongside maxRetries: with the factor left at 2, a
raised attempt count is mostly dead time, since an error carrying headers
but no retry-after is not clamped by RETRY_MAX_DELAY_NO_HEADERS at all.

Refs anomalyco#43596

Assisted-by: Claude Opus 5
Making the constants configurable opened three ways to produce a delay the
scheduler cannot honour. Duration.millis turns every one of them into an
immediate retry, which is the failure mode the attempt cap was added for.

1. maxDelayMs above 2^31-1. That value was the constant precisely because
   setTimeout overflows past it:
     delay() -> 4294967296
     TimeoutOverflowWarning: does not fit into a 32-bit signed integer
     slept 2ms, expected ~49 days
   Bounded in the schema, and normalized in resolve() as well, since the
   plugin config hook mutates the loaded config without revalidation.

2. NaN. A large backoffFactor or jitterFactor overflows the exponential to
   Infinity, and Infinity * 0 jitter is NaN. Reachable at attempt 1 with a
   jitterFactor the schema accepts, and at attempt ~320 with backoffFactor
   10 once maxRetries is -1.

3. Negative. A malformed retry-after already produced a negative delay; the
   HTTP-date branch guards for it, the two numeric branches did not. Bounded
   attempts kept this cheap, -1 does not.

cap() is the single choke point every return path in delay() goes through,
so the finite and non-negative guards live there. Default schedules are
unchanged, pinned by the existing and new tests.

Assisted-by: Claude Opus 5
anomalyco#43596 now also asks for the opposite of what motivated this: turns with
side effects need the processor to never replay a failed provider request.
Zero already did that, since the policy returns done on attempt 1, but
nothing pinned it.

Assisted-by: Claude Opus 5.5
@JoaoBerne
JoaoBerne force-pushed the configurable-retry-policy branch from 4a34561 to c575723 Compare September 23, 2026 16:44

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟡 Changes recommended

Regenerate public SDK/OpenAPI artifacts and normalize invalid maxDelayMs values.

Get a fresh assessment by requesting another Copilot review.

Review effort: Lite
Findings: 1 Medium severity

Open (1)
What changed in this PR

Adds configurable retry policy settings under experimental.retry, preserving default behavior and adding test coverage.

Changes:

  • Configures retry limits, delays, backoff, jitter, and caps.
  • Passes retry configuration into session processing.
  • Adds schema definitions and retry behavior tests.
File Description
packages/​opencode/​test/​session/​retry.test.ts Tests default and customized retry behavior.
packages/​opencode/​src/​session/​retry.ts Implements configurable retry limits and delays.
packages/​opencode/​src/​session/​processor.ts Passes retry configuration into session execution.
packages/​core/​src/​v1/​config/​config.ts Defines the experimental.retry schema.

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Comment on lines +47 to +59
function resolve(tuning?: Tuning) {
return {
maxRetries: tuning?.maxRetries ?? RETRY_MAX_RETRIES,
initialDelayMs: tuning?.initialDelayMs ?? RETRY_INITIAL_DELAY,
backoffFactor: tuning?.backoffFactor ?? RETRY_BACKOFF_FACTOR,
jitterFactor: tuning?.jitterFactor ?? RETRY_JITTER_FACTOR,
// Normalized here so policy() and delay() read the same ceiling. Config
// validation also bounds this, but a plugin config hook mutates the loaded
// config without revalidation, so it cannot be the only guard.
maxDelayMs: Math.min(tuning?.maxDelayMs ?? RETRY_MAX_DELAY, RETRY_MAX_DELAY),
maxDelayNoHeadersMs: tuning?.maxDelayNoHeadersMs ?? RETRY_MAX_DELAY_NO_HEADERS,
}
}

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Right, and it's the argument I'd made in the comment just above it: the plugin hook skips revalidation, so an upper bound alone wasn't enough. Fixed in resolve() rather than for maxDelayMs alone, since every field has the same exposure. Each one now falls back to its default unless it passes the same rule as the schema. That also catches maxRetries: NaN, which the old check read as unlimited. Tests for both.

On regenerating the SDK and openapi.json: generate.yml does that on every push to dev, so I've left the generated files alone.

…'s ceiling

The plugin config hook mutates the loaded config without revalidation, so
schema bounds cannot be the only guard. resolve() only enforced an upper
bound on maxDelayMs: a hook setting it to -1 or NaN made cap() return a
negative or NaN delay, which Duration.millis turns into an immediate retry
(Copilot review). The same exposure applied to every other field, and one
was worse: a NaN maxRetries failed the `>= 0` test and read as unlimited.

Each field now falls back to its default unless it passes the same rule as
the schema. Generated SDK and openapi.json are left to generate.yml, which
regenerates them on every push to dev.

Assisted-by: Claude Opus 5.5
@kleinware

Copy link
Copy Markdown

I just built a similar change to support max retry duration (effectively same as maxRetries*maxDelayNoHeadersMs for long outages). I'd rather have this change as its more robust (my use case only needed a max wait time).

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Configurable retry policy: expose maxRetries / initialDelay / backoffFactor / maxDelay via config

4 participants