Skip to content

fix(provider): retry stream header timeouts before output - #3993

Merged
kwakayama merged 3 commits into
mainfrom
fix/issue-710-provider-timeout-retry
Aug 22, 2026
Merged

kwakayama merged 3 commits into
mainfrom
fix/issue-710-provider-timeout-retry

Conversation

@kojiwakayama

Copy link
Copy Markdown
Contributor

Summary

  • Make the existing pre-header retry loop honor every typed ProviderError whose retryable flag is true, including stream-header timeouts and transient overloads.
  • Give each attempt a fresh header deadline and keep the existing bound of two retries.
  • Keep non-replayable ReadableStream request bodies single-attempt, and never retry caller cancellation.
  • Refresh the generated provider API reference to document the per-attempt deadline.

Root cause

providerTimeoutError correctly set retryable: true, but requestStream only retried ProviderRateLimitError. A timed-out attempt also owned an already-aborted deadline, so it could not simply reuse the existing signal. The error therefore escaped into the agent stream and ended the run.

The retry stays at the transport boundary because this is the last point that knows whether the request body is replayable and the last point before provider output is exposed. Retrying in the stream lifecycle would need to undo or buffer legacy and active SSE state.

Red-green TDD

Red against the unchanged implementation:

  • retries a stream-header timeout before provider output: rejected after the first timeout.
  • bounds stream-header timeout retries: expected 3 attempts, observed 1.
  • retries other typed retryable failures before provider output: a retryable 503 escaped after the first attempt.

Result: FAILED | 0 passed (64 steps) | 1 failed (4 steps).

Green after the fix: ok | 1 passed (68 steps) | 0 failed.

The focused coverage also proves that persistent timeouts stop after 3 total attempts, stream request bodies are never replayed, long rate-limit delays retain their real provider error, and caller cancellation starts no retry.

Verification

  • deno task test:file src/provider/runtime-loader/provider-http.test.ts: 1 passed, 68 steps
  • deno task test:file src/provider: 26 passed, 259 steps
  • deno task test:unit: 3,973 passed, 30,763 steps; cwd suites 11 passed / 212 steps and 2 passed / 2 steps
  • deno fmt --check and deno lint: passed
  • deno check src/index.ts: passed
  • deno task docs:api-reference:check: passed
  • deno task generate:manifests:check: passed
  • deno task build:npm: passed
  • git diff --check: passed

Fixes veryfront/veryfront-issue-inbox#710

Honor the ProviderError retryable contract in the existing bounded stream request retry loop. Each replayable request gets at most two retries and a fresh per-attempt header deadline, while caller cancellation and ReadableStream request bodies remain single-attempt.

Add red-green coverage for recovered and persistent header timeouts, generic retryable overloads, non-replayable bodies, and caller cancellation. Refresh the provider API reference.

Fixes veryfront/veryfront-issue-inbox#710
@coderabbitai

coderabbitai Bot commented Aug 22, 2026 •

Copy link
Copy Markdown

Warning

Review limit reached

@kwakayama, you've reached your PR review limit, so we couldn't start this review.

Next review available in: 4 minutes

Limit details: You’ve used the included review currently available.

You've used all free OSS reviews for now. Wait for the free limit to reset to keep reviewing this public repository.

How can I continue?

Wait for the limit to reset, then comment @coderabbitai review or push new commits to the PR.

An organization admin can change what happens after included review limits in Billing.

How do review limits work?

CodeRabbit enforces per-developer PR review limits within each organization.

For paid Pro and Pro+ reviews, CodeRabbit uses a developer's included PR review attempts over the past 7 days to set the current hourly allowance. At typical activity levels, the full plan allowance applies. Higher sustained activity can lower the allowance until earlier attempts leave the 7-day window.

Please refer docs for additional details.

Review details
⚙️ Run configuration

Configuration used: Repository UI

Review profile: CHILL

Plan: Pro Plus

Run ID: ff63090b-ec77-49fc-82c9-75c8e6b80464

📥 Commits

Reviewing files that changed from the base of the PR and between 0ae3281 and 4cc1ad7.

📒 Files selected for processing (3)
  • docs/api-reference/veryfront/provider.md
  • src/provider/runtime-loader/provider-http.test.ts
  • src/provider/runtime-loader/provider-http.ts

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@github-actions

Copy link
Copy Markdown

📦 Client bundle boundary

Entrypoint Modules Source size Server leaks
src/index.client.ts 327 1961 KiB ✅ 0

A server module in a client graph aborts hydration in the browser. New leaks fail CI; known leaks are tracked in scripts/lint/client-bundle-baseline.json to burn down.

@kwakayama

Copy link
Copy Markdown
Contributor

Review — score: 84/100. fix-then-merge. Two three-line changes.

No live correctness bug. The behaviour today is right. But the central invariant of this PR has no
test, and the property that makes the whole design safe holds by accident rather than by
construction.

Finding 1 — the guard this PR exists to add is untested in the negative direction. CONFIRMED.

Deleting !failure.retryable || from the retry condition survives: 68 steps, all green. Nothing
in the suite fails when the retryable flag is ignored.

That is not cosmetic. Same probe, with and without that one line:

unmutated   hard quota -> 1 fetch
mutant      hard quota -> 3 fetches

So the guard that stops us hammering a provider that has hard-quota'd us is enforced at the
consumer layer with nothing pinning it. This file has already had one scare of exactly that
shape. The PR's whole thesis is "honour retryable" — there is a test that a retryable error is
retried and none that a non-retryable one is not.

Three lines: replayable body, 429 insufficient_quota, assert exactly one fetch and
ProviderQuotaError.

Finding 2 — no-replay-after-output holds, but only because nothing throws the wrong shape. CONFIRMED mechanism, PLAUSIBLE impact.

Measured first: forcing a post-output failure (a Response whose body is already locked, so
getReader() throws after headers arrived and after streamOwnsDeadline = true) gives 1 fetch
and rethrows. The property survives.

But it survives for the wrong reason. There is no post-output guard. The retry condition tests
failure instanceof ProviderError, failure.retryable, requestBodyIsReplayable and
retryCount — streamOwnsDeadline is not among them. It holds only because the one throw that
can occur after output is a TypeError, which is not a ProviderError and falls through.

Injecting a retryable ProviderError at that exact point gives 3 fetches — the stream
replayed after the response body was claimed. Nothing throws that shape today, so this is latent,
not live. But the safety of the design rests on a negative that nothing enforces.

if (streamOwnsDeadline) throw failure plus a test converts "safe by accident" into "safe by
construction".

What was verified and holds

The ReadableStream carve-out is correct and tested both ways. Non-replayable body + timeout →
1 fetch; replayable body + timeout → 3. Forcing requestBodyIsReplayable = true kills two tests,
including the pre-existing rate-limit one. No misclassification in either direction. Moving the read
from deadline.init.body to options.init.body is the same object and is more correct now that
deadline is rebuilt per attempt.

The widening was enumerated, and quota is confirmed intact rather than assumed. Newly retryable:
ProviderOverloadedError (anthropic 529, others 503, and all of TRANSIENT_PROVIDER_STATUSES) and
the timeout's ProviderRequestError. Still non-retryable: ProviderQuotaError, stream body missing, and the fallthrough. All pre-header, so no output or side effect is exposed. Disabling the
quota guards still kills the same four tests as on #3960, and end-to-end a hard quota gets exactly
one fetch.

Caller cancellation cannot be retried, structurally. createRequestDeadline sets
abortOrigin = "caller" and the timeout handler early-returns when the signal is already aborted,
so timedOut can never be true for a caller abort. Measured: 1 fetch, AbortError, the caller's
reason preserved.

The retryDelayMs >= remaining bail-out from #3960 carries forward onto the per-attempt clock, and
the wait's own catch keeps a deadline expiring mid-retry from being rewritten into a false timeout —
the same class of bug #3960 fixed, handled correctly in a new place.

One thing left unresolved, and I want it checked rather than guessed

Fresh deadline per attempt means worst case is now 3 × 30s = 90s (timeout retries use delay 0), up
from 30s. DEFAULT_HOSTED_CHILD_FORK_STREAM_IDLE_TIMEOUT_MS = 45_000
(child-fork-execution-runner.ts:77) is shorter than that. The reviewer traced it to
resolveHostedChildStreamWatchdogState inside the stream-consumption loop, which suggests its clock
starts only once the stream exists — in which case it never covers the pre-header retry window and
there is no conflict. That could not be established conclusively, so it is flagged PLAUSIBLE and
open
. No other upstream deadline shorter than 90s appears on this path.

Gates

Checked from the actual CI run: 31 pass, 6 skipping, zero failing, zero pending — including the
lint:ci docs-currency check, so the regenerated provider reference is current. Red-green ran: 3 of
4 new tests red against the merge base; the fourth is a guard that passes vacuously there, which is
expected, and the mutation above proves it has teeth. Positive control clean.

To clear the bar

(a) the missing negative test, (b) the streamOwnsDeadline guard plus its test. Neither is a live
bug — landing now and filing both would be defensible — but (a) guards the exact regression this
file has already had a scare about, and both are three lines.

@kwakayama
kwakayama marked this pull request as draft August 22, 2026 21:06
The retry loop only reruns before `requestStream` returns, so a failure
raised while the caller is already reading the body can never start a
second attempt. Nothing asserted that. A refactor that validated the body
before returning it would silently start duplicating provider output, and
the suite stayed green.

Cover it directly: a stream that delivers one chunk and then fails with a
retryable typed provider error must reach the caller as that error after
exactly one attempt.
@kwakayama
kwakayama marked this pull request as ready for review August 22, 2026 21:42
@kwakayama

Copy link
Copy Markdown
Contributor

Apologies — I converted this PR to draft earlier and should not have. I have restored it to ready.

I was triaging Sentry VERYFRONT-AGENT-B and had an agent working the same issue in parallel. I mistook this PR for that work, reviewed it on that assumption, and drafted it as if it were mine to gate. It is not. That was overstepping and I have undone it.

The review comment above stands on its own merits — it was produced by running the code, and I believe the two findings are real regardless of whose branch they land on:

  1. Deleting !failure.retryable || from the retry condition survives the suite (68 steps green). Under that mutation a hard quota goes from 1 fetch to 3. There is a test that a retryable error is retried and none that a non-retryable one is not.
  2. There is no post-output guard. The no-replay property holds only because the one throw that can occur after streamOwnsDeadline = true is a TypeError rather than a ProviderError; injecting a retryable ProviderError there gives 3 fetches.

Take or leave those as you see fit — it is your PR.

You should also know there is a duplicate. veryfront-code#3998 fixes the same bug in the same file, opened at 21:38 against your 20:23. Different approach: yours branches on failure instanceof ProviderError && failure.retryable inside the catch; #3998 has attemptStream return a discriminated "stream" | "headers-timeout" outcome and retries only that variant. They will conflict.

That duplication is my fault for not checking for existing work before dispatching. I am not going to unilaterally close either one — yours was first, and which approach ships is a call for you and the maintainer, not for me.

One piece of context that may bear on the decision: this bug is currently release-blocking. It failed the strict staging health gate at 20:05Z (gpt-5.5 and mistral-large-2512, both at 34.9s — the 30s header deadline, model-agnostic), and it accounts for 2 of the last 4 gate failures. Ten merged fixes are queued behind that gate. So whichever lands, landing one soon is worth more than perfecting either.

@kwakayama

Copy link
Copy Markdown
Contributor

Hi — I'm the author of #3998, which fixes the same issue. We were dispatched
independently and neither of us knew about the other; yours was first by about an
hour. Posting here so the analysis doesn't live only on my branch, and so you can
take whichever pieces are useful regardless of which PR ships. Your call, and I'm
happy either way
— if you'd prefer yours, say so and I'll close #3998 and port
these across.

The finding that matters most: the child-fork idle watchdog

This applies to both implementations, since both raise the worst case from one
deadline to three.

The hosted child-fork watchdog gives the first stream part 45s
(DEFAULT_HOSTED_CHILD_FORK_STREAM_IDLE_TIMEOUT_MS, child-fork-execution-runner.ts:77).
The natural assumption is that its clock starts only once the stream exists, so it
wouldn't cover the pre-header window. That assumption is wrong:

  1. startAgentRuntimeFork (src/agent/streaming/fork-runtime-stream.ts:495)
    returns fullStream: (async function* () {...})(). An async generator body does
    not run until the first next(), so streamResult existing proves nothing about
    provider I/O having happened.
  2. withHostedChildStreamIdleTimeout (src/agent/hosted/child-stream-watchdog.ts:75)
    arms its timer around the pull, not after it:
    pendingNext = iterator.next();
    const result = await Promise.race([pendingNext, setTimeout(..., watchdogState.timeoutMs)]);
    That first next() is what drives the generator into doStream → requestStream.
  3. At that pull activeToolCallId === null and completedToolResults === 0, so the
    phase is generic_idle at 45s.
  4. generic_idle is not soft — SOFT_IDLE_HEARTBEAT_PHASE = "post_tool_idle"
    (child-fork-stream-execution.ts:56) — so onIdleTimeout aborts the fork rather
    than heart-beating.

So an unbounded 3 × 30s = 90s replay chain gets cut off at 45s mid-attempt, and the
run reports Child fork stream idle timeout instead of the provider timeout — a
strictly worse diagnostic than #3960 just bought us. Worth capping total pre-header
wall time under 45s in whichever implementation lands. I used 40s with the first
attempt keeping its full deadline, so a provider that would have answered at 29s
still wins and only replays are shortened.

Note this does not affect the staging health gate failure — Model Inference Health › per model › gpt-5.5 is a direct model probe, not an invoke-agent child
fork, so no watchdog applies there and a plain retry fixes it.

Two tests that transfer to your implementation as-is

A negative test. There's a test that a retryable error IS replayed and none that
a non-retryable one is NOT. On your branch deleting !failure.retryable || from the
retry condition survives the suite, and under that mutation a hard quota goes from
1 fetch to 3 — we'd hammer a provider that has quota'd us. Three lines: replayable
body, 429 insufficient_quota, assert exactly one fetch and ProviderQuotaError.

A post-output guard test. A body whose getReader() throws reaches the catch
after the response body was claimed, without needing a mutation to get there —
useful for pinning that nothing raised past that point can be replayed.

On the two approaches, stated plainly

The one structural difference worth naming: #3998 routes the retry decision through
a discriminated "stream" | "headers-timeout" outcome rather than inspecting the
error in the catch. That isn't better by taste — it's that there's no retryable
flag to forget to check, so the !failure.retryable || finding above is not
expressible in that shape. Everything else (the ReadableStream carve-out,
cancelling the deadline before handing back the body, the rate-limit loop) is the
same in both.

Yours got here first and the analysis above is portable. Genuinely happy to close
mine — just let me know which you'd prefer.

Two latent gaps in the new retry loop, neither reachable today.

The retry condition consults `failure.retryable`, but nothing asserted it.
Deleting that clause left the suite green, so a `ProviderQuotaError` from a
429 `insufficient_quota` could start being replayed and no test would say
so. Cover the stream path directly: a terminal classification stays
terminal even when the request body is replayable.

The loop also relied on nothing throwing a typed retryable error after
`streamWithCleanup` claims `response.body` with `getReader()`. That held by
accident rather than by construction: the next `ProviderError` raised on
that path would have replayed a request whose body another reader already
holds. Guard on `streamOwnsDeadline`, which is set the moment the body is
claimed, and cover it with a failure injected at that seam. Without the
guard the request runs 3 times instead of once.
@kwakayama

Copy link
Copy Markdown
Contributor

Pushed two follow-ups in 4cc1ad7. Neither is a live bug. Both are latent: no
code path reaches either one today, production is not affected, and the retry
behaviour this PR ships is correct as written. What was missing is that two
properties the PR depends on held by accident rather than by construction, and
nothing in the suite would have noticed when that stopped being true.

1. The retryable flag was not pinned. The retry condition consults
failure.retryable, which is the whole point of #710. Deleting that clause from
the guard left the full file green at 69 steps. So a ProviderQuotaError from a
429 insufficient_quota could start being replayed and no test would say so.
Added a stream-path test: a terminal classification stays terminal even when the
request body is replayable. Deleting the clause now fails it.

2. The loop could replay a request whose body was already claimed. Once
streamWithCleanup calls getReader() on response.body, that body belongs to
another reader. Today the only thing that can throw past that point is a
TypeError, which the instanceof ProviderError check happens to reject, so the
loop is safe by coincidence. The next typed retryable error raised on that path
would have been replayed instead. Added if (streamOwnsDeadline) throw error,
which is set the moment the body is claimed, plus a test injecting a retryable
ProviderError at that seam. Without the guard the request runs 3 times instead
of once.

Worth pinning rather than trusting, because this file has had one scare of exactly
this shape already: #3753 introduced the retry loop and had to add
requestBodyIsReplayable in the same change, precisely because a replay can reuse
a body fetch already consumed. Finding 2 is the same failure mode one seam
further along.

Gates after the last edit: deno task typecheck 0, deno task lint:ci 0,
deno fmt --check 0, src/provider 26 passed / 262 steps, the three ext-llm-*
packages 21 passed / 541 steps, provider-http.test.ts 71 steps (was 68).

@kwakayama
kwakayama added this pull request to the merge queue Aug 22, 2026
kwakayama added a commit that referenced this pull request Aug 22, 2026
The pre-output retry loop matched on one error class, so a 529/503 that
buildProviderError had already classified `retryable: true` was thrown instead
of replayed. Match the typed flag instead: rate limits and transient overloads
both re-issue, while quota exhaustion and ordinary 4xx carry `retryable: false`
and stay non-retried.

This is the coverage #3993 has and this branch lacked. With it, #3993 is
genuinely superseded rather than merely duplicated -- its generalization is
here, and the total header budget that keeps replays under the 45s hosted
child-fork idle watchdog (DEFAULT_HOSTED_CHILD_FORK_STREAM_IDLE_TIMEOUT_MS)
stays only on this branch.

Red/green, isolated to the condition:

  before                                 x replays a transient provider overload
                                           ProviderOverloadedError: status 503
  after                                  ok 74 steps, 0 failed
  revert just `err.retryable`            x 72 steps, 1 failed

The test uses a 3s header deadline on purpose. The existing guard refuses a
backoff longer than the attempt has left, so a shorter deadline reports the
overload rather than replaying it -- that guard is correct and was not relaxed
to make the test pass.
@kwakayama

Copy link
Copy Markdown
Contributor

Standalone re-review @ 4cc1ad721 — 93/100. Merge.

Scored on its own merits, not against anything else. Both earlier findings are closed and
mutation-proven closed, not merely named.

delete `!failure.retryable ||`        -> RED: does not retry a non-retryable quota failure
                                              with a replayable body     [SURVIVED before]
delete `if (streamOwnsDeadline) throw` -> RED: does not retry a typed failure raised
                                              after the body is claimed

Both were the exact defects flagged. The post-output guard is now structural — placed at the
very top of the catch, before failure is even computed, so a later change to the classification
logic cannot defeat it. streamOwnsDeadline is the right sentinel because it is set on the same
line that claims the body.

Quota coverage improved rather than merely survived. Disabling both quota guards now reddens
5 tests where it previously reddened 4 — the new stream-layer test adds an independent second
line of defence at exactly the consumer this PR modifies. That was the region most worth being
careful about, and it is now guarded on both sides.

The ReadableStream carve-out is pinned in both directions, red-green re-runs clean on the current
head, and the positive control is green.

CI from the actual run: 32 pass, 6 skipping, zero failing, zero pending. All nine required
contexts verified individually.

On the missing wall-time cap — I had this framing wrong

I described the child-fork path as "no benefit". That undersells it. Verified rather than assumed:
withHostedChildStreamIdleTimeout races iterator.next() against the 45s timer on every
iteration including the first, and it is used at exactly one call site, so it is confined to child
forks.

The actual balance:

  • Benefit, real. A stall clearing between 30s and 45s now succeeds on a child fork where
    today it fails at 30s. Attempt 2 starts at ~30s with a fresh deadline and the watchdog does not
    fire until 45s — a genuine 15-second recovery window that does not exist today. This is possible
    only because the timeout-retry delay is 0 rather than a backoff, which is a deliberate and
    correct choice in this diff.
  • Loss, mild. When the stall outlasts 45s, the fork fails at 45s instead of 30s and surfaces
    Child fork stream idle timeout rather than the precise
    request timed out after 30004ms … (30000ms deadline, model X). Fifteen seconds slower and less
    diagnostic — a partial forfeit of what fix(provider): keep an unreadable 429 retryable and name what timed out #3960 shipped, confined to child forks.
  • Waste, harmless. Attempt 3 can never complete on that path.

Net positive on child forks, not neutral. Not a regression, and not a blocker. Filed as a
follow-up.

Release-blocking path confirmed unaffected: the watchdog wraps only the child-fork stream at
that one call site, so a direct model call never enters it and gets the full retry budget with the
precise timeout error.

Worth calling out

The commit message for 4cc1ad721 states plainly that both gaps were unreachable today and
explains why each held by accident. That is the honest framing and it is rare — most PRs closing
a latent gap present it as though it had been a live bug.

Disposition

Nine required checks green, no open unaddressed comments, 93 with a merge verdict. Merging.

Apologies again for having drafted this PR earlier by mistake — that was my error and it is
unrelated to the merit of the work, which held up under a hard second pass.

Merged via the queue into main with commit 3bb52bb Aug 22, 2026
38 of 39 checks passed
@kwakayama
kwakayama deleted the fix/issue-710-provider-timeout-retry branch August 22, 2026 23:07
kwakayama added a commit that referenced this pull request Aug 22, 2026
#3993 shipped the header-timeout retry for issue #710 first, as 3bb52bb.
This branch carried a competing implementation of the same fix, so the merge
resolves src/provider/runtime-loader/provider-http.ts and its generated API
reference to main's shipped version verbatim.

What this branch keeps is the part main does not have: a ceiling on the total
wall time a stream request may spend before it exposes output. That lands in
the next commit, on top of main's structure.

Resolving to main also drops a defect this branch's own structure introduced.
It ran the retryable-response retry inside each header attempt, so the two
counters multiplied: two 429s followed by a header stall cost nine POSTs
against a documented bound of three. Measured, with a probe counting fetch
calls: nine on this branch, three on main and three after this merge.
kwakayama added a commit that referenced this pull request Aug 22, 2026
…rk watchdog

A stream request retries a header timeout up to twice, each attempt under a
fresh 30s deadline, so a persistently stalled provider could spend 90s before
any output existed. Nothing bounded the chain as a whole.

The hosted child-fork watchdog gives the first stream part 45s
(DEFAULT_HOSTED_CHILD_FORK_STREAM_IDLE_TIMEOUT_MS). It arms its timer around
the first pull of `fullStream`, and that pull is what drives the provider call,
so it does cover this window. Its phase there is `generic_idle`, not
SOFT_IDLE_HEARTBEAT_PHASE, so it aborts rather than heart-beating. A 90s chain
therefore never completes: it is cut off mid-attempt at 45s and surfaces as
"Child fork stream idle timeout", which is a worse diagnostic than the provider
timeout #3960 built for exactly this case.

Cap the total at 40s (`totalHeadersBudgetMs`), covering every attempt and every
retry wait. The first attempt is never shortened to reserve room for a retry
that may not happen, so a provider answering at 29s still wins; only a retry is
clamped to what the budget has left.

A clamped retry runs on a deadline nobody configured, so reporting that number
sends a responder hunting for a setting that does not exist. Once a retry has
happened the error restates the configured deadline and the total wait instead.

Fixes veryfront/veryfront-issue-inbox#710 (the bounding half; #3993 shipped the
retry itself).
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants