fix: retry worker response delivery on transient gRPC failures - #365
Merged
Merged
Conversation
Retry completion and abandon responses without re-executing user code. Preserve graceful initial delivery and original-stub ownership while cancelling retries on stop. Disable only worker transport retries to avoid multiplying delivery budgets. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: e0af01a5-0dfa-4e71-a660-c4186e65d7e0
Include both real-socket and persisted-backend delivery specs in the existing emulator group and provide its local connection string so backend cases cannot silently skip. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: e0af01a5-0dfa-4e71-a660-c4186e65d7e0
Match the .NET Azure-managed worker's layered retry behavior. Bound SDK delivery attempts without overriding caller channel options, document the combined budget, and prove ten SDK attempts with fifty actual server calls under a five-attempt channel policy. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: e0af01a5-0dfa-4e71-a660-c4186e65d7e0
Observe cancellation before generating metadata and throughout the wait, retaining the initial-response drain window. Prevent late metadata from sending an RPC and cover pending-work cleanup with deterministic and real-gRPC shutdown regressions. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: e0af01a5-0dfa-4e71-a660-c4186e65d7e0
Route version rejection through a private abandonment primitive shared with the standalone history PR. Allow an explicit cancellation signal to govern initial delivery, retries, and backoff without changing default graceful completion behavior. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: e0af01a5-0dfa-4e71-a660-c4186e65d7e0
Keep completion callbacks typed to CompleteTaskResponse by default, allow the shared response-path mock to model the abandon response, and instantiate AbandonOrchestrationTaskResponse for abandonment. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: e0af01a5-0dfa-4e71-a660-c4186e65d7e0
Preserve metadata cancellation alongside the client wait interceptor-start deferred cancellation fix. Retain both client and worker changelog entries. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: e0af01a5-0dfa-4e71-a660-c4186e65d7e0 Copilot-Session: 186eb204-3a81-49e7-9574-87b9962d5648
Contributor
There was a problem hiding this comment.
🔵 Needs a closer look
Retry, cancellation, reconnect, and transport-layer interactions warrant final human review.
Pull request overview
Adds bounded, cancellable retries for worker response delivery while preserving graceful shutdown and transport retry behavior.
Changes:
- Centralizes response and abandonment retries with backoff and jitter.
- Adds cancellation-aware metadata and RPC handling.
- Expands unit, gRPC, emulator, Azure, documentation, changelog, and CI coverage.
File summaries
| File | Summary |
|---|---|
test/e2e-azuremanaged/worker-response-delivery.spec.ts |
Adds Azure delivery coverage. |
README.md |
Documents retry and shutdown behavior. |
packages/durabletask-js/test/worker-response-delivery.spec.ts |
Adds unit coverage. |
packages/durabletask-js/test/worker-response-delivery-grpc.spec.ts |
Adds real-gRPC coverage. |
packages/durabletask-js/src/worker/task-hub-grpc-worker.ts |
Implements delivery retries and lifecycle handling. |
packages/durabletask-js/src/utils/grpc-helper.util.ts |
Adds cancellation-aware metadata and RPC handling. |
packages/durabletask-js/src/utils/backoff.util.ts |
Adds positive jitter support. |
CHANGELOG.md |
Records the worker delivery fix. |
.github/workflows/dts-e2e-tests.yaml |
Enables delivery test coverage in CI. |
Review details
- Files reviewed: 9/9 changed files
- Comments generated: 0
- Review effort level: Lite
💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.
Contributor
There was a problem hiding this comment.
🔵 Needs a closer look
The concurrent shutdown, cancellation, retry, and channel-replacement lifecycle warrants final human review despite comprehensive coverage.
Review details
- Files reviewed: 9/9 changed files
- Comments generated: 0 new
- Review effort level: Balanced
Keep bounded retries of the computed response and capture retry cancellation at work dispatch. Restore baseline channel retirement and metadata handling; defer unrelated lifecycle and history changes. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: e0af01a5-0dfa-4e71-a660-c4186e65d7e0
9 of 13 tasks
kaibocai (kaibocai)
approved these changes
Sep 11, 2026
wangbill (YunchuWang)
deleted the
yunchuwang-worker-response-retries
branch
September 11, 2026 15:39
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Why is this change needed?
An activity can finish successfully while reporting its result fails. Transport retries do not cover every such failure: after the server sends response headers, gRPC will not retry that call.
42, but its completion report encounters a transient error after response headersThis is response-delivery retry parity with .NET, not a new execution or connection-lifetime policy. Backend work remains at-least-once; this does not promise exactly-once execution or successful recovery from every outage.
What changed?
UNAVAILABLE,UNKNOWN,DEADLINE_EXCEEDED, orINTERNAL; 200 ms exponential backoff capped at 15 seconds before adding 0-20% jitter.Not included: channel retention or per-channel task tracking, changes to the existing 30-second channel-retirement delay, metadata cancellation, history features, new public settings, or workflow changes.
Files to review
task-hub-grpc-worker.tsbackoff.util.tsREADME.md,CHANGELOG.mdIssues / work items
Narrowed in place from the earlier broad implementation in this PR. No companion feature is required.
Project checklist
CHANGELOG.mdAI-assisted code disclosure (required)
Was an AI tool used?
No
Yes, AI helped write parts of this PR
Yes, an AI agent generated most of this PR
Tool: GitHub Copilot.
AI-assisted areas: implementation, tests, documentation.
Changes after initial generation: removed unrelated connection retention, metadata cancellation, additional controllers/maps, future history helpers, and workflow wiring. Kept only the retry policy and necessary JS compatibility plumbing.
AI verification:
Testing
Current head:
da2633cc9bf800998b598dfec45066f65164d746.npm ciunit job explicitly ran both worker response test files.Notes for reviewers
Keeping an old channel alive until all its tasks finish is not the .NET behavior and is deliberately excluded. Retried requests can still fail when a channel closes or a backend token expires.