Skip to content

fix: retry worker response delivery on transient gRPC failures - #365

Merged
wangbill (YunchuWang) merged 8 commits into
mainfrom
yunchuwang-worker-response-retries
Sep 11, 2026
Merged

wangbill (YunchuWang) merged 8 commits into
mainfrom
yunchuwang-worker-response-retries

Conversation

@YunchuWang

@YunchuWang wangbill (YunchuWang) commented Sep 9, 2026 •

Copy link
Copy Markdown
Member

Summary

Why is this change needed?

An activity can finish successfully while reporting its result fails. Transport retries do not cover every such failure: after the server sends response headers, gRPC will not retry that call.

Customer experience Before After
An activity computes 42, but its completion report encounters a transient error after response headers The worker logs the delivery failure and makes no further SDK send. The workflow may not advance until backend recovery/redelivery. The worker resends the same result and completion token. If a resend is accepted, the workflow can advance without this retry loop executing the activity again.
The failure persists or is not retryable Existing error reporting applies. Attempts are bounded; the final failure uses the existing error reporting.

This is response-delivery retry parity with .NET, not a new execution or connection-lifetime policy. Backend work remains at-least-once; this does not promise exactly-once execution or successful recovery from every outage.

What changed?

  • One shared send helper covers existing orchestration, activity, entity, and version-rejection completion/abandon paths.
  • Match the .NET retry loop: at most 10 SDK attempts for UNAVAILABLE, UNKNOWN, DEADLINE_EXCEEDED, or INTERNAL; 200 ms exponential backoff capped at 15 seconds before adding 0-20% jitter.
  • Capture the existing worker-run cancellation signal when work is dispatched. Stop cancels retry waits/RPCs while preserving JS's existing graceful first-response behavior.
  • Preserve configured gRPC transport retries. Ten SDK attempts can therefore involve more than ten network requests.

Not included: channel retention or per-channel task tracking, changes to the existing 30-second channel-retirement delay, metadata cancellation, history features, new public settings, or workflow changes.

Files to review

File Purpose
task-hub-grpc-worker.ts Shared result-send retry loop and existing-run signal plumbing.
backoff.util.ts Small positive-jitter option matching the .NET retry policy; existing defaults are unchanged.
Two worker response test files Policy, existing response paths, shutdown compatibility, and real gRPC recovery.
README.md, CHANGELOG.md Customer behavior and release note.

Issues / work items

Narrowed in place from the earlier broad implementation in this PR. No companion feature is required.


Project checklist

  • Release notes are not required for the next release
    • Otherwise: Notes added to CHANGELOG.md
  • Backport is not required
  • All required tests have been added/updated
  • Breaking change?
    • No public API or protocol changes. Existing transport configuration and channel retirement are preserved.

AI-assisted code disclosure (required)

Was an AI tool used?

  • No

  • Yes, AI helped write parts of this PR

  • Yes, an AI agent generated most of this PR

  • Tool: GitHub Copilot.

  • AI-assisted areas: implementation, tests, documentation.

  • Changes after initial generation: removed unrelated connection retention, metadata cancellation, additional controllers/maps, future history helpers, and workflow wiring. Kept only the retry policy and necessary JS compatibility plumbing.

AI verification:

  • I understand the code and can explain it (human confirmation pending)
  • I verified referenced APIs/types exist and are correct
  • I reviewed edge cases/failure paths
  • I reviewed concurrency/async behavior
  • I checked for unintended breaking or behavior changes

Testing

Current head: da2633cc9bf800998b598dfec45066f65164d746.

  • Current-head CI: 29 successful checks, 1 intentional skip, none pending or failed. Test and Build, DTS Emulator E2E, Sample validation, and both CodeQL workflows passed. The clean npm ci unit job explicitly ran both worker response test files.
  • Real gRPC before/after: with a five-attempt transport retry policy configured, the server sends headers and then fails the completion. Main sends once; this implementation sends twice, reuses the payload/token, and executes the activity once.
  • A separate socket case preserves transport layering: 10 SDK sends, up to 50 server calls with that five-attempt policy.
  • Local targeted run: 207 tests across 12 suites; 32 focused tests rerun after the socket-test refinement. Core build and scoped lint/format passed.
  • Local dependency restoration reused an existing install after the configured feed returned E404; it was not a pristine current-lockfile install. The subsequent clean-install CI above passed at the published head.
  • No new manual Azure or emulator run is claimed.

Notes for reviewers

Keeping an old channel alive until all its tasks finish is not the .NET behavior and is deliberately excluded. Retried requests can still fail when a channel closes or a backend token expires.

Retry completion and abandon responses without re-executing user code. Preserve graceful initial delivery and original-stub ownership while cancelling retries on stop. Disable only worker transport retries to avoid multiplying delivery budgets.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Copilot-Session: e0af01a5-0dfa-4e71-a660-c4186e65d7e0
Include both real-socket and persisted-backend delivery specs in the existing emulator group and provide its local connection string so backend cases cannot silently skip.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Copilot-Session: e0af01a5-0dfa-4e71-a660-c4186e65d7e0
Match the .NET Azure-managed worker's layered retry behavior. Bound SDK delivery attempts without overriding caller channel options, document the combined budget, and prove ten SDK attempts with fifty actual server calls under a five-attempt channel policy.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Copilot-Session: e0af01a5-0dfa-4e71-a660-c4186e65d7e0
Observe cancellation before generating metadata and throughout the wait, retaining the initial-response drain window. Prevent late metadata from sending an RPC and cover pending-work cleanup with deterministic and real-gRPC shutdown regressions.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>

Copilot-Session: e0af01a5-0dfa-4e71-a660-c4186e65d7e0
Route version rejection through a private abandonment primitive shared with the standalone history PR. Allow an explicit cancellation signal to govern initial delivery, retries, and backoff without changing default graceful completion behavior.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Copilot-Session: e0af01a5-0dfa-4e71-a660-c4186e65d7e0
Keep completion callbacks typed to CompleteTaskResponse by default, allow the shared response-path mock to model the abandon response, and instantiate AbandonOrchestrationTaskResponse for abandonment.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Copilot-Session: e0af01a5-0dfa-4e71-a660-c4186e65d7e0
Preserve metadata cancellation alongside the client wait interceptor-start deferred cancellation fix. Retain both client and worker changelog entries.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Copilot-Session: e0af01a5-0dfa-4e71-a660-c4186e65d7e0
Copilot-Session: 186eb204-3a81-49e7-9574-87b9962d5648
Copilot AI lite review requested due to automatic review settings September 9, 2026 23:12

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🔵 Needs a closer look

Retry, cancellation, reconnect, and transport-layer interactions warrant final human review.

Pull request overview

Adds bounded, cancellable retries for worker response delivery while preserving graceful shutdown and transport retry behavior.

Changes:

  • Centralizes response and abandonment retries with backoff and jitter.
  • Adds cancellation-aware metadata and RPC handling.
  • Expands unit, gRPC, emulator, Azure, documentation, changelog, and CI coverage.
File summaries
File Summary
test/e2e-azuremanaged/worker-response-delivery.spec.ts Adds Azure delivery coverage.
README.md Documents retry and shutdown behavior.
packages/durabletask-js/test/worker-response-delivery.spec.ts Adds unit coverage.
packages/durabletask-js/test/worker-response-delivery-grpc.spec.ts Adds real-gRPC coverage.
packages/durabletask-js/src/worker/task-hub-grpc-worker.ts Implements delivery retries and lifecycle handling.
packages/durabletask-js/src/utils/grpc-helper.util.ts Adds cancellation-aware metadata and RPC handling.
packages/durabletask-js/src/utils/backoff.util.ts Adds positive jitter support.
CHANGELOG.md Records the worker delivery fix.
.github/workflows/dts-e2e-tests.yaml Enables delivery test coverage in CI.
Review details
  • Files reviewed: 9/9 changed files
  • Comments generated: 0
  • Review effort level: Lite

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🔵 Needs a closer look

The concurrent shutdown, cancellation, retry, and channel-replacement lifecycle warrants final human review despite comprehensive coverage.

Review details
  • Files reviewed: 9/9 changed files
  • Comments generated: 0 new
  • Review effort level: Balanced

Keep bounded retries of the computed response and capture retry cancellation at work dispatch. Restore baseline channel retirement and metadata handling; defer unrelated lifecycle and history changes.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Copilot-Session: e0af01a5-0dfa-4e71-a660-c4186e65d7e0
@YunchuWang wangbill (YunchuWang) changed the title fix: add bounded cancellable worker response delivery retries fix: retry worker response delivery on transient gRPC failures Sep 10, 2026
@YunchuWang
wangbill (YunchuWang) merged commit eab9340 into main Sep 11, 2026
30 checks passed
@YunchuWang
wangbill (YunchuWang) deleted the yunchuwang-worker-response-retries branch September 11, 2026 15:39
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants