fix: add bounded cancellable worker response delivery retries - #364
wangbill (YunchuWang) wants to merge 6 commits into
Conversation
Retry completion and abandon responses without re-executing user code. Preserve graceful initial delivery and original-stub ownership while cancelling retries on stop. Disable only worker transport retries to avoid multiplying delivery budgets. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: e0af01a5-0dfa-4e71-a660-c4186e65d7e0
Include both real-socket and persisted-backend delivery specs in the existing emulator group and provide its local connection string so backend cases cannot silently skip. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: e0af01a5-0dfa-4e71-a660-c4186e65d7e0
There was a problem hiding this comment.
🟡 Changes recommended
A new unit test mocks abandonTaskOrchestratorWorkItem with the wrong response type, which can mask contract/behavior regressions and should be corrected.
Once you've addressed the issues Copilot identified, you can request another Copilot review.
Pull request overview
This PR adds a bounded, cancellable SDK-level retry loop for worker response delivery (orchestration/activity/entity completion plus version-mismatch failure/abandon), aligning retry semantics with the pinned .NET worker while preventing gRPC transport retries from multiplying the effective retry budget.
Changes:
- Introduces a shared worker response-delivery helper with transient-status retry + capped exponential backoff with jitter, and stop-aware cancellation semantics.
- Forces worker-created gRPC channels to disable transport retries (
grpc.enable_retries: 0) so the worker owns the retry budget. - Adds unit + real-socket gRPC tests and opt-in Azure-managed E2E coverage; documents delivery/shutdown behavior and updates changelog.
File summaries
| File | Description |
|---|---|
| test/e2e-azuremanaged/worker-response-delivery.spec.ts | Opt-in Azure-managed E2E verifying persistence + interceptor-injected delivery faults without re-executing user code. |
| README.md | Documents the worker response delivery retry policy and shutdown/cancellation behavior. |
| packages/durabletask-js/test/worker-startup.spec.ts | Updates expectations for worker channel options to include disabled transport retries. |
| packages/durabletask-js/test/worker-response-delivery.spec.ts | Adds unit tests for retry policy, cancellation behavior, and stub-retirement interactions. |
| packages/durabletask-js/test/worker-response-delivery-grpc.spec.ts | Adds real grpc-js socket tests covering post-header failures, stop behavior, and retry budget enforcement. |
| packages/durabletask-js/src/worker/task-hub-grpc-worker.ts | Implements response delivery retry helper, stop-aware signals, and per-stub pending-work tracking for safe stub retirement. |
| packages/durabletask-js/src/utils/backoff.util.ts | Adds “positive” jitter strategy used by worker delivery retries. |
| CHANGELOG.md | Notes the new worker response delivery retry behavior and transport-retry override. |
Review details
Suppressed comments (1)
packages/durabletask-js/test/worker-response-delivery.spec.ts:368
- This mock always replies with CompleteTaskResponse, even when the spied method is abandonTaskOrchestratorWorkItem (which should respond with AbandonOrchestrationTaskResponse). Returning the correct response type keeps the test faithful to the gRPC contract and avoids masking future response-handling logic changes.
const respond = typeof optionsOrCallback === "function" ? optionsOrCallback : callback!;
respond(requests.length === 1 ? grpcError(grpc.status.INTERNAL) : null, new pb.CompleteTaskResponse());
return unaryCall();
- Files reviewed: 9/9 changed files
- Comments generated: 1
- Review effort level: Lite
💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.
Match the .NET Azure-managed worker's layered retry behavior. Bound SDK delivery attempts without overriding caller channel options, document the combined budget, and prove ten SDK attempts with fifty actual server calls under a five-attempt channel policy. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: e0af01a5-0dfa-4e71-a660-c4186e65d7e0
Observe cancellation before generating metadata and throughout the wait, retaining the initial-response drain window. Prevent late metadata from sending an RPC and cover pending-work cleanup with deterministic and real-gRPC shutdown regressions. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: e0af01a5-0dfa-4e71-a660-c4186e65d7e0
Route version rejection through a private abandonment primitive shared with the standalone history PR. Allow an explicit cancellation signal to govern initial delivery, retries, and backoff without changing default graceful completion behavior. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: e0af01a5-0dfa-4e71-a660-c4186e65d7e0
Keep completion callbacks typed to CompleteTaskResponse by default, allow the shared response-path mock to model the abandon response, and instantiate AbandonOrchestrationTaskResponse for abandonment. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: e0af01a5-0dfa-4e71-a660-c4186e65d7e0
Summary
What changed?
UNAVAILABLE,UNKNOWN,DEADLINE_EXCEEDED, andINTERNAL: 200 ms exponential backoff, capped at 15 seconds before +0–20% jitter. Permanent/exhausted failures retain existing error logs._abandonOrchestrationWorkItem(stub, completionToken, signal?), backed by the same delivery loop. An explicit signal governs initial attempts, retries and backoff; omission preserves graceful-initial/retry-stop behavior. This provides the shared primitive for the independently based history PR, not another retry framework.Why is this change needed?
Azure-managed workers already configure transport retries, normally for UNAVAILABLE. Core completion/abandon sites lack a bounded SDK delivery loop, including after headers commit and for additional .NET transient statuses. This closes that gap without changing persisted history or the backend's at-least-once work-item delivery contract.
Pinned reference:
microsoft/durabletask-dotnet@bc2bc12ca5ee3a12a6e633250ade5efe6ae90ef4, processorExecuteWithRetryAsync, internal options, andGrpcBackoff.cs. Its Azure-managedCreateChannelalso retainsGrpcRetryPolicyDefaults.DefaultServiceConfigbeneath that loop, confirming the layered retry behavior.Base: upstream main
28730dfaedaecd696468cae2e2365153b230fa56. Final tested head:4ce281c32f46ca806fcd39acc8f86b2a8b1e7abe. All final-head CI workflows passed. The coordinator's principal source re-review approved this exact head, including the shared-abandonment delta; no remaining source blocker.Issues / work items
Project checklist
CHANGELOG.mdAI-assisted code disclosure (required)
Was an AI tool used? (select one)
If AI was used:
AI verification (required if AI was used):
Testing
Automated tests
npx --no-install jest --runInBand --runTestsByPath ... --detectOpenHandles. Covers affected worker/versioning/entity/tracing/host-integration/backoff behavior, Azure-managed builder/retry/options/endpoint behavior, and client regressions for the shared unary helper.npm run build:core,npm run build:azuremanaged, changed TS ESLint/Prettier andgit diff --check: passed. Changelog/workflow formatting passed previously and those files are unchanged by the latest delta. README's pre-existing whole-file formatting was not rewritten.Manual validation (only if runtime/behavior changed)
workerresponsetask hub, Azure CLI authentication. No Azure resources created/deleted; no combined-run hub access.npx --no-install jest --runInBand --runTestsByPath test/e2e-azuremanaged/worker-response-delivery.spec.ts --detectOpenHandles, with the dedicatedWORKER_DELIVERY_CONNECTION_STRING. 2 passed.Real Azure baseline — no injection
response-delivery-2ba5d9c0-a76e-49e1-905d-1985050d0a41.@deliverycounter-response-delivery-2ba5d9c0-a76e-49e1-905d-1985050d0a41@counter.Completed, output42, entity state42; activity/entity each executed once.Real Azure persisted outcome with controlled client-side faults
response-delivery-1b138691-3b2a-40aa-893e-4b2db2ba5c41.@deliverycounter-response-delivery-1b138691-3b2a-40aa-893e-4b2db2ba5c41@counter.Completed, output42, entity state42; activity/entity each executed once.Notes for reviewers
private async _abandonOrchestrationWorkItem(stub, completionToken, signal?): Promise<void>. Keep this PR's retry-backed helper body and both history-PR callers. The history-fetch-error caller supplies its immediate-stop signal; version rejection omits it. Do not retain the history PR's standalone direct-callWithMetadatahelper body in the combined tree, or that path bypasses response retries. Combined history-error + transient-abandon fault coverage is owned by integration; this standalone PR does not add history fetching.npx --no-install lint-stagedcheck passed when run directly before committing.