feat(worker): align history streaming with .NET - #363
wangbill (YunchuWang) merged 11 commits into
Conversation
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: e0af01a5-0dfa-4e71-a660-c4186e65d7e0
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: e0af01a5-0dfa-4e71-a660-c4186e65d7e0
There was a problem hiding this comment.
🟡 Changes recommended
The new worker history streaming test leaks global OpenTelemetry state by disabling tracing instead of restoring the previous tracer provider, which can make other test suites flaky.
Once you've addressed the issues Copilot identified, you can request another Copilot review.
Pull request overview
This PR adds worker-side support for service-selected orchestration history streaming by hydrating full past history via StreamInstanceHistory before version checks, tracing, and replay, and it introduces new tests (local gRPC + opt-in Azure) to validate correctness and shutdown/transport-failure behavior.
Changes:
- Advertises
WORKER_CAPABILITY_HISTORY_STREAMINGinGetWorkItemsand hydrates streamed orchestration history before replay. - Adds shutdown/cleanup handling for in-flight history hydration and abandons work items on incomplete history transport to enable redelivery.
- Adds comprehensive local gRPC tests plus an opt-in bounded Azure E2E test; updates changelog.
File summaries
| File | Description |
|---|---|
packages/durabletask-js/src/worker/task-hub-grpc-worker.ts |
Implements streamed history hydration, cancellation tracking, capability advertisement, and abandon-on-transport-failure behavior. |
packages/durabletask-js/test/worker-history-streaming.spec.ts |
Adds local ephemeral gRPC server tests covering streaming ordering, failures/redelivery, shutdown cancellation, and trace continuity. |
test/e2e-azuremanaged/worker-history-streaming.spec.ts |
Adds opt-in Azure-managed negotiation test that passively observes real service-selected streaming and validates durable outcome replay. |
CHANGELOG.md |
Documents the new worker history streaming/hydration behavior and its failure handling semantics. |
Review details
- Files reviewed: 4/4 changed files
- Comments generated: 1
- Review effort level: Lite
💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: e0af01a5-0dfa-4e71-a660-c4186e65d7e0
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: e0af01a5-0dfa-4e71-a660-c4186e65d7e0
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: e0af01a5-0dfa-4e71-a660-c4186e65d7e0 Copilot-Session: 7b853e01-1e42-41c2-92a7-e2bbab87d215
Reuse the shared gRPC cancellation prerequisite from 7219129 and cover late metadata settlement after a history transport failure. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: e0af01a5-0dfa-4e71-a660-c4186e65d7e0 Copilot-Session: 7b853e01-1e42-41c2-92a7-e2bbab87d215
Move the unchanged helper before the orchestration execution JSDoc so the independently based response-retry implementation overlaps at the same insertion point. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: e0af01a5-0dfa-4e71-a660-c4186e65d7e0 Copilot-Session: 7b853e01-1e42-41c2-92a7-e2bbab87d215
Preserve upstream client wait cancellation and restore the exact upstream gRPC helper. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: e0af01a5-0dfa-4e71-a660-c4186e65d7e0 Copilot-Session: 7b853e01-1e42-41c2-92a7-e2bbab87d215
Complete non-shutdown history failures as Failed without running partial history. Remove history-specific abandonment and its shared helper, restore the upstream version-rejection flow, and keep completion cancellation scoped to streamed work items. Update loopback failure and shutdown coverage without changing the existing service-negotiation E2E scenario or client-history cases. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: e0af01a5-0dfa-4e71-a660-c4186e65d7e0 Copilot-Session: 7b853e01-1e42-41c2-92a7-e2bbab87d215
Preserve dispatch-time retry signals and graceful initial responses for inline work. Deliver streamed-history failures and completions through the existing retry policy while retaining immediate history shutdown cancellation. Add focused loopback coverage for same-response retries and shutdown during initial delivery/backoff. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: e0af01a5-0dfa-4e71-a660-c4186e65d7e0 Copilot-Session: 7b853e01-1e42-41c2-92a7-e2bbab87d215
Use the dispatch-captured run signal for initial response RPCs, retries, and backoff. Remove the first-send exemption and streamed-versus-inline signal policy. Preserve history hydration and retry behavior. Update all response-type cancellation coverage, real gRPC initial-call cases, and shutdown documentation. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: e0af01a5-0dfa-4e71-a660-c4186e65d7e0 Copilot-Session: 7b853e01-1e42-41c2-92a7-e2bbab87d215
There was a problem hiding this comment.
🔵 Needs a closer look
It changes core replay hydration and shutdown behavior across every worker response path, warranting final human review despite comprehensive coverage.
Review details
- Files reviewed: 7/7 changed files
- Comments generated: 0 new
- Review effort level: Balanced
Summary
Why is this change needed?
The .NET worker supports service-selected history streaming; the JS worker did not. Client history queries are a different operation and do not supply the execution-specific replay history required by a worker. As explicitly requested in review, this PR also uses .NET-style cancellation for every worker response send, without a first-send exception.
HistoryStreaming; the service normally sends history inline with the work item.ExecutionStartedThis is a worker protocol capability, not a claim that all large histories previously failed, a fixed Azure size threshold, or reduced replay memory usage.
What changed?
HistoryStreaming.StreamInstanceHistorywith the original instance ID, execution ID, andForWorkItemProcessing=true.ExecutionStartedin past or new events, including a valid initial work item with empty past history.initialSignaland the streamed-versus-inline first-send policy. All response sends and retry backoff use one signal captured when work is dispatched; existing retry statuses, attempt limits, delays, and same-response delivery are unchanged.Reference: .NET BuildRuntimeStateAsync and its orchestrator failure handling.
Response cancellation reference: .NET completion RPC and retry loop.
Not included: a new recovery policy, a shared abandonment abstraction, changes to the global unary metadata helper or response retry policy, channel-lifetime changes, additional capabilities, or new public configuration.
Issues / work items
Narrowed in place to .NET worker-history parity, with uniform worker response cancellation explicitly requested in review. No companion PR is required.
Project checklist
CHANGELOG.mdAI-assisted code disclosure (required)
Was an AI tool used?
No
Yes, AI helped write parts of this PR
Yes, an AI agent generated most of this PR
Tool: GitHub Copilot.
AI-assisted areas: worker implementation, tests, release note, merge conflict resolution.
Changes after initial generation: removed the JS-specific abandon-on-history-failure policy and global unary metadata cancellation changes; kept the history-streaming capability and .NET failure semantics. Following review, removed the initial-send cancellation exception and dual-signal interface instead of retaining a JS compatibility policy.
AI verification:
Testing
Current head:
baf52c26f45762e34ec3036743ad88506767dd9f, based on maineab934075d939c9d86dead6b2a981df1eb7c66d1(response retries from #365 are already merged).4f34e10. Coverage includes already-aborted work, stop during initial metadata or an active RPC, every response type, and old work finishing after restart. Existing retry and streamed-history cases remain covered.7a41ddb, including service-selected multi-chunk history. Its Test and Build, DTS Emulator E2E, and Sample validation results are not claimed as validation of the new all-response shutdown policy.Real Azure DTS validation at
baf52c2(2026-09-11)Executed against the public Azure DTS production service in West US 2 using a newly isolated Consumption scheduler/task hub and Azure CLI identity, not an emulator or mocked backend. Production source and the existing history suite remained unchanged at the head above. The two additional response scenarios used an archived, uncommitted test-only overlay.
false,true,true. Reads used the same instance/execution andForWorkItemProcessing=true: first 1 chunk/3 events/1,573,273 bytes, then 2 chunks/6 events/2,359,827 bytes. The persisted result was Completed with{"indices":[0],"payloadBytes":786432}.{"recovered":true,"value":42}.INTERNALbefore remote dispatch, then forwarded the retry to Azure. Observed 2 SDK attempts, 1 actual forwarded request, identical response/token hashes, 1 activity execution, and persisted Completed with{"value":42}.The response scenarios passed 2/2, for 14/14 across these two suites. These tests do not claim an Azure-generated outage, a failure after server acceptance, or rollback of an already accepted completion. The cancellation scenario controls timing on the client but uses actual Azure acquisition, redelivery, and persistence.
An initial access-only run hit HTTP 403 because the runner's outbound NAT address rotates. After correcting only the isolated scheduler's narrow IP allowlist, authenticated preflight and the fresh runs above passed; the initial failures are not counted as feature validation. Final authenticated reads confirmed all 24 tracked instance IDs absent: 14 completed/purged roots or children and 10 earlier rejected scheduling attempts.
Cleanup completed: all workers/clients exited, the test-only overlay was archived and removed from the worktree, and the task-owned role assignment, task hub, scheduler, and resource group were deleted. Azure confirmed the resource group absent at
2026-09-11T20:39:38Z. No customer task hubs or instances were used.Notes for reviewers
The previous description's shared-helper integration instructions for PR #364/#365 no longer apply. Main now contains #365; this PR uses its existing response retry path and uniformly applies the captured cancellation signal from the first send. Hydration precedes JS's existing version-check/replay path; it does not claim identical tracing placement or the entire worker shutdown lifecycle to .NET.
The global unary helper remains identical to main. Metadata generation and user code are not interrupted by stop. If response metadata returns after stop, the aborted signal prevents RPC dispatch; waiting for that metadata can still outlast the bounded shutdown wait. Cancellation cannot undo a response already accepted by the backend.