Skip to content

feat(worker): align history streaming with .NET - #363

Merged
wangbill (YunchuWang) merged 11 commits into
microsoft:mainfrom
YunchuWang:yunchuwang-worker-history-streaming
Sep 11, 2026
Merged

wangbill (YunchuWang) merged 11 commits into
microsoft:mainfrom
YunchuWang:yunchuwang-worker-history-streaming

Conversation

@YunchuWang

@YunchuWang wangbill (YunchuWang) commented Sep 8, 2026 •

Copy link
Copy Markdown
Member

Summary

Why is this change needed?

The .NET worker supports service-selected history streaming; the JS worker did not. Client history queries are a different operation and do not supply the execution-specific replay history required by a worker. As explicitly requested in review, this PR also uses .NET-style cancellation for every worker response send, without a first-send exception.

Customer experience Before After
A workflow accumulates history The worker does not advertise HistoryStreaming; the service normally sends history inline with the work item. The service may select streamed history. The worker fetches the complete history before replaying the same orchestrator code.
A non-shutdown streamed-history read fails or lacks ExecutionStarted Worker history streaming was not supported on main. Match .NET: report a Failed orchestration result rather than replay partial history or invent an abandon/redelivery recovery policy. If the backend accepts that result, the instance becomes Failed.
The worker stops during a history read No worker history stream existed to cancel. Cancel the history read and prevent late metadata from launching a completion RPC. An RPC already received by the backend cannot be guaranteed to be undone.
The worker stops before or during a response send Initial sends were exempt from the worker stop signal; retries and backoff were cancellable. Initial sends, retries, and backoff all use the same dispatch-captured signal, for inline/streamed orchestrations, activities, entities, and version-failure/rejection responses. Work finishing after stop cannot send a response, including after restart.

This is a worker protocol capability, not a claim that all large histories previously failed, a fixed Azure size threshold, or reduced replay memory usage.

What changed?

  • Advertise only HistoryStreaming.
  • For a service-selected streamed work item, call StreamInstanceHistory with the original instance ID, execution ID, and ForWorkItemProcessing=true.
  • Preserve event order, new events, and the completion token. Replace past events only after the stream finishes successfully.
  • Accept ExecutionStarted in past or new events, including a valid initial work item with empty past history.
  • Match .NET's non-shutdown hydration-failure outcome. Keep history-read cancellation and resource cleanup local to the new history stream.
  • Remove initialSignal and the streamed-versus-inline first-send policy. All response sends and retry backoff use one signal captured when work is dispatched; existing retry statuses, attempt limits, delays, and same-response delivery are unchanged.

Reference: .NET BuildRuntimeStateAsync and its orchestrator failure handling.

Response cancellation reference: .NET completion RPC and retry loop.

Not included: a new recovery policy, a shared abandonment abstraction, changes to the global unary metadata helper or response retry policy, channel-lifetime changes, additional capabilities, or new public configuration.

Issues / work items

Narrowed in place to .NET worker-history parity, with uniform worker response cancellation explicitly requested in review. No companion PR is required.


Project checklist

  • Release notes are not required for the next release
    • Otherwise: Notes added to CHANGELOG.md
  • Backport is not required
  • All required tests have been added/updated
  • Breaking change?
    • Intentional shutdown behavior change: initial response sends are now cancellable for all work items, not just retries. Already-running user code is not stopped, but results produced after stop are not sent. No public API or stored-history format changes.

AI-assisted code disclosure (required)

Was an AI tool used?

  • No

  • Yes, AI helped write parts of this PR

  • Yes, an AI agent generated most of this PR

  • Tool: GitHub Copilot.

  • AI-assisted areas: worker implementation, tests, release note, merge conflict resolution.

  • Changes after initial generation: removed the JS-specific abandon-on-history-failure policy and global unary metadata cancellation changes; kept the history-streaming capability and .NET failure semantics. Following review, removed the initial-send cancellation exception and dual-signal interface instead of retaining a JS compatibility policy.

AI verification:

  • I understand the code and can explain it (human confirmation pending)
  • I verified referenced APIs/types exist and are correct
  • I reviewed edge cases/failure paths
  • I reviewed concurrency/async behavior
  • I checked for unintended breaking or behavior changes

Testing

Current head: baf52c26f45762e34ec3036743ad88506767dd9f, based on main eab934075d939c9d86dead6b2a981df1eb7c66d1 (response retries from #365 are already merged).

  • Fresh targeted run: 128 tests across six suites passed: response delivery, real-gRPC response delivery, history streaming, worker versioning, startup, and stream recovery.
  • Initial-send regressions first failed against 4f34e10. Coverage includes already-aborted work, stop during initial metadata or an active RPC, every response type, and old work finishing after restart. Existing retry and streamed-history cases remain covered.
  • Real-loopback cases observe initial RPC cancellation and prevent late metadata from dispatching a first response.
  • Core build, scoped ESLint/Prettier, and normal commit hooks passed.
  • Historical backend evidence only: the full history suite passed 12/12 against an isolated DTS emulator at 7a41ddb, including service-selected multi-chunk history. Its Test and Build, DTS Emulator E2E, and Sample validation results are not claimed as validation of the new all-response shutdown policy.
  • Current-head CI is reported in the PR Checks tab.

Real Azure DTS validation at baf52c2 (2026-09-11)

Executed against the public Azure DTS production service in West US 2 using a newly isolated Consumption scheduler/task hub and Azure CLI identity, not an emulator or mocked backend. Production source and the existing history suite remained unchanged at the head above. The two additional response scenarios used an archived, uncommitted test-only overlay.

Scenario Result and observed evidence
Entire existing history E2E suite 12/12 passed. Covered worker streaming plus client history for activities, timers, sub-orchestrations, external events, ordering, and event sequences.
Service-selected worker history streaming Actual work-item flags were false,true,true. Reads used the same instance/execution and ForWorkItemProcessing=true: first 1 chunk/3 events/1,573,273 bytes, then 2 chunks/6 events/2,359,827 bytes. The persisted result was Completed with {"indices":[0],"payloadBytes":786432}.
Stop before first response; restart and real redelivery Held response metadata, stopped and restarted the same worker object, then released metadata. The old run made zero completion RPCs. A replacement worker received the same execution from Azure and persisted Completed with {"recovered":true,"value":42}.
Same-response retry and real persistence Injected one client-side INTERNAL before remote dispatch, then forwarded the retry to Azure. Observed 2 SDK attempts, 1 actual forwarded request, identical response/token hashes, 1 activity execution, and persisted Completed with {"value":42}.

The response scenarios passed 2/2, for 14/14 across these two suites. These tests do not claim an Azure-generated outage, a failure after server acceptance, or rollback of an already accepted completion. The cancellation scenario controls timing on the client but uses actual Azure acquisition, redelivery, and persistence.

An initial access-only run hit HTTP 403 because the runner's outbound NAT address rotates. After correcting only the isolated scheduler's narrow IP allowlist, authenticated preflight and the fresh runs above passed; the initial failures are not counted as feature validation. Final authenticated reads confirmed all 24 tracked instance IDs absent: 14 completed/purged roots or children and 10 earlier rejected scheduling attempts.

Cleanup completed: all workers/clients exited, the test-only overlay was archived and removed from the worktree, and the task-owned role assignment, task hub, scheduler, and resource group were deleted. Azure confirmed the resource group absent at 2026-09-11T20:39:38Z. No customer task hubs or instances were used.


Notes for reviewers

The previous description's shared-helper integration instructions for PR #364/#365 no longer apply. Main now contains #365; this PR uses its existing response retry path and uniformly applies the captured cancellation signal from the first send. Hydration precedes JS's existing version-check/replay path; it does not claim identical tracing placement or the entire worker shutdown lifecycle to .NET.

The global unary helper remains identical to main. Metadata generation and user code are not interrupted by stop. If response metadata returns after stop, the aborted signal prevents RPC dispatch; waiting for that metadata can still outlast the bounded shutdown wait. Cancellation cannot undo a response already accepted by the backend.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>

Copilot-Session: e0af01a5-0dfa-4e71-a660-c4186e65d7e0
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Copilot-Session: e0af01a5-0dfa-4e71-a660-c4186e65d7e0
@YunchuWang
wangbill (YunchuWang) marked this pull request as ready for review September 8, 2026 16:29
Copilot AI lite review requested due to automatic review settings September 8, 2026 16:29

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 Changes recommended

The new worker history streaming test leaks global OpenTelemetry state by disabling tracing instead of restoring the previous tracer provider, which can make other test suites flaky.

Once you've addressed the issues Copilot identified, you can request another Copilot review.

Pull request overview

This PR adds worker-side support for service-selected orchestration history streaming by hydrating full past history via StreamInstanceHistory before version checks, tracing, and replay, and it introduces new tests (local gRPC + opt-in Azure) to validate correctness and shutdown/transport-failure behavior.

Changes:

  • Advertises WORKER_CAPABILITY_HISTORY_STREAMING in GetWorkItems and hydrates streamed orchestration history before replay.
  • Adds shutdown/cleanup handling for in-flight history hydration and abandons work items on incomplete history transport to enable redelivery.
  • Adds comprehensive local gRPC tests plus an opt-in bounded Azure E2E test; updates changelog.
File summaries
File Description
packages/durabletask-js/src/worker/task-hub-grpc-worker.ts Implements streamed history hydration, cancellation tracking, capability advertisement, and abandon-on-transport-failure behavior.
packages/durabletask-js/test/worker-history-streaming.spec.ts Adds local ephemeral gRPC server tests covering streaming ordering, failures/redelivery, shutdown cancellation, and trace continuity.
test/e2e-azuremanaged/worker-history-streaming.spec.ts Adds opt-in Azure-managed negotiation test that passively observes real service-selected streaming and validates durable outcome replay.
CHANGELOG.md Documents the new worker history streaming/hydration behavior and its failure handling semantics.
Review details
  • Files reviewed: 4/4 changed files
  • Comments generated: 1
  • Review effort level: Lite

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Comment thread packages/durabletask-js/test/worker-history-streaming.spec.ts
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Copilot-Session: e0af01a5-0dfa-4e71-a660-c4186e65d7e0
@YunchuWang
wangbill (YunchuWang) marked this pull request as draft September 8, 2026 16:45
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Copilot-Session: e0af01a5-0dfa-4e71-a660-c4186e65d7e0
@YunchuWang
wangbill (YunchuWang) marked this pull request as ready for review September 8, 2026 17:11
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Copilot-Session: e0af01a5-0dfa-4e71-a660-c4186e65d7e0
Copilot-Session: 7b853e01-1e42-41c2-92a7-e2bbab87d215
@YunchuWang
wangbill (YunchuWang) marked this pull request as draft September 8, 2026 17:23
@YunchuWang
wangbill (YunchuWang) marked this pull request as ready for review September 8, 2026 17:33
Reuse the shared gRPC cancellation prerequisite from 7219129 and cover late metadata settlement after a history transport failure.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Copilot-Session: e0af01a5-0dfa-4e71-a660-c4186e65d7e0
Copilot-Session: 7b853e01-1e42-41c2-92a7-e2bbab87d215
@YunchuWang
wangbill (YunchuWang) marked this pull request as draft September 8, 2026 17:45
@YunchuWang
wangbill (YunchuWang) marked this pull request as ready for review September 8, 2026 17:53
Move the unchanged helper before the orchestration execution JSDoc so the independently based response-retry implementation overlaps at the same insertion point.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Copilot-Session: e0af01a5-0dfa-4e71-a660-c4186e65d7e0
Copilot-Session: 7b853e01-1e42-41c2-92a7-e2bbab87d215
@YunchuWang
wangbill (YunchuWang) marked this pull request as draft September 8, 2026 17:59
@YunchuWang
wangbill (YunchuWang) marked this pull request as ready for review September 8, 2026 18:06
Preserve upstream client wait cancellation and restore the exact upstream gRPC helper.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Copilot-Session: e0af01a5-0dfa-4e71-a660-c4186e65d7e0
Copilot-Session: 7b853e01-1e42-41c2-92a7-e2bbab87d215
Complete non-shutdown history failures as Failed without running partial history. Remove history-specific abandonment and its shared helper, restore the upstream version-rejection flow, and keep completion cancellation scoped to streamed work items.

Update loopback failure and shutdown coverage without changing the existing service-negotiation E2E scenario or client-history cases.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Copilot-Session: e0af01a5-0dfa-4e71-a660-c4186e65d7e0
Copilot-Session: 7b853e01-1e42-41c2-92a7-e2bbab87d215
@YunchuWang wangbill (YunchuWang) changed the title feat(worker): hydrate streamed orchestration history before replay feat(worker): align history streaming with .NET Sep 10, 2026
Preserve dispatch-time retry signals and graceful initial responses for inline work. Deliver streamed-history failures and completions through the existing retry policy while retaining immediate history shutdown cancellation.

Add focused loopback coverage for same-response retries and shutdown during initial delivery/backoff.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Copilot-Session: e0af01a5-0dfa-4e71-a660-c4186e65d7e0
Copilot-Session: 7b853e01-1e42-41c2-92a7-e2bbab87d215
Comment thread packages/durabletask-js/src/worker/task-hub-grpc-worker.ts
Comment thread packages/durabletask-js/src/worker/task-hub-grpc-worker.ts Outdated
Comment thread packages/durabletask-js/src/worker/task-hub-grpc-worker.ts
Use the dispatch-captured run signal for initial response RPCs, retries, and backoff. Remove the first-send exemption and streamed-versus-inline signal policy. Preserve history hydration and retry behavior.

Update all response-type cancellation coverage, real gRPC initial-call cases, and shutdown documentation.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Copilot-Session: e0af01a5-0dfa-4e71-a660-c4186e65d7e0
Copilot-Session: 7b853e01-1e42-41c2-92a7-e2bbab87d215
@YunchuWang
wangbill (YunchuWang) requested a balanced review from Copilot September 11, 2026 20:48

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🔵 Needs a closer look

It changes core replay hydration and shutdown behavior across every worker response path, warranting final human review despite comprehensive coverage.

Review details
  • Files reviewed: 7/7 changed files
  • Comments generated: 0 new
  • Review effort level: Balanced

@YunchuWang
wangbill (YunchuWang) merged commit 254d9c1 into microsoft:main Sep 11, 2026
26 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants