Summary
@workflow/world-local can treat duplicate processing of the same hook_created operation as a real token conflict. The resulting event log contains hook_created followed by hook_conflict for the exact same hook correlation id, token, and run id. On replay, the hook consumer sees the hook_conflict and rejects with HookConflictError, even though no other workflow or hook owned the token.
This looks related to #1665, but the failure below is specifically about idempotency for the same hook creation in the local world storage path.
Versions observed
@workflow/core@5.0.0-beta.12
@workflow/world-local@5.0.0-beta.13
Event log smoking gun
A single run produced this sequence around the failing turn:
hook_created hook_...PTJ token wrun_...:turn-completion:2
hook_conflict hook_...PTJ token wrun_...:turn-completion:2 conflictingRunId wrun_...
The important part is that both events use the same:
correlationId / hook id
- hook token
runId
So this is not two different workflows racing for one token. It is duplicate processing of the same hook creation being recorded as a conflict with itself.
Why this happens
In @workflow/world-local/dist/storage/events-storage.js, the hook_created branch atomically claims the hook token with an exclusive token file:
const tokenClaimed = await writeExclusive(constraintPath, JSON.stringify({
token: hookData.token,
hookId: data.correlationId,
runId: effectiveRunId,
}));
But when that exclusive write fails, the existing token-claim read path does not preserve enough identity to decide whether the existing claim is the same hook or a different hook. In particular, the persisted claim includes hookId, but the schema used when reading the claim drops it.
That makes duplicate/replayed/concurrent processing of the same hook creation non-idempotent.
There is an existing contrast in the step path: step_created has explicit duplicate-correlation protection and throws EntityConflictError, which the runtime treats as a benign concurrent replay outcome. The hook path needs the same same-entity idempotency behavior.
Suggested patch shape
Make same-hook token claims idempotent in @workflow/world-local.
The token claim schema should preserve hookId when reading an existing claim:
const HookTokenClaimSchema = z.object({
hookId: z.string().optional(),
runId: z.string(),
});
Then, when writeExclusive(constraintPath, ...) fails in the hook_created branch:
const existingClaim = await readHookTokenClaim(constraintPath);
if (
existingClaim?.runId === effectiveRunId &&
existingClaim.hookId === data.correlationId
) {
throw new EntityConflictError(`Hook "${data.correlationId}" already created`);
}
// Otherwise this is a real token conflict with another hook/run.
// Keep writing hook_conflict here.
This preserves real conflict behavior while making duplicate same-hook creation idempotent.
Note on core behavior
An earlier version of this report suggested changing @workflow/core so a handler that sees duplicate hook creation returns no pending steps. I no longer think that is the right suggested fix shape.
Current core appears to use a broader recovery pattern: handleSuspension returns all pending steps, inline execution is gated by createdStepCorrelationIds, and the runtime queues pending steps with idempotencyKey: step.correlationId. That may be intentional crash recovery for a handler that created step_created but died before queueing the step.
So this issue should focus on the concrete world-local storage bug: same (runId, hookId, token) must not become hook_conflict. Any question about whether core's pending-step enqueue behavior needs stronger guarantees should be handled separately with careful analysis of queue idempotency and crash-recovery semantics.
Expected behavior
Duplicate processing of the same hook creation should be idempotent:
- same
runId + same hookId + same token: no hook_conflict
- different
runId or different hookId with the same token: still write hook_conflict
Regression tests that would cover this
-
world-local storage test:
- create a
hook_created event
- concurrently or repeatedly create the same
hook_created event with the same runId, correlationId, and token
- assert no
hook_conflict is appended
-
world-local negative test:
- create a hook with token
T
- create a different hook id or run id with token
T
- assert
hook_conflict is still appended
Summary
@workflow/world-localcan treat duplicate processing of the samehook_createdoperation as a real token conflict. The resulting event log containshook_createdfollowed byhook_conflictfor the exact same hook correlation id, token, and run id. On replay, the hook consumer sees thehook_conflictand rejects withHookConflictError, even though no other workflow or hook owned the token.This looks related to #1665, but the failure below is specifically about idempotency for the same hook creation in the local world storage path.
Versions observed
@workflow/core@5.0.0-beta.12@workflow/world-local@5.0.0-beta.13Event log smoking gun
A single run produced this sequence around the failing turn:
The important part is that both events use the same:
correlationId/ hook idrunIdSo this is not two different workflows racing for one token. It is duplicate processing of the same hook creation being recorded as a conflict with itself.
Why this happens
In
@workflow/world-local/dist/storage/events-storage.js, thehook_createdbranch atomically claims the hook token with an exclusive token file:But when that exclusive write fails, the existing token-claim read path does not preserve enough identity to decide whether the existing claim is the same hook or a different hook. In particular, the persisted claim includes
hookId, but the schema used when reading the claim drops it.That makes duplicate/replayed/concurrent processing of the same hook creation non-idempotent.
There is an existing contrast in the step path:
step_createdhas explicit duplicate-correlation protection and throwsEntityConflictError, which the runtime treats as a benign concurrent replay outcome. The hook path needs the same same-entity idempotency behavior.Suggested patch shape
Make same-hook token claims idempotent in
@workflow/world-local.The token claim schema should preserve
hookIdwhen reading an existing claim:Then, when
writeExclusive(constraintPath, ...)fails in thehook_createdbranch:This preserves real conflict behavior while making duplicate same-hook creation idempotent.
Note on core behavior
An earlier version of this report suggested changing
@workflow/coreso a handler that sees duplicate hook creation returns no pending steps. I no longer think that is the right suggested fix shape.Current core appears to use a broader recovery pattern:
handleSuspensionreturns all pending steps, inline execution is gated bycreatedStepCorrelationIds, and the runtime queues pending steps withidempotencyKey: step.correlationId. That may be intentional crash recovery for a handler that createdstep_createdbut died before queueing the step.So this issue should focus on the concrete
world-localstorage bug: same(runId, hookId, token)must not becomehook_conflict. Any question about whether core's pending-step enqueue behavior needs stronger guarantees should be handled separately with careful analysis of queue idempotency and crash-recovery semantics.Expected behavior
Duplicate processing of the same hook creation should be idempotent:
runId+ samehookId+ same token: nohook_conflictrunIdor differenthookIdwith the same token: still writehook_conflictRegression tests that would cover this
world-localstorage test:hook_createdeventhook_createdevent with the samerunId,correlationId, and tokenhook_conflictis appendedworld-localnegative test:TThook_conflictis still appended