Skip to content

fix: exit the database-maintenance worker test harness on a detached-spawn error (#9059) - #9063

Merged
atomantic merged 1 commit into
mainfrom
claim/issue-9059
Sep 28, 2026
Merged

atomantic merged 1 commit into
mainfrom
claim/issue-9059

Conversation

@atomantic

Copy link
Copy Markdown
Owner

What

Fixes a real hang in server/lib/databaseMaintenanceWorker.test.js's launch() test helper that most likely explains the "worker exit not observed within 30s on Windows CI" flake.

Root cause

launch()'s inline child script attaches both an 'error' and a 'close' handler to the spawnDatabaseMaintenanceWorker() handle, plus a released poll interval used to detect the release-admission outcome. That interval only clears once the operation's control directory disappears — which is not the case for a refusal (the control dir stays under the fence).

The 'close' handler calls process.exit() to force the harness process to exit even while that interval is still running. The 'error' handler did not — it only set process.exitCode, leaving the still-running interval to keep the process alive indefinitely. If spawnDatabaseMaintenanceWorker's detached-spawn machinery ever emits 'error' instead of 'close' (e.g. the Windows two-hop PowerShell supervisor missing its PID deadline under CI load), the harness process never exits on its own. The only thing that eventually reaps it is the outer run() spawn's own timeout: 30_000 (spawn(..., { timeout: 30_000 })), which SIGTERMs the hung process — producing exactly the observed exitCode: null around 30s instead of a real status.

Fix

child.on('error', …) now also calls process.exit(), matching the 'close' handler, so a detached-spawn error surfaces immediately as a real exit code/stderr instead of stalling the harness for the full outer timeout.

Test plan

  • server/node_modules/.bin/vitest run lib/databaseMaintenanceWorker.test.js — 8/8 passing locally (macOS).
  • npm run pregate — green (40 files / 594 tests).
  • Local codex review (provider:codex, low effort) — no findings.

This was not reproduced on Windows CI directly (the original failure is a suspected flake), but the fix removes a concrete, demonstrable hang that the existing code left in place for any 'error' outcome, and does not raise any timeout.

Closes #9059

…or (#9059)

launch()'s inline child script only set process.exitCode on the handle's
'error' event, unlike its 'close' handler which also calls process.exit().
The reusable 'released' poll interval never clears while the operation's
control dir still exists (true for every non-release outcome, including a
refusal), so an 'error' instead of 'close' left the harness process running
forever with no application-level timeout — only the outer run() spawn's own
30s timeout eventually reaped it via SIGTERM, surfacing as exitCode: null
instead of a real status. That matches the Windows CI symptom: the two-hop
PowerShell supervisor chain spawnDatabaseMaintenanceWorker drives on Windows
can miss its PID deadline under load and emit 'error', which this harness
bug then stretched into a spurious 30s timeout instead of a fast, readable
failure.
@atomantic
atomantic merged commit 1302f8e into main Sep 28, 2026
9 checks passed
@atomantic
atomantic deleted the claim/issue-9059 branch September 28, 2026 08:56
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

databaseMaintenanceWorker.test.js: worker exit not observed within 30s on Windows CI (suspected flake)

1 participant