Skip to content

fix(world-postgres): ignore stream rows written after the first EOF - #3712

Merged
VaguelySerious merged 2 commits into
vercel:mainfrom
himself65:fix/world-postgres-read-after-eof
Sep 14, 2026
Merged

VaguelySerious merged 2 commits into
vercel:mainfrom
himself65:fix/world-postgres-read-after-eof

Conversation

@himself65

Copy link
Copy Markdown
Contributor

Description

streams.get() in @workflow/world-postgres closes the controller on the EOF row during its catch-up loop and keeps iterating. If the stream has rows after its first EOF — a producer that retried its terminal write after a lost ACK or an overlapping attempt, so the frame was appended and the stream closed a second time — the next row hits controller.enqueue() on a closed controller. Node throws ERR_INVALID_STATE (Invalid state: Controller is already closed) out of start(), the stream errors, and every chunk still queued is discarded: the reader sees an empty, errored stream instead of the data written before the EOF.

We hit this in production: a long-running agent run pushed its output through writeToStream/closeStream; an ingest retry duplicated the terminal frame (data ×3, EOF ×3, data ×2, EOF ×2 in workflow_stream_chunks), and the workflow's finalize step — which reads the stream from 0 — persisted the run as failed with empty output even though all the data was there. On our instance 143 streams had rows after their EOF and 1809 had more than one EOF row. Reproduces on 4.3.3, 4.3.4 and main.

The fix tracks the first EOF and ignores every later row (data or EOF). It does not change behaviour for well-formed streams.

How did you test your changes?

  • New unit test packages/world-postgres/src/streamer.test.ts drives createStreamer against a fake pool/drizzle (the pg Client is stubbed so no LISTEN socket is opened). It writes five data rows, an EOF, then a duplicate data row + EOF, and expects the stream to drain exactly the five rows. Without the fix it rejects with TypeError: Invalid state: Controller is already closed.
  • Several data rows before the EOF are deliberate: the catch-up loop is synchronous, so all but the first sit in the controller's queue when the EOF closes it, and it is that non-empty queue that the post-EOF enqueue blows away. With a single queued chunk the consumer drains it first by microtask order and the failure does not reproduce.
  • tsc --noEmit and biome check on the package are clean (Biome's pre-existing noExcessiveCognitiveComplexity warning on enqueue goes from 16 to 20 with the extra guard; getChunks already warns at 21).

PR Checklist - Required to merge

  • 📦 pnpm changeset was run to create a changelog for this PR (@workflow/world-postgres: patch)
  • 🔒 DCO sign-off passes (run git commit --signoff on your commits)
  • 📝 Ping @vercel/workflow in a comment once the PR is ready, and the above checklist is complete

@himself65
himself65 requested a review from a team as a code owner August 21, 2026 18:28
@changeset-bot

changeset-bot Bot commented Aug 21, 2026 •

Copy link
Copy Markdown

🦋 Changeset detected

Latest commit: 434b214

The changes in this PR will be included in the next version bump.

This PR includes changesets to release 1 package
Name Type
@workflow/world-postgres Patch

Not sure what this means? Click here to learn what changesets are.

Click here if you're a maintainer who wants to add another changeset to this PR

@vercel

vercel Bot commented Aug 21, 2026

Copy link
Copy Markdown
Contributor

@himself65 is attempting to deploy a commit to the Vercel Labs Team on Vercel.

A member of the Team first needs to authorize it.

@karthikscale3 karthikscale3 left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

See inline comment.

Comment thread packages/world-postgres/src/streamer.ts Outdated
controller.enqueue(new Uint8Array(msg.data));
}
if (msg.eof) {
closed = true;

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for submitting a fix for this. One edge case remains: because the offset > 0 branch runs before this EOF handling, a positive startIndex can consume the first EOF as if it were a data chunk. With five data rows + EOF + retried data/EOF, get(..., 6) skips the first EOF and returns the post-EOF duplicate (or hangs if no later EOF arrives). Could we handle EOF before decrementing offset—or only apply the offset branch when !msg.eof—and add a regression test for a start index beyond the valid data count?

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Good catch, thanks. Fixed in d70e0e8 by exempting EOF rows from the offset branch (if (offset > 0 && !msg.eof)) — offsets count data chunks, matching getInfo's tailIndex, so the marker is never consumed. Added a regression test for a start index at (5) and past (6) the data count with a post-EOF duplicate and no trailing EOF; the 6 case hung until timeout before the fix.

@karthikscale3 karthikscale3 left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

See inline comment.

// index at or past the data count must still close the
// stream rather than consume the marker and then hang, or
// surface rows written after it.
if (offset > 0 && !msg.eof) {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

One remaining edge case: a negative startIndex is calculated below from chunks.length, which includes rows after the first EOF. For A…E, EOF, duplicate-E, get(..., -1) computes an offset from seven rows, then reaches the first EOF without returning the final valid chunk. Could we derive dataCount from the first EOF (for example, with findIndex) and add a startIndex = -1 regression test?

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks — fixed in 4319027: dataCount now comes from chunks.findIndex((c) => c.eof) (falling back to chunks.length with no EOF), so a negative start index resolves against the same rows enqueue delivers. Regression test added for -1 (→ e) and -2 (→ d, e) on A…E + EOF + duplicate-E; -1 returned nothing before the fix.

@himself65
himself65 force-pushed the fix/world-postgres-read-after-eof branch from 4319027 to 6a05ac4 Compare August 22, 2026 15:45
@VaguelySerious

Copy link
Copy Markdown
Member

@himself65 in order for us to merge this, commits must have verified signatures. Can you squash and re-push a signed commit?

@VaguelySerious
VaguelySerious force-pushed the fix/world-postgres-read-after-eof branch from 8d31d56 to 3e35119 Compare September 14, 2026 21:21
`streams.get()`'s catch-up loop closed the controller on the first EOF
row and kept iterating. A stream with rows after its first EOF — a
producer that retried its terminal write after a lost ACK or an
overlapping attempt, so the frame was appended and the stream closed
again — then hit `controller.enqueue()` on a closed controller. Node
threw ERR_INVALID_STATE, the stream errored, and every chunk still
queued was discarded: a reader saw an empty, errored stream instead of
the data written before the EOF.

Track the first EOF and ignore everything written after it:
- `enqueue()` now ignores rows once the stream has reached a terminal
  state, via the same reader-lifecycle `cleanedUp` flag main already
  uses to detach a completed reader's listener.
- The EOF marker itself is exempt from the start-index offset, so a
  start index at or past the data count still closes the stream
  instead of consuming the marker and hanging.
- A negative start index resolves against the rows before the first
  EOF, not every row in the table.
- `getChunks()` and `getInfo()` read the same table and had the same
  gap: both counted and returned rows written after the first EOF.
  They're now bounded by it too, via a shared `findFirstEofChunkId`
  lookup.

Co-authored-by: Peter Wielander <peter.wielander@vercel.com>
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Signed-off-by: Alex Yang <himself65@outlook.com>
Signed-off-by: Peter Wielander <peter.wielander@vercel.com>
@VaguelySerious
VaguelySerious force-pushed the fix/world-postgres-read-after-eof branch from 3e35119 to 209aa44 Compare September 14, 2026 21:45

@VaguelySerious VaguelySerious left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

AI review: no blocking issues

AI Review: Note

Out of scope for this PR (different package, not in this diff), but packages/world-local/src/streamer.ts around line 643 has the same class of bug this PR fixes for world-postgres: its negative-startIndex resolution only excludes a single trailing EOF chunk file, not every row written after the first EOF. Nothing in world-local's write/close prevents writing after close, so a retried terminal write there would inflate dataChunkCount and resolve a negative startIndex against the wrong position. Worth a follow-up issue if that producer-retry scenario is possible for world-local consumers too.

// that retries a terminal write can append data and EOF rows after it;
// `getChunks`/`getInfo` bound their queries by this so those rows are
// ignored the same way `streams.get()` ignores them.
const findFirstEofChunkId = async (

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

AI Review: Note

getChunks() and getInfo() read the same workflow_stream_chunks table as streams.get() and had the same gap: both counted/returned rows written after the first EOF (a retried terminal write never errors these two, it just leaks a duplicate chunk into a page and inflates tailIndex by one per duplicate). Verified against the original PR revision with a real-Postgres integration test (getChunks() returned ["a","b","c","d","e","e"] instead of 5 chunks, getInfo().tailIndex was 5 instead of 4) before this fix. Added findFirstEofChunkId and bounded both queries by it, plus test/stream-post-eof-duplicates.test.ts covering both against a real Postgres container.

Comment thread .changeset/world-postgres-read-after-eof.md Outdated
Signed-off-by: Peter Wielander <mittgfu@gmail.com>
@VaguelySerious
VaguelySerious merged commit 5f723b3 into vercel:main Sep 14, 2026
5 of 23 checks passed
github-actions Bot added a commit that referenced this pull request Sep 14, 2026
…3712)

Signed-off-by: Alex Yang <himself65@outlook.com>
Signed-off-by: Peter Wielander <peter.wielander@vercel.com>
Co-authored-by: Peter Wielander <peter.wielander@vercel.com>
Co-authored-by: Peter Wielander <mittgfu@gmail.com>
Signed-off-by: Alex Yang <himself65@outlook.com>
@github-actions

Copy link
Copy Markdown
Contributor

Backport PR opened against stable: #4171. Merge conflicts were resolved by AI — please review carefully. (backport job run)

VaguelySerious added a commit that referenced this pull request Sep 14, 2026
…3712) (#4171)

Signed-off-by: Alex Yang <himself65@outlook.com>
Signed-off-by: Peter Wielander <peter.wielander@vercel.com>
Co-authored-by: Peter Wielander <peter.wielander@vercel.com>
Co-authored-by: Peter Wielander <mittgfu@gmail.com>
Co-authored-by: Alex Yang <himself65@outlook.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants