Skip to content

fix(agent): stop one oversized model call context disabling a run's mirror - #3496

Merged
kwakayama merged 3 commits into
mainfrom
fix/oversize-run-event-truncate
Aug 9, 2026
Merged

kwakayama merged 3 commits into
mainfrom
fix/oversize-run-event-truncate

Conversation

@kwakayama

@kwakayama kwakayama commented Aug 9, 2026 •

Copy link
Copy Markdown
Contributor

Fixes veryfront/veryfront-issue-inbox#421.

Problem

assertSupportedEventSize threw before anything was written, and the sink's catch disposes the mirror on any error:

} catch (error) {
  input.mirror.dispose()
  throw error
}

The damage is what happens next. The mirror's own methods return early when disabled rather than throwing (run-chunk-mirror.ts:180-199), and the main streaming path goes through handleChunk. So after one oversized context the run kept executing and every later event was dropped in silence — no throw, no log, no metric — while the run could still report completed.

That is the "run finishes with a hole in its history and the UI waits forever" shape, with nothing anywhere recording that it happened.

What is deliberately unchanged

The gate stays. This sink is the persist-context-before-dispatch mechanism from #3309. Its oversize rejection is not only a size check — it is what stops a model being called when the call cannot be recorded. rejects append failures and a request over the general body limit asserts dispatches === 0, and that assertion is untouched and still passing.

An earlier draft of this change let the truncated context through and allowed the dispatch. That quietly traded a fail-closed audit guarantee for run completion, which is a bigger decision than this issue. Reverted.

What changes

Only what an operator is left holding after a refusal:

  • The attempt is on the record. The context is reduced to fit and persisted, stamped truncated: true with originalByteLength and omittedMessageCount. It cannot be mistaken for a faithful record — that is the whole point of the stamps, since a truncated model call context still parses as a valid one.
  • The mirror survives. The run's later events, including its failure, still persist. This is the actual defect in chore: bump version to 0.1.21 #421.
  • The refusal explains itself. The error names the real size, the limit, that the model was not dispatched, and what to reduce. A warn carries the same fields structured.

Reduction keeps the newest messages, which is what someone reading a failed run wants: text is clamped in decreasing passes, then the oldest messages are dropped, then a guaranteed-fit fallback keeps only the envelope.

Tests

11 passing. Updated two that encoded the old behaviour, and added:

  • the mirror stays usable and later events still persist after a refusal
  • the error is operator-actionable (states size, limit, non-dispatch, and the audit record)
  • a fit is still guaranteed when message count alone blows the budget and text clamping cannot help

The pre-existing dispatches === 0 assertion is kept verbatim and extended to also assert the mirror was not disposed.

deno task verify:quick exit 0.

Note on this issue's history

#421's description was written from too shallow a read and has been corrected three times on the issue: the abort path is narrower than I first claimed, truncation machinery already existed, and private run events are exempt from it by design. The defect that survived all of that is the one fixed here — the silent loss of everything after one oversized event.

Summary by CodeRabbit

  • Bug Fixes
    • Oversized run events are now saved in a truncated form instead of being discarded.
    • Truncation preserves key audit details, including original size and omitted-message information.
    • Clear errors are shown when model dispatch is refused after an oversized event is persisted.
    • The event mirror remains available after this refusal, while model dispatch stays suppressed.

…irror

An oversized context threw before anything was written, and the sink's catch
disposes the mirror on any error. Because the mirror's own handleChunk and
appendEvents return early when disabled rather than throwing, every later event
in that run was then dropped in silence — no throw, no log, no metric — while
the run carried on and could still report completed.

Keep the gate. The model must not be dispatched when its context cannot be
recorded faithfully, and the pre-existing dispatches === 0 contract still holds.
What changes is what an operator is left with:

- the context is reduced to fit and persisted, stamped truncated with
  originalByteLength and omittedMessageCount, so the attempt is on the record
  and can never be mistaken for a faithful one
- the mirror survives, so the run's later events, including its failure, persist
- the refusal names the actual size, the limit, that the model was not called,
  and what to reduce; a warn carries the same fields

Newest messages are kept: text is clamped in decreasing passes, then the oldest
messages are dropped, then a guaranteed-fit fallback keeps only the envelope.

Refs veryfront/veryfront-issue-inbox#421
@kwakayama
kwakayama requested a review from kojiwakayama as a code owner August 9, 2026 10:37
@coderabbitai

coderabbitai Bot commented Aug 9, 2026 •

Copy link
Copy Markdown

Review Change Stack

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Repository UI

Review profile: CHILL

Plan: Pro Plus

Run ID: 3abbe59c-21cc-478d-8dbd-1c04870a6036

📥 Commits

Reviewing files that changed from the base of the PR and between 3e65e69 and b7ae274.

📒 Files selected for processing (2)
  • src/agent/hosted/durable-run-event-sink.test.ts
  • src/agent/hosted/project-reference-resolver.test.ts
🚧 Files skipped from review as they are similar to previous changes (1)
  • src/agent/hosted/durable-run-event-sink.test.ts

📝 Walkthrough

Walkthrough

Oversized durable run events are reduced to fit the append limit, persisted with truncation metadata, and then refused for model dispatch. Tests verify persistence, error details, mirror usability, and fail-closed dispatch behavior. The resolver test narrows request header access.

Changes

Durable event handling

Layer / File(s) Summary
Oversized event reduction
src/agent/hosted/durable-run-event-sink.ts
The sink truncates message text by UTF-8 byte budget, removes older messages, and uses an omission fallback when required. It records original size and omitted-message count.
Persistence and dispatch refusal
src/agent/hosted/durable-run-event-sink.ts
The sink persists the resolved event, logs truncation details, and rejects model dispatch after persistence.
Oversized event validation
src/agent/hosted/durable-run-event-sink.test.ts
Tests verify append-budget compliance, production normalization, truncation metadata, refusal messages, mirror usability, and suppressed model dispatch.

Request mock typing

Layer / File(s) Summary
Request header narrowing
src/agent/hosted/project-reference-resolver.test.ts
The request mock narrows RequestInit before reading headers and continues to verify the authorization header.

Estimated code review effort: 4 (Complex) | ~45 minutes

Sequence Diagram(s)

sequenceDiagram
  participant DurableRunEventSink
  participant Append
  participant agentLogger
  participant ModelDispatch
  DurableRunEventSink->>Append: persist truncated audit record
  Append-->>DurableRunEventSink: persistence succeeds
  DurableRunEventSink->>agentLogger: log size and omission details
  DurableRunEventSink--x ModelDispatch: refuse oversized dispatch
Loading

Possibly related PRs

Suggested reviewers: kojiwakayama, copilot

🚥 Pre-merge checks | ✅ 5
✅ Passed checks (5 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title clearly and concisely describes the main fix: preventing one oversized model-call context from disabling the run mirror.
Docstring Coverage ✅ Passed No functions found in the changed files to evaluate docstring coverage. Skipping docstring coverage check.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
✨ Finishing Touches
📝 Generate docstrings
  • Create stacked PR
  • Commit on current branch
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch fix/oversize-run-event-truncate

Comment @coderabbitai help to get the list of available commands.

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: f0fd860ff3

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment on lines +111 to +113
truncated: true,
originalByteLength,
omittedMessageCount,

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Allow truncation metadata through private-event validation

Whenever a model-call context exceeds the append limit, these added fields make the audit event invalid in the production mirror path. ConversationRunChunkMirror.appendEvents passes it through prepareConversationRunExternalEvents, where isPrivateConversationRunEvent permits only type, messages, and tools; normalization therefore throws Invalid private run event shape before enqueueing anything, and this sink catches that error and disposes the mirror. The new tests miss the regression because their stub appendEvents bypasses normalization. Update the private-event schema and its consumers to accept the audit metadata, or encode the truncation without adding rejected fields.

Useful? React with 👍 / 👎.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Correct, and severe — fixed in 3e65e69.

private-run-event.ts:23 is a strict allowlist:

return Object.keys(value).every((key) => key === "type" || key === "messages" || key === "tools")

So the three fields I added would have failed isPrivateConversationRunEvent, normalizeConversationRunEvent would have thrown Invalid private run event shape inside appendEvents, this sink would have caught that as a persistence failure, and the mirror would have been disposed. The fix would have caused the exact loss it exists to prevent, on every oversized context.

Your point about the tests is the reason it got through: the stub appendEvents never normalizes, so nothing exercised the real path.

The marker now lives in a leading system message instead, so the event keeps exactly the three allowed keys. It still states the original size, the limit, how many messages were omitted, and that the record is an excerpt rather than the context that was sent.

Guarded two ways now: a test that pushes the persisted record through prepareConversationRunExternalEvents — the real production path — and direct isPrivateConversationRunEvent assertions. Verified load-bearing: adding a single stray top-level key fails three assertions.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 2

🧹 Nitpick comments (3)
src/agent/hosted/durable-run-event-sink.ts (2)

102-103: 🚀 Performance & Scalability | 🔵 Trivial | ⚡ Quick win

Fix the comment and drop the duplicate size measurement.

Two small points in this function:

  • Line 116 states that each pass halves the budget. Line 120 divides by four. Correct the comment.
  • Line 103 recomputes the append byte length that resolvePersistableEvent already computed at line 181. Each call serializes the whole oversized event, which is at least 10 MiB. Pass the known value in.
♻️ Proposed refactor
-function truncatePrivateRunEventToLimit(event: Record<string, unknown>): Record<string, unknown> {
-  const originalByteLength = getPrivateRunEventAppendRequestByteLength(event);
+function truncatePrivateRunEventToLimit(
+  event: Record<string, unknown>,
+  originalByteLength: number,
+): Record<string, unknown> {
   const messages = Array.isArray(event.messages) ? event.messages : [];
-  // Clamp text progressively; each pass halves the per-part budget.
+  // Clamp text progressively; each pass quarters the per-part budget.

Update the call site:

-  const truncated = truncatePrivateRunEventToLimit(event);
+  const truncated = truncatePrivateRunEventToLimit(event, requestByteLength);

Also applies to: 116-121

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@src/agent/hosted/durable-run-event-sink.ts` around lines 102 - 103, Update
truncatePrivateRunEventToLimit to accept the already computed append-request
byte length from resolvePersistableEvent, and use that value instead of calling
getPrivateRunEventAppendRequestByteLength again. Correct the truncation-loop
comment to describe division by four rather than halving the budget, while
preserving the existing truncation behavior.

79-89: 🚀 Performance & Scalability | 🔵 Trivial | 💤 Low value

Text clamping skips string content and non-text parts.

truncateMessageTextParts only reduces parts that carry a string text field inside an array content. Two common shapes escape clamping:

  • content as a plain string.
  • Non-text parts, such as base64 media or tool-result payloads.

Correctness still holds, because the message-removal loop and the envelope fallback guarantee a fit. The cost is wasted work: when the oversize comes from those shapes, all five clamping passes run and serialize the full event without reducing anything. Handling string content is a small addition that covers a frequent case.

♻️ Optional: clamp string `content`
 function truncateMessageTextParts(message: unknown, maxTextBytes: number): unknown {
-  if (!isRecord(message) || !Array.isArray(message.content)) return message;
+  if (!isRecord(message)) return message;
+  if (typeof message.content === "string") {
+    return { ...message, content: truncateTextToBytes(message.content, maxTextBytes) };
+  }
+  if (!Array.isArray(message.content)) return message;
   return {
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@src/agent/hosted/durable-run-event-sink.ts` around lines 79 - 89, Update
truncateMessageTextParts to also truncate message.content when it is a string,
using truncateTextToBytes with maxTextBytes. Preserve the existing
array-of-parts handling and leave non-text parts unchanged.
src/agent/hosted/durable-run-event-sink.test.ts (1)

210-224: 🚀 Performance & Scalability | 🔵 Trivial | 💤 Low value

This test is expensive to run.

The event holds 120,000 messages. The reducer maps all of them and serializes the full event once per clamping pass, five times, before message removal starts. That is several hundred megabytes of transient allocation and a large CPU cost for one unit test.

The assertion needs a message count that exceeds the budget on its own, so the count must stay large. You can still cut the cost by lowering the per-message text to a few characters and raising the count only as far as the budget requires. Measure the current runtime first, and keep the test as is if it stays fast.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@src/agent/hosted/durable-run-event-sink.test.ts` around lines 210 - 224,
Reduce the expensive fixture in the “guarantees a fit when message count alone
exceeds the budget” test while preserving its purpose: use only a few characters
per message and the minimum message count needed to exceed the persistence
budget. Measure the test runtime first and leave it unchanged if it remains
fast; otherwise adjust the existing Array.from message fixture without changing
the rejection assertion.
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@src/agent/hosted/durable-run-event-sink.ts`:
- Around line 152-160: Make the fallback in the event-clamping logic truly fit
the byte limit: after constructing the current fallback, measure it and, if
still oversized, return a minimal envelope containing only the required event
identity fields, replacement system message, and empty tools when applicable.
Add a test that makes a non-messages field such as providerOptions or metadata
oversized and verifies the final result is within
MAX_CONVERSATION_RUN_EVENT_APPEND_REQUEST_BYTES.
- Around line 106-114: Update the private-event validation used by
prepareConversationRunExternalEvents and isPrivateConversationRunEvent to accept
the truncation metadata fields truncated, originalByteLength, and
omittedMessageCount, while preserving validation of all existing event fields so
stamped audit records pass normalization and persist.

---

Nitpick comments:
In `@src/agent/hosted/durable-run-event-sink.test.ts`:
- Around line 210-224: Reduce the expensive fixture in the “guarantees a fit
when message count alone exceeds the budget” test while preserving its purpose:
use only a few characters per message and the minimum message count needed to
exceed the persistence budget. Measure the test runtime first and leave it
unchanged if it remains fast; otherwise adjust the existing Array.from message
fixture without changing the rejection assertion.

In `@src/agent/hosted/durable-run-event-sink.ts`:
- Around line 102-103: Update truncatePrivateRunEventToLimit to accept the
already computed append-request byte length from resolvePersistableEvent, and
use that value instead of calling getPrivateRunEventAppendRequestByteLength
again. Correct the truncation-loop comment to describe division by four rather
than halving the budget, while preserving the existing truncation behavior.
- Around line 79-89: Update truncateMessageTextParts to also truncate
message.content when it is a string, using truncateTextToBytes with
maxTextBytes. Preserve the existing array-of-parts handling and leave non-text
parts unchanged.
🪄 Autofix

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Repository UI

Review profile: CHILL

Plan: Pro Plus

Run ID: 6a85e4bb-caee-4112-a3ca-76fd7740bab0

📥 Commits

Reviewing files that changed from the base of the PR and between 9d7dde4 and f0fd860.

📒 Files selected for processing (2)
  • src/agent/hosted/durable-run-event-sink.test.ts
  • src/agent/hosted/durable-run-event-sink.ts

Comment thread src/agent/hosted/durable-run-event-sink.ts Outdated
Comment thread src/agent/hosted/durable-run-event-sink.ts Outdated
…ields

isPrivateConversationRunEvent allows only type, messages and tools
(private-run-event.ts:23). The truncated, originalByteLength and
omittedMessageCount fields the previous commit added made the audit record fail
that check, so normalizeConversationRunEvent would have thrown "Invalid private
run event shape" inside appendEvents, the sink would have caught it as a
persistence failure, and the mirror would have been disposed — reintroducing the
exact loss this change exists to prevent.

The stub appendEvents in the tests never normalizes, so nothing caught it.

Move the marker into a leading system message instead: it states the original
size, the limit, how many messages were omitted, and that the record is an
excerpt rather than the context that was sent. The event keeps exactly the three
allowed keys.

Also: the last-resort fallback now drops tools and asserts it fits rather than
spreading the original event and hoping, and the reused append-request byte
length is passed in instead of recomputed over a 10 MiB payload.

Guarded by a test that runs the persisted record through
prepareConversationRunExternalEvents — the real production path — and by direct
isPrivateConversationRunEvent assertions. Adding any stray top-level key fails
them.
Three problems the ci (lint) chain caught that local --no-check runs could not:

- appended[0][0] and the notice-message walk violated noUncheckedIndexedAccess;
  both go through small helpers that fail loudly instead of indexing blind
- the sink returns void | Promise<void>, so .then() is not available; the
  normalization test is async/await now
- assertRejects yields unknown, so error.message needed assertInstanceOf first

Importing the real private-event validator into this test widened the type graph
enough to surface a latent error in project-reference-resolver.test.ts, where
init is a union of RequestInit variants and .headers is not on every branch.
Narrowed with an `in` check rather than dropping the import: asserting against a
hand-copied allowlist instead of the real validator is what let the shape bug
through in the first place.

Full ci (lint) chain and verify:quick both clean.
@kwakayama
kwakayama added this pull request to the merge queue Aug 9, 2026
Merged via the queue into main with commit a46ead2 Aug 9, 2026
31 checks passed
@kwakayama
kwakayama deleted the fix/oversize-run-event-truncate branch August 9, 2026 11:30
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant