Skip to content

fix(openai): bound tool-call arguments by bytes, not provider chunking - #3521

Merged
kojiwakayama merged 1 commit into
mainfrom
fix/openai-tool-arg-fragment-budget
Aug 9, 2026
Merged

kojiwakayama merged 1 commit into
mainfrom
fix/openai-tool-arg-fragment-budget

Conversation

@kojiwakayama

Copy link
Copy Markdown
Contributor

Fixes veryfront-issue-inbox#432.

Symptom

Every scheduled tick of agentic-job-submission-processing in production produces a failed child run:

terminal_error_code:    agent-provider-error
terminal_error_message: veryfront-cloud request failed: invalid successful stream
                        (tool call arguments exceeded 4096 fragments)

The parent run reports completed while the child fails, so it is invisible from run status — the schedule looks healthy from the top.

Root cause

appendOpenAIStreamToolArgument counted every fragment against MAX_OPENAI_STREAM_TOOL_ARGUMENT_FRAGMENTS, including the ones carrying bytes:

if (budget.fragments >= MAX_OPENAI_STREAM_TOOL_ARGUMENT_FRAGMENTS) return "fragments";

The two budgets are wildly out of proportion. At typical streaming delta sizes of 5–20 UTF-8 bytes, 4096 fragments is roughly 20–80 KB — about 2–8% of the 1 MiB that MAX_OPENAI_STREAM_TOOL_ARGUMENT_BYTES nominally allows. The fragment cap therefore always fired first, leaving the byte limit effectively unreachable.

Fragment count is also the wrong thing to hold a caller to: it is chosen by the provider's tokenizer, not by the size or complexity of the arguments the agent produced. The same logical tool call passed or failed depending on how OpenAI happened to chunk it.

Fix

I did not guess at the cap's intent — the code's own tests state it. Both provider streams exercise it by flooding arguments: "" / delta: "". Zero-byte fragments never advance the byte budget, so they are the only ones that can arrive without bound, and they are what the cap is for.

So only zero-byte fragments count against it. Content is bounded by bytes, and since every kept fragment is at least one byte, the chunk array cannot outgrow the byte budget either.

Behaviour preserved

Both existing flood tests still trip the cap, unchanged — they flood empty fragments. The guard's real purpose is intact; only the false positive on ordinary traffic is gone.

Tests

Adds direct unit tests for the budget, which had none — it was covered only indirectly through the two stream parsers, which is precisely why a cap that rejected ordinary traffic went unnoticed.

The new test fails before this change with:

AssertionError: Values are not equal: expected no limit for 16384 small fragments

Also pinned: empty-fragment floods still rejected, the byte budget still enforced, and multi-byte characters counted by UTF-8 length rather than code units.

Verification

10/10 tests pass in extensions/ext-llm-openai. lint, fmt, deno check clean.

The 5 failures in src/provider on a full run are pre-existing and unrelated — I stashed this change and reproduced them identically on a clean tree (model-registry, veryfront-cloud/provider; they look network-dependent).

Follow-up worth separate triage

A child run can fail while its parent reports completed. That is why this ran hourly for as long as it did without surfacing.

Every scheduled run of at least one production project has been failing
hourly with:

  veryfront-cloud request failed: invalid successful stream
  (tool call arguments exceeded 4096 fragments)

`appendOpenAIStreamToolArgument` counted every fragment against
`MAX_OPENAI_STREAM_TOOL_ARGUMENT_FRAGMENTS`, including the ones that
carried bytes. At typical streaming delta sizes of 5-20 UTF-8 bytes, 4096
fragments is roughly 20-80 KB -- about 2-8% of the 1 MiB that
`MAX_OPENAI_STREAM_TOOL_ARGUMENT_BYTES` nominally allows. The fragment cap
therefore always fired first and the byte budget was unreachable.

Fragment count is also the wrong thing to hold a caller to: it is chosen
by the provider's tokenizer, not by the size or complexity of the
arguments the agent produced, so the same tool call passed or failed
depending on how the stream happened to be chunked.

Both provider streams' existing tests show what the cap is actually for --
each floods `arguments: ""` / `delta: ""`. Zero-byte fragments never
advance the byte budget, so they are the only ones that can arrive without
bound. Count only those against the cap. Content is then bounded by bytes,
and since every kept fragment is at least one byte, the chunk array cannot
outgrow the byte budget either.

Behaviour preserved: both existing flood tests still trip the cap
unchanged, because they flood empty fragments.

Adds direct unit tests for the budget, which had none -- it was only
covered indirectly through the two stream parsers. The new test fails
before this change with "expected no limit for 16384 small fragments".

Fixes veryfront-issue-inbox#432. Note the parent run reports `completed`
while the child fails, so this was invisible from run status; that
reporting gap is worth a separate look.
@kojiwakayama
kojiwakayama requested a review from kwakayama as a code owner August 9, 2026 22:03
@coderabbitai

coderabbitai Bot commented Aug 9, 2026

Copy link
Copy Markdown

Warning

Review limit reached

@kojiwakayama, you've reached your PR review limit, so we couldn't start this review.

Next review available in: 4 minutes

Enable usage-based reviews in Billing to review now. Otherwise, wait until the next included review is available.
You're only billed for reviews past your plan's rate limits ($0.25/file).

How can I continue?

After more reviews become available, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

To avoid repeated limits, reduce automatic review volume by pausing incremental auto-reviews earlier, using label-based review opt-in, excluding WIP or generated PR titles, or requesting reviews manually when the PR is ready. If your team needs uninterrupted high-volume reviews, an organization admin can enable usage-based reviews.

How do review limits work?

CodeRabbit enforces per-developer PR review limits for each organization. Most developers receive the normal plan review availability.

For paid Pro and Pro+ PR reviews, CodeRabbit uses adaptive limits for sustained high-volume activity. When a developer's recent PR review activity reaches the 95th percentile or higher among CodeRabbit users, additional reviews become available more gradually as earlier reviews age out of the rolling window.

Please refer docs for additional details.

Review details
⚙️ Run configuration

Configuration used: Repository UI

Review profile: CHILL

Plan: Pro Plus

Run ID: d71b7f1b-4212-4f97-bffb-968b7fa349ac

📥 Commits

Reviewing files that changed from the base of the PR and between 59d159c and c7481c7.

📒 Files selected for processing (2)
  • extensions/ext-llm-openai/src/openai-tool-input.test.ts
  • extensions/ext-llm-openai/src/openai-tool-input.ts

Comment @coderabbitai help to get the list of available commands.

@kojiwakayama
kojiwakayama enabled auto-merge August 9, 2026 22:03

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: c7481c7464

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

// unreachable. Fragment count is chosen by the provider's tokenizer rather
// than by the caller, so the same tool call passed or failed depending on how
// the stream happened to be chunked.
if (fragmentBytes === 0) {

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Bound retained nonempty argument chunks

When an OpenAI-compatible provider emits very small nonempty deltas, this condition bypasses the fragment cap entirely, so both stream parsers can retain up to 1,048,576 separately allocated strings before the byte budget fires. The payload is limited to 1 MiB, but the array and per-string overhead can consume tens of megabytes per concurrent request, whereas the previous implementation retained at most 4,096 chunks. Allow provider-controlled chunking without retaining every fragment separately, for example by coalescing chunks while maintaining the byte limit.

Useful? React with 👍 / 👎.

@kojiwakayama
kojiwakayama added this pull request to the merge queue Aug 9, 2026
Merged via the queue into main with commit bfd16a1 Aug 9, 2026
31 checks passed
@kojiwakayama
kojiwakayama deleted the fix/openai-tool-arg-fragment-budget branch August 9, 2026 22:22
@kojiwakayama kojiwakayama mentioned this pull request Aug 9, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant