Skip to content

fix(agent): resolve undated model ids to their dated sibling's output cap - #4514

Merged
kojiwakayama merged 4 commits into
mainfrom
fix/model-output-cap-undated-ids
Sep 17, 2026
Merged

kojiwakayama merged 4 commits into
mainfrom
fix/model-output-cap-undated-ids

Conversation

@kojiwakayama

Copy link
Copy Markdown
Contributor

Why this is the demo blocker

MODEL_MAX_OUTPUT_TOKENS is keyed by exact model-id string. Any id not present falls through to FALLBACK_MODEL_MAX_OUTPUT_TOKENS = 4_096 — silently, with no warning and no telemetry.

The table carries anthropic/claude-haiku-4-5-20251001 at 64000 but not the undated anthropic/claude-haiku-4-5, so two spellings of one model differ 16×. The staging project agentic-email-processing-outlook has four agent definitions written with the undated id.

That silent 4096 cap is the head of the demo failure chain:

  1. This bug — undated id → 4096 instead of 64000.
  2. The cap truncates model output mid-tool_use.
  3. veryfront-issue-inbox#1467 — the Anthropic stream parser reads that truncation as a malformed stream rather than a legitimate max_tokens stop.
  4. The run surfaces the opaque Provider replay turn failed before its boundary — the incident in veryfront-issue-inbox#1475.

#1467 makes the failure diagnosable and #1475 hardens the terminal transition, but neither stops the demo failing. This does.

Fix

  • Undated ids resolve to their dated sibling's budget, derived automatically from the table rather than hand-maintained in a second list. Where several dated snapshots collapse to one undated id, the minimum wins — an undated spelling never promises more than the snapshot it may resolve to.
  • Lookup is case-insensitive and still strips the veryfront-cloud/ prefix.
  • An unknown id is now loud. It logs the model id and the applied limit, so a missing entry shows up in logs and traces instead of resurfacing later as a malformed provider stream. (Reported as max_output_limit because the log redactor masks any context key containing token.)

Guard against recurrence

The table's only previous protection was a MAINTENANCE: comment. Three tests now pin it against the in-repo VERYFRONT_CLOUD_CHAT_MODELS catalog, so a newly added model fails CI instead of silently getting 4096:

  • every catalog model has an explicit budget, not just the fallback
  • every dated catalog model gives its undated id the same budget
  • every thinking model clears a 16k floor — reasoning streams ahead of the answer, so a low cap truncates before any answer is produced

Audited the rest of the table, as asked

  • mistral/mistral-large-2512 sits at 1_024 against the api catalog's 128_000. Left alone deliberately — 4f109ba93 ("Fit Mistral Large chat within current quota") set it as an operational cap so hosted default-chat fits a 20k TPM deployment ceiling; raising it would reintroduce RateLimitReached. The thinking-model test documents the exemption so it is not "fixed" by mistake. Its standing directive is to replace the cap with registry-backed per-deployment quota metadata.
  • gpt-realtime, gpt-realtime-mini and gemini-embedding-001 are absent, but they are not chat models and do not reach this path.
  • gemini-3-flash-preview, gemini-3.1-flash-lite and kimi-k2 are in the table but not the chat catalog — harmless extra coverage.

Still needed separately

veryfront-api repeats the pattern independently: src/shared/ai/model-catalog.ts:557 lists aliases: ['haiku', 'anthropic/claude-haiku-4-5-20251001'] with the undated id absent, so getModelMaxEstimatedOutputTokens() falls back to DEFAULT_MAX_ESTIMATED_OUTPUT_TOKENS = 16_384. That needs its own PR.

Refs veryfront/veryfront-issue-inbox#1480

… cap

MODEL_MAX_OUTPUT_TOKENS is keyed by exact model id, so
`anthropic/claude-haiku-4-5` missed the table and silently took the
4_096 fallback while `anthropic/claude-haiku-4-5-20251001` took 64_000.
The truncated output then surfaces far away as a malformed provider
stream.

Lookups now normalize case and strip a trailing `-YYYYMMDD` snapshot
date, so a snapshot and its undated id share a ceiling. Two snapshots
that collapse to one id keep the lower ceiling, so an undated id can
never raise a snapshot's ceiling. This is not a family fallback: ids
without a snapshot date stay separate.

A model the table does not cover is now loud: it logs the model id and
the applied limit instead of returning 4_096 in silence.

Refs veryfront/veryfront-issue-inbox#1480

@greptile-apps greptile-apps Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Your trial has ended. Reactivate Greptile to resume code reviews.

@coderabbitai

coderabbitai Bot commented Sep 17, 2026 •

Copy link
Copy Markdown

Warning

Review limit reached

Next included review available in 33 minutes.

Check out review usage here.

View limit details

Limit details: You’ve used the included review currently available.

You've used all free OSS reviews for now. Wait for the free limit to reset to keep reviewing this public repository.

Learn how review limits work.

Review configuration:

⚙️ Run configuration

Configuration used: Repository UI

Review profile: CHILL

Plan: Advanced

Run ID: 0bc48ad4-cb98-427b-9492-629b9d9a1cf0

📥 Commits

Reviewing files that changed from the base of the PR and between 881c70e and 9e7ec07.

📒 Files selected for processing (2)
  • src/agent/runtime/constants.test.ts
  • src/agent/runtime/constants.ts

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@chatgpt-codex-connector

Copy link
Copy Markdown

You have reached your Codex usage limits for security reviews. Please try again later.

@github-actions

Copy link
Copy Markdown

📦 Client bundle boundary

Entrypoint Modules Source size Server leaks
src/index.client.ts 289 2312 KiB ✅ 0

A server module in a client graph aborts hydration in the browser. New leaks fail CI; known leaks are tracked in scripts/lint/client-bundle-baseline.json to burn down.

@gitar-bot

gitar-bot Bot commented Sep 17, 2026 •

Copy link
Copy Markdown

Gitar is working

Gitar

@kojiwakayama

Copy link
Copy Markdown
Contributor Author

@codex review

@chatgpt-codex-connector

Copy link
Copy Markdown

Codex Review: Didn't find any major issues. You're on a roll.

Reviewed commit: 7cb1267726

ℹ️ About Codex in GitHub

Codex has been enabled to automatically review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

When you sign up for Codex through ChatGPT, Codex can also answer questions or update the PR, like "@codex address that feedback".

@codecov

codecov Bot commented Sep 17, 2026

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.

📢 Thoughts on this report? Let us know!

Case-insensitivity was incomplete: the `veryfront-cloud/` prefix was stripped
case-sensitively BEFORE lowercasing, so `Veryfront-Cloud/anthropic/claude-haiku-4-5`
missed the table and took the 4096 fallback. Lowercase first, then strip.

The unknown-model warning fired on every agent step, not once per run. A project
on `local/llama-3.3-70b` logged up to DEFAULT_MAX_STEPS warnings per run forever.
It now warns once per distinct id (bounded set, so an attacker-supplied id cannot
grow it) and stays quiet for prefixes the cloud catalog is not meant to cover.

The warning test depended on the ambient LOG_LEVEL and failed under
`LOG_LEVEL=error`. It now pins the level and resets the dedup set, so it asserts
the same thing in any shell and in CI.
@chatgpt-codex-connector

Copy link
Copy Markdown

You have reached your Codex usage limits for security reviews. Please try again later.

@greptile-apps greptile-apps Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Your trial has ended. Reactivate Greptile to resume code reviews.

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 4875b07301

ℹ️ About Codex in GitHub

Codex has been enabled to automatically review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

When you sign up for Codex through ChatGPT, Codex can also answer questions or update the PR, like "@codex address that feedback".

Comment thread src/agent/runtime/constants.ts Outdated
@kojiwakayama

Copy link
Copy Markdown
Contributor Author

@codex review

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 4875b07301

ℹ️ About Codex in GitHub

Codex has been enabled to automatically review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

When you sign up for Codex through ChatGPT, Codex can also answer questions or update the PR, like "@codex address that feedback".

Comment thread src/agent/runtime/constants.ts Outdated
@kojiwakayama

Copy link
Copy Markdown
Contributor Author

@codex review

@chatgpt-codex-connector

Copy link
Copy Markdown

Codex Review: Didn't find any major issues. Chef's kiss.

Reviewed commit: 4875b07301

ℹ️ About Codex in GitHub

Codex has been enabled to automatically review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

When you sign up for Codex through ChatGPT, Codex can also answer questions or update the PR, like "@codex address that feedback".

…d memory

Two Codex P2s on veryfront-code#4514, both real.

The dedup set stopped INSERTING at 256 ids but still returned true, so every id
past the cap warned on every call -- precisely the log flood the cap existed to
prevent. It now stops warning once the cap is reached.

A model id is caller-supplied and can be as large as the request body allows, so
an entry-count cap did not bound memory: 256 ids could retain ~256 MiB. The
retained key is now truncated to 200 characters; the warning still reports the
id the caller sent.

Both are pinned by tests. The cap test fails against the previous shape.
@chatgpt-codex-connector

Copy link
Copy Markdown

You have reached your Codex usage limits for security reviews. Please try again later.

@greptile-apps greptile-apps Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Your trial has ended. Reactivate Greptile to resume code reviews.

@kojiwakayama

Copy link
Copy Markdown
Contributor Author

@codex review

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 73eb54b59b

ℹ️ About Codex in GitHub

Codex has been enabled to automatically review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

When you sign up for Codex through ChatGPT, Codex can also answer questions or update the PR, like "@codex address that feedback".

Comment thread src/agent/runtime/constants.ts Outdated

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 73eb54b59b

ℹ️ About Codex in GitHub

Codex has been enabled to automatically review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

When you sign up for Codex through ChatGPT, Codex can also answer questions or update the PR, like "@codex address that feedback".

Comment thread src/agent/runtime/constants.ts
Codex P2 follow-up on veryfront-code#4514. My previous fix bounded only the
retained deduplication key, so the warning still emitted the caller-controlled
model id in full: 256 distinct ids could still produce ~256 MiB of logs, and
arbitrary untrusted request content reached log subscribers through the model
context key.

AGENTS.md forbids raw request payloads in logs and asks for a redaction marker,
so an oversized id is now cut at the same bound with its dropped length named.

My earlier test asserted the full value was logged -- it pinned the defect. It
now asserts the emitted value is bounded and carries the marker, and fails if
the raw id is logged again.
@kojiwakayama

Copy link
Copy Markdown
Contributor Author

@codex review

@greptile-apps greptile-apps Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Your trial has ended. Reactivate Greptile to resume code reviews.

@chatgpt-codex-connector

Copy link
Copy Markdown

You have reached your Codex usage limits for security reviews. Please try again later.

@chatgpt-codex-connector

Copy link
Copy Markdown

Codex Review: Didn't find any major issues. Nice work!

Reviewed commit: 9e7ec07187

ℹ️ About Codex in GitHub

Codex has been enabled to automatically review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

When you sign up for Codex through ChatGPT, Codex can also answer questions or update the PR, like "@codex address that feedback".

@sonarqubecloud

Copy link
Copy Markdown

@kojiwakayama
kojiwakayama added this pull request to the merge queue Sep 17, 2026
@github-merge-queue
github-merge-queue Bot removed this pull request from the merge queue due to failed status checks Sep 17, 2026
@kojiwakayama
kojiwakayama added this pull request to the merge queue Sep 17, 2026
Merged via the queue into main with commit 53b775c Sep 17, 2026
59 checks passed
@kojiwakayama
kojiwakayama deleted the fix/model-output-cap-undated-ids branch September 17, 2026 16:43
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant