fix(agent): resolve undated model ids to their dated sibling's output cap - #4514
Conversation
… cap MODEL_MAX_OUTPUT_TOKENS is keyed by exact model id, so `anthropic/claude-haiku-4-5` missed the table and silently took the 4_096 fallback while `anthropic/claude-haiku-4-5-20251001` took 64_000. The truncated output then surfaces far away as a malformed provider stream. Lookups now normalize case and strip a trailing `-YYYYMMDD` snapshot date, so a snapshot and its undated id share a ceiling. Two snapshots that collapse to one id keep the lower ceiling, so an undated id can never raise a snapshot's ceiling. This is not a family fallback: ids without a snapshot date stay separate. A model the table does not cover is now loud: it logs the model id and the applied limit instead of returning 4_096 in silence. Refs veryfront/veryfront-issue-inbox#1480
There was a problem hiding this comment.
Your trial has ended. Reactivate Greptile to resume code reviews.
|
Warning Review limit reachedNext included review available in 33 minutes. View limit detailsLimit details: You’ve used the included review currently available. You've used all free OSS reviews for now. Wait for the free limit to reset to keep reviewing this public repository. Review configuration: ⚙️ Run configurationConfiguration used: Repository UI Review profile: CHILL Plan: Advanced Run ID: 📒 Files selected for processing (2)
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
|
You have reached your Codex usage limits for security reviews. Please try again later. |
📦 Client bundle boundary
A server module in a client graph aborts hydration in the browser. New leaks fail CI; known leaks are tracked in |
|
@codex review |
|
Codex Review: Didn't find any major issues. You're on a roll. Reviewed commit: ℹ️ About Codex in GitHubCodex has been enabled to automatically review pull requests in this repo. Reviews are triggered when you
If Codex has suggestions, it will comment; otherwise it will react with 👍. When you sign up for Codex through ChatGPT, Codex can also answer questions or update the PR, like "@codex address that feedback". |
Codecov Report✅ All modified and coverable lines are covered by tests. 📢 Thoughts on this report? Let us know! |
Case-insensitivity was incomplete: the `veryfront-cloud/` prefix was stripped case-sensitively BEFORE lowercasing, so `Veryfront-Cloud/anthropic/claude-haiku-4-5` missed the table and took the 4096 fallback. Lowercase first, then strip. The unknown-model warning fired on every agent step, not once per run. A project on `local/llama-3.3-70b` logged up to DEFAULT_MAX_STEPS warnings per run forever. It now warns once per distinct id (bounded set, so an attacker-supplied id cannot grow it) and stays quiet for prefixes the cloud catalog is not meant to cover. The warning test depended on the ambient LOG_LEVEL and failed under `LOG_LEVEL=error`. It now pins the level and resets the dedup set, so it asserts the same thing in any shell and in CI.
|
You have reached your Codex usage limits for security reviews. Please try again later. |
There was a problem hiding this comment.
Your trial has ended. Reactivate Greptile to resume code reviews.
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 4875b07301
ℹ️ About Codex in GitHub
Codex has been enabled to automatically review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
When you sign up for Codex through ChatGPT, Codex can also answer questions or update the PR, like "@codex address that feedback".
|
@codex review |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 4875b07301
ℹ️ About Codex in GitHub
Codex has been enabled to automatically review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
When you sign up for Codex through ChatGPT, Codex can also answer questions or update the PR, like "@codex address that feedback".
|
@codex review |
|
Codex Review: Didn't find any major issues. Chef's kiss. Reviewed commit: ℹ️ About Codex in GitHubCodex has been enabled to automatically review pull requests in this repo. Reviews are triggered when you
If Codex has suggestions, it will comment; otherwise it will react with 👍. When you sign up for Codex through ChatGPT, Codex can also answer questions or update the PR, like "@codex address that feedback". |
…d memory Two Codex P2s on veryfront-code#4514, both real. The dedup set stopped INSERTING at 256 ids but still returned true, so every id past the cap warned on every call -- precisely the log flood the cap existed to prevent. It now stops warning once the cap is reached. A model id is caller-supplied and can be as large as the request body allows, so an entry-count cap did not bound memory: 256 ids could retain ~256 MiB. The retained key is now truncated to 200 characters; the warning still reports the id the caller sent. Both are pinned by tests. The cap test fails against the previous shape.
|
You have reached your Codex usage limits for security reviews. Please try again later. |
There was a problem hiding this comment.
Your trial has ended. Reactivate Greptile to resume code reviews.
|
@codex review |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 73eb54b59b
ℹ️ About Codex in GitHub
Codex has been enabled to automatically review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
When you sign up for Codex through ChatGPT, Codex can also answer questions or update the PR, like "@codex address that feedback".
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 73eb54b59b
ℹ️ About Codex in GitHub
Codex has been enabled to automatically review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
When you sign up for Codex through ChatGPT, Codex can also answer questions or update the PR, like "@codex address that feedback".
Codex P2 follow-up on veryfront-code#4514. My previous fix bounded only the retained deduplication key, so the warning still emitted the caller-controlled model id in full: 256 distinct ids could still produce ~256 MiB of logs, and arbitrary untrusted request content reached log subscribers through the model context key. AGENTS.md forbids raw request payloads in logs and asks for a redaction marker, so an oversized id is now cut at the same bound with its dropped length named. My earlier test asserted the full value was logged -- it pinned the defect. It now asserts the emitted value is bounded and carries the marker, and fails if the raw id is logged again.
|
@codex review |
There was a problem hiding this comment.
Your trial has ended. Reactivate Greptile to resume code reviews.
|
You have reached your Codex usage limits for security reviews. Please try again later. |
|
Codex Review: Didn't find any major issues. Nice work! Reviewed commit: ℹ️ About Codex in GitHubCodex has been enabled to automatically review pull requests in this repo. Reviews are triggered when you
If Codex has suggestions, it will comment; otherwise it will react with 👍. When you sign up for Codex through ChatGPT, Codex can also answer questions or update the PR, like "@codex address that feedback". |
|



Why this is the demo blocker
MODEL_MAX_OUTPUT_TOKENSis keyed by exact model-id string. Any id not present falls through toFALLBACK_MODEL_MAX_OUTPUT_TOKENS = 4_096— silently, with no warning and no telemetry.The table carries
anthropic/claude-haiku-4-5-20251001at 64000 but not the undatedanthropic/claude-haiku-4-5, so two spellings of one model differ 16×. The staging projectagentic-email-processing-outlookhas four agent definitions written with the undated id.That silent 4096 cap is the head of the demo failure chain:
tool_use.max_tokensstop.Provider replay turn failed before its boundary— the incident in veryfront-issue-inbox#1475.#1467 makes the failure diagnosable and #1475 hardens the terminal transition, but neither stops the demo failing. This does.
Fix
veryfront-cloud/prefix.max_output_limitbecause the log redactor masks any context key containingtoken.)Guard against recurrence
The table's only previous protection was a
MAINTENANCE:comment. Three tests now pin it against the in-repoVERYFRONT_CLOUD_CHAT_MODELScatalog, so a newly added model fails CI instead of silently getting 4096:Audited the rest of the table, as asked
mistral/mistral-large-2512sits at1_024against the api catalog's128_000. Left alone deliberately —4f109ba93("Fit Mistral Large chat within current quota") set it as an operational cap so hosted default-chat fits a 20k TPM deployment ceiling; raising it would reintroduceRateLimitReached. The thinking-model test documents the exemption so it is not "fixed" by mistake. Its standing directive is to replace the cap with registry-backed per-deployment quota metadata.gpt-realtime,gpt-realtime-miniandgemini-embedding-001are absent, but they are not chat models and do not reach this path.gemini-3-flash-preview,gemini-3.1-flash-liteandkimi-k2are in the table but not the chat catalog — harmless extra coverage.Still needed separately
veryfront-api repeats the pattern independently:
src/shared/ai/model-catalog.ts:557listsaliases: ['haiku', 'anthropic/claude-haiku-4-5-20251001']with the undated id absent, sogetModelMaxEstimatedOutputTokens()falls back toDEFAULT_MAX_ESTIMATED_OUTPUT_TOKENS = 16_384. That needs its own PR.Refs veryfront/veryfront-issue-inbox#1480