fix(cloudflare): correct overstated context windows - #8314
Merged
Merged
Conversation
Some listings advertise a larger window than the serving backend accepts, so
a client that fits its output to the listed window still sends requests the
provider rejects.
Workers AI: glm-5.3, glm-5.3-flash and deepseek-v4-flash-0731 are listed at
1,310,720 tokens by /ai/models/search and the model pages, but the backend
rejects any request over 1,048,576 ("Requested token count exceeds the model's
maximum context length of 1048576 tokens"). The deepseek-v4-flash-0731 model
page already says 1,048,576. #7454 capped these, then the automated sync in
#7640 restored 1,310,720 because it re-derives context from search on every
run. Pin the served window in the sync (it only ever lowers a window) and
correct the three files.
AI Gateway: anthropic/claude-sonnet-4.5 was curated to 1,000,000 since the
generator port, but the gateway rejects prompts over 200,000 tokens
("prompt is too long: 231024 tokens > 200000 maximum") and the model page
says 200,000. Drop the curated override so the catalog value applies.
Contributor
|
No actionable findings. |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Some Cloudflare listings advertise a larger context window than the serving backend accepts. A client that fits its output to the listed window then sends requests the provider rejects, and because the requested output is the whole window, even a one-message session fails.
Workers AI
glm-5.3,glm-5.3-flashanddeepseek-v4-flash-0731are listed at 1,310,720 tokens by/ai/models/search(and theglm-5.3/glm-5.3-flashmodel pages), but the backend rejects any request over 1,048,576:The
deepseek-v4-flash-0731model page already says 1,048,576.#7454 capped these to 1,048,576, but the automated sync in #7640 restored 1,310,720 the same day, because the sync re-derives
contextfrom search on every run and only preserveslimit.output. This PR corrects the three files and pins the served window in the sync. The pin only ever lowers a window, so it is a no-op if Cloudflare fixes search.AI Gateway
anthropic/claude-sonnet-4.5has been curated to 1,000,000 since the generator port (#4922). The gateway rejects prompts over 200,000 tokens (prompt is too long: 231024 tokens > 200000 maximum), and the model page says 200,000. This drops the curated override so the catalog value applies.Verification
bun test packages/core/test/cloudflare-workers-ai.test.tspasses, including two new tests; the pinning test fails without the change.glm-5.3andglm-5.3-flashunchanged (it previously reverted both).bun run cloudflare-ai-gateway:generate --checkreports up to date.bun validatepasses.Not included
deepseek-v4-flash-0731sync still shows a price update (Cloudflare lowered its price). That is unrelated and left for the bot.glm-4.7-flash,gpt-5*on the gateway) are safe-side and left alone.