Skip to content

fix(cloudflare): correct overstated context windows - #8314

Merged
rekram1-node merged 1 commit into
devfrom
workers-ai-context
Sep 28, 2026
Merged

rekram1-node merged 1 commit into
devfrom
workers-ai-context

Conversation

@rekram1-node

@rekram1-node rekram1-node commented Sep 28, 2026 •

Copy link
Copy Markdown
Collaborator

Some Cloudflare listings advertise a larger context window than the serving backend accepts. A client that fits its output to the listed window then sends requests the provider rejects, and because the requested output is the whole window, even a one-message session fails.

Workers AI

glm-5.3, glm-5.3-flash and deepseek-v4-flash-0731 are listed at 1,310,720 tokens by /ai/models/search (and the glm-5.3 / glm-5.3-flash model pages), but the backend rejects any request over 1,048,576:

Requested token count exceeds the model's maximum context length of 1048576 tokens.
You requested a total of 1048589 tokens: 13 tokens from the input messages and 1048576 tokens for the completion.

The deepseek-v4-flash-0731 model page already says 1,048,576.

#7454 capped these to 1,048,576, but the automated sync in #7640 restored 1,310,720 the same day, because the sync re-derives context from search on every run and only preserves limit.output. This PR corrects the three files and pins the served window in the sync. The pin only ever lowers a window, so it is a no-op if Cloudflare fixes search.

AI Gateway

anthropic/claude-sonnet-4.5 has been curated to 1,000,000 since the generator port (#4922). The gateway rejects prompts over 200,000 tokens (prompt is too long: 231024 tokens > 200000 maximum), and the model page says 200,000. This drops the curated override so the catalog value applies.

Verification

  • Live probes against Workers AI and the AI Gateway produced the errors quoted above.
  • bun test packages/core/test/cloudflare-workers-ai.test.ts passes, including two new tests; the pinning test fails without the change.
  • A live dry-run of the Workers AI sync now leaves glm-5.3 and glm-5.3-flash unchanged (it previously reverted both).
  • bun run cloudflare-ai-gateway:generate --check reports up to date.
  • bun validate passes.

Not included

  • The deepseek-v4-flash-0731 sync still shows a price update (Cloudflare lowered its price). That is unrelated and left for the bot.
  • Windows Cloudflare understates (for example glm-4.7-flash, gpt-5* on the gateway) are safe-side and left alone.

Some listings advertise a larger window than the serving backend accepts, so
a client that fits its output to the listed window still sends requests the
provider rejects.

Workers AI: glm-5.3, glm-5.3-flash and deepseek-v4-flash-0731 are listed at
1,310,720 tokens by /ai/models/search and the model pages, but the backend
rejects any request over 1,048,576 ("Requested token count exceeds the model's
maximum context length of 1048576 tokens"). The deepseek-v4-flash-0731 model
page already says 1,048,576. #7454 capped these, then the automated sync in
#7640 restored 1,310,720 because it re-derives context from search on every
run. Pin the served window in the sync (it only ever lowers a window) and
correct the three files.

AI Gateway: anthropic/claude-sonnet-4.5 was curated to 1,000,000 since the
generator port, but the gateway rejects prompts over 200,000 tokens
("prompt is too long: 231024 tokens > 200000 maximum") and the model page
says 200,000. Drop the curated override so the catalog value applies.
@github-actions

Copy link
Copy Markdown
Contributor

No actionable findings.

@github-actions github-actions Bot added the reviewer: ready Automated review found no actionable items label Sep 28, 2026
@rekram1-node
rekram1-node merged commit c36b8da into dev Sep 28, 2026
2 checks passed
@rekram1-node
rekram1-node deleted the workers-ai-context branch September 28, 2026 23:59
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

reviewer: ready Automated review found no actionable items

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant