feat(ai): gate Chat Completions body fields by provider - #44710
Merged
Merged
Conversation
Derive store, stream usage, strict mode, and tool stream from provider/baseURL when explicit compatibility is not set. OpenAI keeps store:false and strict:false by default; DeepSeek/Moonshot/Together/ Nvidia/Cerebras/Chutes/ZAI/Cloudflare/Vercel omit store and use max_tokens. Add compatibility overrides for the gated fields.
Vercel AI Gateway is a pass-through proxy: docs list max_tokens only but upstream OpenAI reasoning models (gpt-5/o1/o3) require max_completion_tokens (vercel/ai#7863, #11828). Forcing max_tokens for provider vercel-ai-gateway breaks that. Leave to compatibility override / default max_completion_tokens instead.
rekram1-node
force-pushed
the
chat-compat-gates
branch
from
August 24, 2026 16:50
14fefaa to
ddda292
Compare
infinitetrooper
added a commit
to tessaryai/tessary
that referenced
this pull request
Sep 9, 2026
…pencode to 1.18.30 (#26) ## What broke Every Layer-2 triage run against `gpt-5.6-luna` was failing. The symptom was maximally misleading: ``` {"is_error":true,"error":"opencode produced no usable reply", "usage":{"input_tokens":0,"output_tokens":0,...}} ``` Zero tokens, valid key. That reads like a credential or connectivity problem, and it is neither. ## Why `@ai-sdk/openai-compatible` always sends `max_tokens`. OpenAI's GPT-5.x reasoning line rejects it: ``` Unsupported parameter: 'max_tokens' is not supported with this model. Use 'max_completion_tokens' instead. ``` OpenCode folds that 400 into an empty assistant turn, so a hard request-shape failure surfaced to us as an empty reply with no spend. `@ai-sdk/openai` rewrites the field per model ([`openai-chat-language-model.ts`](https://github.com/vercel/ai/blob/main/packages/openai/src/chat/openai-chat-language-model.ts)) and is the only one of the two that can: knowing which models need it means knowing the models, which is exactly what "any OpenAI-compatible endpoint" cannot assume. Upstream tracks this as [opencode#45223](anomalyco/opencode#45223), still open with no scheduled fix, so this does not wait on them. ## What changed **`fix`** — `openAiCompatProviderBlock` selects `@ai-sdk/openai` for the `OPENAI` mode, `@ai-sdk/openai-compatible` for everything else. One line plus its rationale, and the two test assertions that pinned the old value. **`chore`** — `opencode-ai` and `@opencode-ai/sdk` 1.18.13 → 1.18.30 across both recipes, `package.json`, and the lockfile. Independent of the fix; it stands on being current. ## For the reviewer **The bump does not fix the bug, and that was measured rather than assumed.** 1.18.30 reproduces the failure identically while on the compat adapter, because [opencode#44710](anomalyco/opencode#44710 provider gating does not recognise a custom `openai-compatible` block like `openai-direct`. Worth knowing before anyone concludes the bump was the cure. **Why only the OPENAI mode moved.** The other modes are not the same case. `api.x.ai/v1` and `generativelanguage.googleapis.com/v1beta/openai/` are deliberately maintained OpenAI-*compatible* surfaces with a stated compatibility contract. OpenAI's `/v1` is the origin of the shape and owes nobody compatibility, which is precisely how it drifted. Pointing the compat adapter at a compat endpoint remains correct. `@ai-sdk/xai` is a closer call than the others: it defaults to the same `api.x.ai/v1` base URL already in use and is a standalone implementation, so it is nearly a drop-in. It was left alone because there is no xAI key here to verify it with, and swapping a working provider blind inside an unrelated fix is how the next quiet zero-token failure gets introduced. Worth doing as its own change. Gemini is genuinely different: `@ai-sdk/google` uses a different endpoint path and an `x-goog-api-key` header instead of Bearer, so it is a credential-plumbing refactor, not a swap. **The rest of the codebase was audited for the same pattern and is clean.** The Java backend cannot hit this defect: `.maxTokens(4096)` appears exactly once in the whole tree ([`ChatModelFactory.java:515`](backend/llm-runtime/src/main/java/ai/tessary/llm/ChatModelFactory.java#L515)), on the Anthropic builder, because Anthropic's API requires it. No OpenAI-wire path sets a token cap at all. Its OpenAI path already uses OpenAI's own LangChain4j client. ## Verification - Reproduced and fixed against a real key inside the agent image, on both 1.18.13 and 1.18.30: compat adapter throws the 400, native package replies. - Confirmed the key, auth, egress and model access were never the problem (`GET /v1/models` → 200, model present). - Launcher suite 42/42. - **Not yet run end-to-end through a rebuilt launcher against a live triage job.** That is the remaining gap. 🤖 Generated with [Claude Code](https://claude.com/claude-code) --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
jinhuang712
pushed a commit
to jinhuang712/opencode
that referenced
this pull request
Sep 26, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Derive
store,stream_options.include_usage,strict,tool_stream, andmax_tokensfrom provider/baseURL when explicitcompatibilityis not set.store:falseonly for OpenAI-like providers (omit for DeepSeek/Moonshot/Together/Nvidia/Cerebras/Chutes/ZAI/Cloudflare/Vercel/XAI/Opencode/AntLing)supportsStrictMode=falsefor Moonshot/Together/Cloudflare AI Gateway/Nvidia (omitstrict)zaiToolStreamfor ZAI except glm-4.5 variantsmax_tokensfor non-standard providers,max_completion_tokensotherwiseLanguageModelCompatibilityoverrides for the gated fieldsFixes 2 recorded cassettes and 3 unit test expectations (
compile,tool-runtime,openai-chat,openai-compatible-chat).