Skip to content

feat(ai): gate Chat Completions body fields by provider - #44710

Merged
rekram1-node merged 5 commits into
v2from
chat-compat-gates
Aug 24, 2026
Merged

rekram1-node merged 5 commits into
v2from
chat-compat-gates

Conversation

@rekram1-node

Copy link
Copy Markdown
Collaborator

Derive store, stream_options.include_usage, strict, tool_stream, and max_tokens from provider/baseURL when explicit compatibility is not set.

  • store:false only for OpenAI-like providers (omit for DeepSeek/Moonshot/Together/Nvidia/Cerebras/Chutes/ZAI/Cloudflare/Vercel/XAI/Opencode/AntLing)
  • supportsStrictMode=false for Moonshot/Together/Cloudflare AI Gateway/Nvidia (omit strict)
  • zaiToolStream for ZAI except glm-4.5 variants
  • max_tokens for non-standard providers, max_completion_tokens otherwise
  • Add LanguageModelCompatibility overrides for the gated fields

Fixes 2 recorded cassettes and 3 unit test expectations (compile, tool-runtime, openai-chat, openai-compatible-chat).

@rekram1-node
rekram1-node changed the base branch from dev to v2 August 24, 2026 14:54
Derive store, stream usage, strict mode, and tool stream from
provider/baseURL when explicit compatibility is not set. OpenAI keeps
store:false and strict:false by default; DeepSeek/Moonshot/Together/
Nvidia/Cerebras/Chutes/ZAI/Cloudflare/Vercel omit store and use
max_tokens. Add compatibility overrides for the gated fields.
Vercel AI Gateway is a pass-through proxy: docs list max_tokens only
but upstream OpenAI reasoning models (gpt-5/o1/o3) require
max_completion_tokens (vercel/ai#7863, #11828). Forcing max_tokens for
provider vercel-ai-gateway breaks that. Leave to compatibility override
/ default max_completion_tokens instead.
@rekram1-node
rekram1-node merged commit 6a68739 into v2 Aug 24, 2026
6 of 7 checks passed
@rekram1-node
rekram1-node deleted the chat-compat-gates branch August 24, 2026 17:40
infinitetrooper added a commit to tessaryai/tessary that referenced this pull request Sep 9, 2026
…pencode to 1.18.30 (#26)

## What broke

Every Layer-2 triage run against `gpt-5.6-luna` was failing. The symptom
was maximally misleading:

```
{"is_error":true,"error":"opencode produced no usable reply",
 "usage":{"input_tokens":0,"output_tokens":0,...}}
```

Zero tokens, valid key. That reads like a credential or connectivity
problem, and it is neither.

## Why

`@ai-sdk/openai-compatible` always sends `max_tokens`. OpenAI's GPT-5.x
reasoning line rejects it:

```
Unsupported parameter: 'max_tokens' is not supported with this model.
Use 'max_completion_tokens' instead.
```

OpenCode folds that 400 into an empty assistant turn, so a hard
request-shape failure surfaced to us as an empty reply with no spend.

`@ai-sdk/openai` rewrites the field per model
([`openai-chat-language-model.ts`](https://github.com/vercel/ai/blob/main/packages/openai/src/chat/openai-chat-language-model.ts))
and is the only one of the two that can: knowing which models need it
means knowing the models, which is exactly what "any OpenAI-compatible
endpoint" cannot assume. Upstream tracks this as
[opencode#45223](anomalyco/opencode#45223),
still open with no scheduled fix, so this does not wait on them.

## What changed

**`fix`** — `openAiCompatProviderBlock` selects `@ai-sdk/openai` for the
`OPENAI` mode, `@ai-sdk/openai-compatible` for everything else. One line
plus its rationale, and the two test assertions that pinned the old
value.

**`chore`** — `opencode-ai` and `@opencode-ai/sdk` 1.18.13 → 1.18.30
across both recipes, `package.json`, and the lockfile. Independent of
the fix; it stands on being current.

## For the reviewer

**The bump does not fix the bug, and that was measured rather than
assumed.** 1.18.30 reproduces the failure identically while on the
compat adapter, because
[opencode#44710](anomalyco/opencode#44710
provider gating does not recognise a custom `openai-compatible` block
like `openai-direct`. Worth knowing before anyone concludes the bump was
the cure.

**Why only the OPENAI mode moved.** The other modes are not the same
case. `api.x.ai/v1` and
`generativelanguage.googleapis.com/v1beta/openai/` are deliberately
maintained OpenAI-*compatible* surfaces with a stated compatibility
contract. OpenAI's `/v1` is the origin of the shape and owes nobody
compatibility, which is precisely how it drifted. Pointing the compat
adapter at a compat endpoint remains correct.

`@ai-sdk/xai` is a closer call than the others: it defaults to the same
`api.x.ai/v1` base URL already in use and is a standalone
implementation, so it is nearly a drop-in. It was left alone because
there is no xAI key here to verify it with, and swapping a working
provider blind inside an unrelated fix is how the next quiet zero-token
failure gets introduced. Worth doing as its own change. Gemini is
genuinely different: `@ai-sdk/google` uses a different endpoint path and
an `x-goog-api-key` header instead of Bearer, so it is a
credential-plumbing refactor, not a swap.

**The rest of the codebase was audited for the same pattern and is
clean.** The Java backend cannot hit this defect: `.maxTokens(4096)`
appears exactly once in the whole tree
([`ChatModelFactory.java:515`](backend/llm-runtime/src/main/java/ai/tessary/llm/ChatModelFactory.java#L515)),
on the Anthropic builder, because Anthropic's API requires it. No
OpenAI-wire path sets a token cap at all. Its OpenAI path already uses
OpenAI's own LangChain4j client.

## Verification

- Reproduced and fixed against a real key inside the agent image, on
both 1.18.13 and 1.18.30: compat adapter throws the 400, native package
replies.
- Confirmed the key, auth, egress and model access were never the
problem (`GET /v1/models` → 200, model present).
- Launcher suite 42/42.
- **Not yet run end-to-end through a rebuilt launcher against a live
triage job.** That is the remaining gap.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
jinhuang712 pushed a commit to jinhuang712/opencode that referenced this pull request Sep 26, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant