Skip to content

Add chat-completions and output-budget knobs for OpenAI-compatible en… - #727

Closed
shpaw415 wants to merge 1 commit into
CopilotKit:mainfrom
shpaw415:zai-glm-compat
Closed

shpaw415 wants to merge 1 commit into
CopilotKit:mainfrom
shpaw415:zai-glm-compat

Conversation

@shpaw415

@shpaw415 shpaw415 commented Oct 3, 2026

Copy link
Copy Markdown

…dpoints

The built-in Bots resolve every model through the runtime's provider strings. On the OpenAI side that means the Responses API, which api.openai.com serves and a compatible endpoint that implements only /chat/completions answers with a 404 on every message. And a model the SDK does not recognise runs under a 4096-token output cap, which a thinking model can spend entirely on reasoning, ending the turn with no text and no tool calls.

OPENBOT_OPENAI_CHAT_COMPLETIONS=true builds the chat-completions client directly for the openai provider, with the same resolved key, base URL and model name the string form uses. OPENBOT_MAX_OUTPUT_TOKENS names an explicit output budget for the built-in Bots.

Both default to unset, which is the behaviour before this change.

What this changes

Where it runs

OpenBot is deployed as several server processes behind a load balancer, serving a whole company.
Consecutive requests from the same person reach different processes, and the process that answered a
WebSocket upgrade is rarely the one that answers the next call on that conversation.

State that outlives a single request therefore has to be shared, or the change works on one machine
and stops working the moment there are two, without saying so. That failure is worse than not
shipping the feature: it passes review, passes CI, passes a local demo, and only surfaces as a Bot
that forgets, a question nobody can answer, or a boundary that never fires.

Answer these even when the answer is "none":

  • New state that outlives a request? Where does it live? A Map or Set held in a module or
    a factory closure does not count as somewhere.
  • What happens on the second replica? Name the concrete outcome, not "should be fine".
  • Anything serialised? Say what stops two processes doing it at once. A unique index, a
    conditional update, or an advisory lock are answers. A check-then-write is not.
  • Anything fanned out to a browser? Say how it reaches a socket held by another process.
  • New listener, port, or schedule? Say how it is reached through the same ingress as the API,
    and what a hundred copies of it do.

Postgres is already there and is the default answer to all of the above: a table, a unique index, a
conditional update, LISTEN/NOTIFY for fan-out.

Boundary and audit

  • Every acting call still goes through the gateway: resolve, decide, audit, then act.
  • New refusals and new failures each write a row.
  • Nothing new is trusted from the client that the server can resolve itself.

Changelog

  • A line in CHANGELOG.md under Unreleased, or a sentence on why a deployment behaves no
    differently afterwards.

Proof

…dpoints

The built-in Bots resolve every model through the runtime's provider
strings. On the OpenAI side that means the Responses API, which
api.openai.com serves and a compatible endpoint that implements only
/chat/completions answers with a 404 on every message. And a model the
SDK does not recognise runs under a 4096-token output cap, which a
thinking model can spend entirely on reasoning, ending the turn with
no text and no tool calls.

OPENBOT_OPENAI_CHAT_COMPLETIONS=true builds the chat-completions
client directly for the openai provider, with the same resolved key,
base URL and model name the string form uses. OPENBOT_MAX_OUTPUT_TOKENS
names an explicit output budget for the built-in Bots.

Both default to unset, which is the behaviour before this change.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant