Skip to content
Closed
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
13 changes: 13 additions & 0 deletions .env.example
Original file line number Diff line number Diff line change
Expand Up @@ -194,6 +194,19 @@ OPENAI_API_KEY=
#
# OPENAI_BASE_URL=

# Serve the built-in Bots' model calls over `/chat/completions` instead of the OpenAI Responses
# API. Needed for a compatible endpoint that implements only chat completions — Z.AI's coding
# endpoint is one — because every `openai/<model>` model string otherwise goes to `/responses`
# and comes back 404. The key, base URL and model name are the same ones the string form uses.
# OPENBOT_OPENAI_CHAT_COMPLETIONS=true

# An explicit output-token budget for the built-in Bots. Models the AI SDK does not recognise
# run in a compatibility mode capped at 4096 output tokens, which a thinking model can spend
# entirely on reasoning: the turn then ends with no text and no tool calls. Set a budget the
# model fits in (32768 is comfortable for GLM-5.3-Flash with thinking). Unset keeps the SDK's
# own default for models it knows.
# OPENBOT_MAX_OUTPUT_TOKENS=32768

# Optional dictation, independent of every agent/chat provider. Set provider, base URL,
# and model together. The endpoint must implement POST /audio/transcriptions under this
# base URL (include /v1 if needed). No chat endpoint or API key is inherited.
Expand Down
47 changes: 8 additions & 39 deletions bun.lock

Some generated files are not rendered by default. Learn more about how customized files appear on GitHub.

2 changes: 2 additions & 0 deletions docs/configuration.md
Original file line number Diff line number Diff line change
Expand Up @@ -56,6 +56,8 @@ at `agent-langgraph` on a laptop.
| `BOT_MODEL` | the Bot's row in the spec file | Model for whichever Bot is starting. Unset, it comes from that Bot's row in [the provider spec file](#the-provider-spec-file); the provider fallbacks are `gpt-5.5`, `claude-sonnet-4-5`, and `gemini-2.5-flash`. |
| `AGENT_BOT_MODEL` | `gpt-5.5` | Model for the proof-of-concept Bot (`agent-bot`), kept separate because it speaks `/v1/chat/completions` directly and refuses a model it cannot use. |
| `BOT_RESPONSES_API` | `false` | Makes `agent-langgraph` use the OpenAI Responses API. |
| `OPENBOT_OPENAI_CHAT_COMPLETIONS` | unset | Serves the built-in Bots over `/chat/completions` instead of the Responses API, for endpoints that implement only the older one. |
| `OPENBOT_MAX_OUTPUT_TOKENS` | unset | Explicit output-token budget for the built-in Bots; lifts the 4096-token compatibility cap on models the SDK does not know. |
| `BOT_REASONING_EFFORT` | unset (provider default) | OpenAI and the Responses API only: one of `none`, `minimal`, `low`, `medium`, `high`, `xhigh`, `max`. `agent-langgraph` refuses to start on any other value, on a non-`openai` provider, or without the Responses API. |
| `AGENT_STALL_TIMEOUT_MS` | unset (off) | How long a Bot's stream may produce nothing before the turn is ended for it. |
| `AGENT_TOOL_TOKEN` | unset; `start.sh` generates one | The secret a framework Bot presents when it calls a granted tool back through this server. |
Expand Down
1 change: 1 addition & 0 deletions server/package.json
Original file line number Diff line number Diff line change
Expand Up @@ -20,6 +20,7 @@
"@better-auth/scim": "1.7.1",
"@better-auth/sso": "^1.7.1",
"@composio/core": "^0.18.1",
"@ai-sdk/openai": "3.0.99",
"@copilotkit/runtime": "1.73.1",
"@mastra/client-js": "^1.43.0",
"@mastra/core": "^1.64.0",
Expand Down
59 changes: 58 additions & 1 deletion server/src/copilot.ts
Original file line number Diff line number Diff line change
@@ -1,5 +1,6 @@
import type { BaseEvent, Message, RunAgentInput } from "@ag-ui/client";
import { AbstractAgent, HttpAgent } from "@ag-ui/client";
import { createOpenAI } from "@ai-sdk/openai";
import type { BuiltInAgentConfiguration } from "@copilotkit/runtime/v2";
import {
BuiltInAgent,
Expand Down Expand Up @@ -202,6 +203,59 @@ export function normalizeModelBaseUrls(
}
}

/**
* The model a built-in Bot answers on, as the runtime's `provider/model` string unless a
* deployment needs the chat-completions client built here instead.
*
* WHY AN INSTANCE IS EVER NEEDED. @copilotkit/runtime's resolveModel turns every `openai/<model>`
* string into the OpenAI **Responses API** — correct for api.openai.com, which serves `/responses`,
* and a 404 on every OpenAI-compatible endpoint that implements only `/chat/completions` (Z.AI's
* coding endpoint among them). No model-string form selects chat completions, so
* `OPENBOT_OPENAI_CHAT_COMPLETIONS=true` builds the chat model directly, with the same resolved
* key, the same `OPENAI_BASE_URL`, and the same model name the string form would have used: one
* line moves a whole deployment onto an endpoint the Responses-shaped call cannot reach.
*
* Chat completions is the older API and carries no reasoning summaries or server-side state, so
* this stays opt-in rather than becoming the default for every unknown base URL.
*/
function builtInModelSpecifier(
model: RuntimeModel,
apiKey: string | null,
environment: Record<string, string | undefined> = process.env,
) {
if (
environment.OPENBOT_OPENAI_CHAT_COMPLETIONS !== "true" ||
model.provider !== "openai"
) {
return `${model.provider}/${model.defaultModel}`;
}
return createOpenAI({
apiKey: apiKey ?? environment.OPENAI_API_KEY,
...(environment.OPENAI_BASE_URL?.trim()
? { baseURL: environment.OPENAI_BASE_URL.trim() }
: {}),
}).chat(model.defaultModel);
}

/**
* An explicit output-token budget for the built-in Bots.
*
* The AI SDK runs models it does not recognise — every name outside its per-provider tables, which
* is every model behind a compatible endpoint — in a compatibility mode that caps output at 4096
* tokens. A thinking model whose reasoning alone runs past that ceiling spends its whole budget on
* thinking and returns no text and no tool calls: the turn ends silently, which reads as a Bot
* that ignored you. `OPENBOT_MAX_OUTPUT_TOKENS` names a budget the model actually fits in; unset
* keeps the SDK's own behaviour for models it knows.
*/
function maxOutputTokensForEnvironment(
environment: Record<string, string | undefined> = process.env,
): number | undefined {
const raw = environment.OPENBOT_MAX_OUTPUT_TOKENS?.trim();
if (!raw) return undefined;
const parsed = Number.parseInt(raw, 10);
return Number.isFinite(parsed) && parsed > 0 ? parsed : undefined;
}

export function runtimeModelForEnvironment(
packageModel: RuntimeModel,
environment: Record<string, string | undefined> = process.env,
Expand Down Expand Up @@ -399,8 +453,11 @@ export function builtInAgentConfiguration(
})) ?? []),
];

const outputBudget = maxOutputTokensForEnvironment();
return {
model: planModel ?? `${model.provider}/${model.defaultModel}`,
model: planModel ?? builtInModelSpecifier(model, apiKey),
// See maxOutputTokensForEnvironment: absent for models the SDK knows, explicit for the rest.
...(outputBudget !== undefined ? { maxOutputTokens: outputBudget } : {}),
/*
* The package's role, then the person's own standing instructions, then what this Bot actually
* holds, then the computer.
Expand Down