You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
Model API translation: one gateway per deployment now, no translation where avoidable later #464
Codex speaks only OpenAI's Responses API. Self-hosted model servers speak Chat Completions reliably, but their Responses support is incomplete (tool calls in particular). The kernel bridges the gap with a translator: today a per-session LiteLLM sidecar ("chat bridge") on self-hosted endpoints, and on one deployment a shared LiteLLM gateway. Translation is a moving part between every harness upgrade and every model server:
A harness upgrade can add request shapes the translator does not know. Codex 0.160 added a sub-agent tool and a web-search tool; the per-session bridge rejected them and every Codex session on a chat-only endpoint failed with "unsupported bridge request" (fixed by keeping those tools off, Codex: no built-in sub-agents; web search per setting (fixes chat-endpoint sessions) #463).
Two translation paths behaved differently for the same harness (the shared gateway let the new tools through), so a change was tested on one path and broke the other.
A translator can silently change what the model sees (dropped or rewritten tools, reasoning, streaming quirks), which matters for measured experiments.
Direction
Near term (0.4): one model gateway per deployment instead of a sidecar per session: one pinned service on the same path for every backend, with per-role credentials, request logging and token accounting, and the strict "translate only what is faithful" check as a gateway rule. Design note to follow.
Long term: no translation where avoidable.
Prefer model servers whose Responses API handles tool calls natively (track vLLM and SGLang Responses support for the open models we serve); the gateway then passes requests through.
Where a harness has a native API that the server already speaks (Claude Code on an Anthropic-compatible endpoint, as the on-prem GLM fleet does today), use it directly.
Keep a small conformance check per (harness version, server version) pair: a scripted session exercising tool calls, resume and streaming, run when either side is upgraded, so an incompatibility is found before a fleet runs into it.
Harness upgrades become a checklist item: diff the request shapes a new harness version sends (tools, input item types, content parts) against what each path accepts, as done for Codex 0.160.
Done when
One translation component (or none) per deployment, identical across sites.
Conformance check runs on every harness or server pin change.
Open models served through native Responses where the server supports it.
Problem
Codex speaks only OpenAI's Responses API. Self-hosted model servers speak Chat Completions reliably, but their Responses support is incomplete (tool calls in particular). The kernel bridges the gap with a translator: today a per-session LiteLLM sidecar ("chat bridge") on self-hosted endpoints, and on one deployment a shared LiteLLM gateway. Translation is a moving part between every harness upgrade and every model server:
Direction
Done when