Skip to content

Model API translation: one gateway per deployment now, no translation where avoidable later #464

Description

@renmengye

Problem

Codex speaks only OpenAI's Responses API. Self-hosted model servers speak Chat Completions reliably, but their Responses support is incomplete (tool calls in particular). The kernel bridges the gap with a translator: today a per-session LiteLLM sidecar ("chat bridge") on self-hosted endpoints, and on one deployment a shared LiteLLM gateway. Translation is a moving part between every harness upgrade and every model server:

  • A harness upgrade can add request shapes the translator does not know. Codex 0.160 added a sub-agent tool and a web-search tool; the per-session bridge rejected them and every Codex session on a chat-only endpoint failed with "unsupported bridge request" (fixed by keeping those tools off, Codex: no built-in sub-agents; web search per setting (fixes chat-endpoint sessions) #463).
  • Two translation paths behaved differently for the same harness (the shared gateway let the new tools through), so a change was tested on one path and broke the other.
  • A translator can silently change what the model sees (dropped or rewritten tools, reasoning, streaming quirks), which matters for measured experiments.

Direction

  1. Near term (0.4): one model gateway per deployment instead of a sidecar per session: one pinned service on the same path for every backend, with per-role credentials, request logging and token accounting, and the strict "translate only what is faithful" check as a gateway rule. Design note to follow.
  2. Long term: no translation where avoidable.
    • Prefer model servers whose Responses API handles tool calls natively (track vLLM and SGLang Responses support for the open models we serve); the gateway then passes requests through.
    • Where a harness has a native API that the server already speaks (Claude Code on an Anthropic-compatible endpoint, as the on-prem GLM fleet does today), use it directly.
    • Keep a small conformance check per (harness version, server version) pair: a scripted session exercising tool calls, resume and streaming, run when either side is upgraded, so an incompatibility is found before a fleet runs into it.
  3. Harness upgrades become a checklist item: diff the request shapes a new harness version sends (tools, input item types, content parts) against what each path accepts, as done for Codex 0.160.

Done when

  • One translation component (or none) per deployment, identical across sites.
  • Conformance check runs on every harness or server pin change.
  • Open models served through native Responses where the server supports it.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions