Skip to content
This repository was archived by the owner on Apr 26, 2026. It is now read-only.

feat: selective tool proxying, session isolation, MCP bridging, and streaming fixes - #16

Open
khalilgharbaoui wants to merge 402 commits into
unixfox:masterfrom
khalilgharbaoui:master
Open

khalilgharbaoui wants to merge 402 commits into
unixfox:masterfrom
khalilgharbaoui:master

Conversation

@khalilgharbaoui

@khalilgharbaoui khalilgharbaoui commented Apr 24, 2026 •

Copy link
Copy Markdown

Summary

This PR brings 18 commits that address the three known limitations listed in the original README and add significant new functionality. The changes fall into four areas:

1. Selective Tool Proxy — route dangerous tools through opencode's permission system

The headline feature. Claude CLI normally executes tools (Bash, Edit, Write, WebFetch) internally, bypassing opencode's permission UI entirely. This PR adds a proxyTools option that selectively disables Claude's built-in tools and replaces them with equivalent MCP proxy tools hosted by an in-process HTTP server.

How it works:

  • An embedded MCP server starts on 127.0.0.1 (random port) when proxyTools is configured.
  • For each proxied tool, --disallowedTools <ToolName> is passed to the CLI.
  • Claude calls the MCP proxy tool instead → the plugin emits a client-executed tool-call to opencode → opencode runs the real tool with its native permission checks → the result flows back to Claude.
  • Non-proxied tools (Read, Glob, Grep, etc.) remain fully native to Claude CLI for performance.

New files: src/proxy-mcp.ts (MCP server), src/proxy-broker.ts (pause/resume broker).

Supported proxy tools: Bash, Edit, Write, WebFetch.

Config:

{
  "options": {
    "proxyTools": ["Bash", "Edit", "Write", "WebFetch"]
  }
}

2. Session isolation — no more cross-chat interference

Sessions are now keyed by (cwd, model, x-session-affinity) instead of just (cwd, model). The x-session-affinity header is set by opencode on LLM calls to third-party providers, so two simultaneous chats in the same project get separate CLI processes. An LRU cap (16 processes) prevents subprocess accumulation.

3. MCP config auto-bridging — one config, not two

The plugin now auto-discovers opencode.json / opencode.jsonc (via cwd, OPENCODE_CONFIG, OPENCODE_CONFIG_DIR, $XDG_CONFIG_HOME/opencode) and translates its mcp block into Claude CLI's --mcp-config format. Local servers get type: \"stdio\", remote servers get type: \"http\", disabled servers are skipped. This means MCP servers configured in opencode are automatically available to Claude CLI without maintaining a separate ~/.claude/settings.json.

New file: src/mcp-bridge.ts.

Config overrides: bridgeOpencodeMcp (default true), mcpConfig (extra paths), strictMcpConfig.

4. Streaming correctness fixes

  • Empty content sentinel (0736306, 5def53c): replaced \"(continue)\" with \"(empty)\" so the model doesn't interpret the sentinel as an instruction to resume the previous turn.
  • Tool-execution semantics (33cb03a): TodoWrite and WebSearch are now forwarded as client-executed (not provider-executed), so opencode's todo UI and search results populate correctly.
  • Object-shaped tools (c665524): opencode sometimes passes tools as an object map rather than an array; the scope classifier now handles both.
  • CLI error surfacing (09db874): if Claude returns only a result message with error text (rate limit, auth failure), it's now emitted as visible text instead of a blank turn.
  • Control request handling (70badf9): can_use_tool control requests get immediate control_response replies with configurable allow/deny policy, preventing stream deadlocks.
  • Per-iteration usage (6d126c3, refined in 4af2a96): uses usage.iterations[-1] instead of cumulative totals and computes inputTokens.total = noCache + cacheRead + cacheWrite, preventing inflated context estimates and fixing cache-aware token accounting.
  • Per-block text emission (6d126c3, refined in 4af2a96): each text content block gets its own text-start/delta/text-end lifecycle so partial text is preserved on stream abort.
  • Result fallback timing (6d126c3, refined in 4af2a96, tightened in ce5701c): a 5-second timeout closes the stream gracefully if the CLI emits content but never sends a result event. The timer is now only armed on assistant text without tool use, abort starts a grace period instead of closing immediately, and the non-streaming path now honors proxied tools consistently.
  • Anthropic cache metadata (4af2a96): emits providerMetadata.anthropic.cacheCreationInputTokens so OpenCode can display cache write tokens correctly.
  • Lazy cwd resolution (ce3eb26): provider init no longer freezes process.cwd(), so each request resolves cwd at call time.

Other improvements

  • AI SDK v3 compatibility (0ae354c)
  • Reasoning effort levels (93d610c): --thinking-effort passthrough for low/medium/high/xhigh/max
  • Image input support (93d610c, hardened in 4af2a96): base64 image parts forwarded to Claude CLI, plus MIME allowlist, robust data URI parsing, and early rejection of unsupported remote URL images
  • --permission-mode passthrough (ea27f17)
  • Windows compatibility (6d126c3): shell: process.platform === \"win32\" on both spawn sites so claude.cmd works on Windows
  • Comprehensive README rewrite with architecture diagrams, config reference, and proxy documentation

Relationship to other open PRs

This PR subsumes or addresses the core concerns of several other open PRs. We developed these independently and discovered many of the same issues:

Open PR Author What it does How this PR addresses it
#6 @simonseo Stream finish handling + effort passthrough + session scoping by effort We fix stream finish (result fallback timer, per-block text), pass --thinking-effort, and scope sessions by x-session-affinity header. We intentionally keep our variant-based effort approach rather than model-suffix ergonomics.
#9 @nbalzotti Windows shell:true for .cmd spawn Included in 6d126c3 — same fix on both spawn sites.
#12 @Aptul9 AI SDK V3 migration, per-iteration usage, per-block text, result fallback timer, providerExecuted flag, empty content We independently implemented all of these and then tightened the last details in 4af2a96: V3 spec (0ae354c), lastIterationUsage via iterations[-1] (6d126c3), cache-aware totals + noCache (4af2a96), per-block text emission (6d126c3), smarter fallback timing + abort grace (4af2a96), providerExecuted (33cb03a, refined in 4af2a96), empty content sentinel (0736306, 5def53c).
#13 @Aptul9 Image support in user messages Included in 93d610c; hardened in 4af2a96 with supported MIME allowlist, robust data URI parsing, and remote URL rejection.
#15 @waveywaves --effort flag via provider option Included in 93d610c — reasoning effort passthrough.

PR #4 is only partially addressed here. Commit ce3eb26 adopts the safe cross-platform piece by resolving cwd lazily per request instead of freezing process.cwd() at provider initialization. We intentionally did not adopt the desktop-specific SQLite/session lookup fallback, request-option sessionID/cwd plumbing, or hard-coded path logic from #4, so #4 remains distinct draft work for desktop-specific cwd recovery.

PR #14 is independent and useful, but it's a standalone migration utility rather than a runtime plugin improvement.

Issues addressed


Commits (chronological)

  1. 0ae354c fix: make claude-code provider compatible with AI SDK v3
  2. 93d610c feat: add reasoning effort levels and image input support
  3. 0736306 fix: use neutral sentinel instead of "(continue)" for empty user content
  4. 33cb03a fix: correct tool-execution semantics for opencode-hosted tools
  5. 5def53c fix: use "(empty)" sentinel matching provider's parenthetical meta-note convention
  6. ea27f17 feat: expose --mcp-config passthrough and fix known-limitations wording
  7. 1941685 feat: fix session sharing and auto-bridge opencode MCP config to Claude CLI
  8. 70badf9 feat: handle Claude control-request permissions in stream-json mode
  9. 09db874 fix: surface CLI error text from stream-json result messages
  10. c665524 fix: detect object-shaped tools when choosing stream scope
  11. a663266 fix: emit Claude-compatible MCP transport types in bridge
  12. 4145493 feat: proxy Bash through opencode tools and permissions
  13. 9230421 feat: proxy Edit and Write through opencode tools
  14. 820cc22 feat: proxy WebFetch through opencode tools and permissions
  15. 6d126c3 fix: per-iteration usage, per-block text emission, result fallback timer, Windows spawn
  16. 4af2a96 fix: refine usage accounting, text emission, fallback timing, and image handling
  17. ce5701c fix: honor proxied tools in doGenerate and tighten fallback handling
  18. ce3eb26 fix: resolve cwd lazily per request

Test plan

  • tsc --noEmit passes
  • tsup build passes
  • opencode run \"hi\" -m claude-code/claude-sonnet-4-6 returns visible output (or explicit rate-limit text, not blank)
  • Proxy Bash: Claude calls mcp__opencode_proxy__bash, opencode executes, result flows back
  • Proxy Edit: Claude calls mcp__opencode_proxy__edit, opencode executes file diff
  • Proxy Write: Claude calls mcp__opencode_proxy__write, opencode writes file
  • Proxy WebFetch: proxied MCP tool is exposed and wired through the same selective proxy path
  • Proxy with bash: ask permission rule: opencode's permission.asked fires, auto-rejected in headless mode
  • MCP bridge: opencode MCP config translated to Claude CLI format (local → stdio, remote → http, disabled → skipped)
  • Session isolation: two chats with different x-session-affinity headers get separate CLI processes
  • Empty content: whitespace-only messages produce \"(empty)\" sentinel, not blank or \"(continue)\"
  • Rate-limit: 429 responses surface visible error text instead of blank turn
  • Per-iteration usage: usage.iterations[-1] used when present, falls back to cumulative
  • Cache-aware input totals: inputTokens.total includes cache read/write, noCache is populated
  • Windows: shell: true gated on process.platform === \"win32\"
  • Image handling: supported MIME types accepted, malformed data URIs and remote URLs rejected early
  • Lazy cwd resolution: provider no longer freezes init-time process.cwd()

Breaking changes

None. All new features are opt-in via config. Default behavior is unchanged from upstream.

Known limitations

  • Proxy tool set: only Bash, Edit, Write, and WebFetch are supported. More can be added when opencode gains matching built-in executors.
  • Non-proxied tools bypass opencode permissions: Read, Glob, Grep, etc. remain native to Claude CLI for performance.
  • Claude upstream bug #34046: Claude CLI does not emit can_use_tool control requests for built-in tools. The selective proxy approach works around this entirely.

@emreycolakoglu

Copy link
Copy Markdown

@khalilgharbaoui would you consider publishing to npm yourself? this repo is likely dead. I'm looking forward to use your fixes but I couldn't use it locally (clone + build).

@khalilgharbaoui

Copy link
Copy Markdown
Author

@emreycolakoglu yep — I went ahead and published a maintained fork. It is on npm now, no clone/build needed:

Just add it to the plugin array in your opencode.json — the README has the up-to-date install/config and a few quirks worth knowing about (selective tool proxying, MCP bridge discovery order, MultiEdit pass-through, plan mode handling). Worth a quick read before wiring it up.

Includes the fixes from this PR plus a couple of regressions I hit afterwards (empty-text-block 400s, variant selection on model pick, lazy cwd resolution). Issues / PRs welcome over on the fork.

khalilgharbaoui and others added 24 commits June 10, 2026 17:02
The CLI treats --session-id as 'create a NEW session with this UUID'
and exits with 'Session ID ... is already in use' whenever a transcript
for that ID already exists on disk — so every respawn that tried to
continue a session (MCP hot reload, eviction, crash recovery) failed,
cleared the session, and fell back to re-injecting history as text.
Verified against the real CLI: --session-id reuse fails with no live
process holding the ID; --resume continues under the same session ID.

Also catch the lowercase 'No conversation found with session ID' error
that --resume prints for a purged transcript, so a stale remembered ID
still self-heals on the next turn.
The proxy server's close() rejects any in-flight tools/call so the HTTP
handler doesn't hang, but that rejection was logged at WARN — surfacing
a yellow TUI bubble on every normal teardown (process exit, abort kill,
MCP hot-reload respawn, compaction). By the time close() runs, the
owning Claude process is gone or being replaced, so nobody can consume
the response; the rejection is pure cleanup.

Extract the expected-cleanup classification into an exported
isExpectedCleanupError(), add the server-closed message to it, and share
the message string via SERVER_CLOSED_MESSAGE so the classifier and
close() cannot drift apart. Genuine errors still log at WARN.
Task subagents and long bash builds were getting killed at the flat
10-minute proxy ceiling, even when the caller passed a larger bash
input.timeout. A Task proxy timeout fired mid-subagent, Claude believed
its dispatch had failed, "scheduled a wake-up" (an affordance that can't
fire headless/proxy), and the eventual result was dropped.

resolveProxyCallTimeoutMs layers flat -> per-tool (task 60m, question
30m) -> proxyToolTimeoutMs override -> bash input.timeout (max only,
never undercuts). Both timeout sites (proxy-mcp handler + broker) share
the one resolver so they never race. The Task timeout message tells the
model not to schedule a wake-up or defer. Resolved values clamp to
Node's 2^31-1 ms timer max.
Claude CLI validates every tools/call response against the MCP result
schema and rejects JSON-RPC error envelopes as malformed. All three
error paths (unknown tool, kind:error results, broker rejections via
the outer catch) now return results with isError: true. Found live by
@jknlsn (2026-07-04); enforced by their test-proxy-mcp.ts suite.
Also reconciles the --mcp-config client ceiling with per-tool deadlines
via resolveProxyClientCeilingMs.
A reused claude --print child can go silent on stdout after a fresh-turn
envelope write. This was masked before the per-tool proxy timeout fix
because the flat 10-minute ceiling ended the turn first; now that a long
proxy-blocked task call blocks and returns successfully, resuming the
reused child afterwards can leave it producing nothing (live ses_0cfc0da6
step 8 — idle, 0% CPU, no network, no error, needed a manual Esc).

Add a start watchdog (doStream, fresh-turn path only; default 90s, env
CLAUDE_CODE_START_WATCHDOG_MS). On first fire it respawns the child via
respawnActiveProcess, which kills the wedged child but reuses its proxy
server, system-prompt file, and mcp hash (handles baked into cliArgs) and
appends --session-id so the conversation resumes transparently. A second
fire ends the turn with an error so the next turn spawns fresh. Complementary
to the existing inactivity watchdog, which deliberately skips the
pre-content gap.
Re-audited the plugin against opencode 1.18.5: nothing we depend on
broke, but opencode still hands plugins no version. Since the plugin
runs inside opencode's process, process.execPath is the opencode binary,
so probe it for --version (cached, guarded on basename so a bun-run
source checkout reports unknown instead of Bun's version).

Audit findings and the new 1.18.5 surface (v2 plugin API, cost tiers,
tool.definition, compaction hooks) are documented in AGENTS.md and
tracked in #24.
khalilgharbaoui and others added 30 commits September 28, 2026 06:39
* Surface skipped MCP entries and plan usage

Two measurement-first roadmap items on Claude Code 2.1.280.

Digest-turn measurement: no code changed. Meridian's hidden
"digest" turn does not exist on this architecture; what a proxied
tool does cost is one extra API call per claude process, and it is
a ToolSearch, because the CLI defers MCP tools behind it. It
amortises across calls and the only lever that removes it
(ENABLE_TOOL_SEARCH=0) measured 2.5 to 4 times the cost. Numbers
in docs/agents-history.md #g166, summary in the README and skill.

Free CLI diagnostics: mcp_server_errors is now parsed off the
system/init frame and warned once per name:type, because a skipped
--mcp-config entry is absent from mcp_servers rather than listed
broken, and the entry that goes missing can be the plugin's own
proxy. The doctor gains that section plus an opt-in Plan usage
section (/claude-code-doctor usage) carrying the CLI's own free
/cost answer. /advisor does not exist headless and /reload-plugins
is an action rather than a diagnostic, so neither got code; both
verdicts are recorded in #g167.

* Spawn the usage probe with the turn spawn env
* Add a per-agent prompt cache TTL

The CLI's subagent steering knobs reach nothing here: every process this
plugin spawns is a main conversation, because `Task` is disallowed by
default so opencode runs the subagent instead. Measured on 2.1.280.

The main-conversation knob does reach, so an agent can declare
`cacheTtl: 5m` (or `defaultSubagentCacheTtl` for all discovered
subagents) and short-lived workers stop writing 1-hour caches they
never read.

* Note that cacheTtl is headless only
opencode reads a finish part's usage as how full the context is and
auto-compacts when input + output + cache read + cache write reaches
context - maxOutputTokens. On the headless transport every finish
reported the terminal result's usage, which the CLI sums over every API
call in the turn. A 14-call turn whose real context grew 121K -> 180K
reported 2,050,806 cache-read tokens (the sum of its per-call reads),
and opencode compacted four times in 32 minutes with the real context
never above 204K.

toUsage meant to avoid this by preferring the last entry of
usage.iterations, but that field is not a per-tool-loop breakdown: the
CLI (2.1.281) copies it from the last API response's server-side
sampling iterations, normally [], so the turn sum was used every time.

Track the newest assistant frame's message.usage whose input + cache
counts are non-zero (a <synthetic> frame never overwrites it) and report
that call's input side with the result's turn-total output, the same
convention the interactive transport already uses (416bef0). With no
call seen the result's usage is reported as before, and a tool-call
finish with no result yet still reports zeros, so opencode cannot
compact with a proxied call parked. anthropic.cacheCreationInputTokens
now matches the reported cache write, because opencode falls back to it
when the usage carries none.

providerMetadata["claude-code"] and turnStats keep the turn totals, so
the CLI's true cost stays available; opencode's own cost figure for a
multi-call turn now counts only the last call's input and cache.
Measure opencode's native background task mode and surface it: gate the
`background` field on the host's own advertised task schema, add
task_status and task_cancel for collect and cancel, and document what
was measured.
)

The interactive transport put the stop reason in the synthesized
result's subtype, so a clean turn arrived as subtype "end_turn" and
every completed turn finished as an error, which also suppressed
turnStats. A real headless result keeps subtype "success" and carries
the stop reason in a top-level stop_reason field.

Also name the transcript directory from the cwd's real path, as the
CLI does, so a symlinked cwd (on macOS any /tmp path) no longer tails
a file that is never written.
…doctor (#69)

`task_status` and `task_cancel` took their no-client branches on every
opencode 2 host: the V1-shaped client shim exposed `session.get` only, so
`fetchSessionReplies` and `abortSession` had nothing to call.

The shim now answers `session.messages` from V2's `session.context` and
`session.abort` from `session.interrupt`, the two routes a plugin actually
gets there. V2 offers a plugin no all-sessions run-state map, so
`isBackgroundTaskRunning` falls back to the child's own transcript, and
`abortSession` reads the boolean both majors return so a refused cancel
cannot render as a successful one. Neither envelope is the same on both
majors, so the background note is per dialect.

Adds a Background subagents section to /claude-code-doctor: whether
`background` was offered and the two tools registered, the opencode major,
what decided it, and the tasks this process collected or cancelled.

Verified live on opencode 2.0.16.
* Forget a probe killed by its own deadline

A 5s execFile deadline on 'claude --version' and '--help' cached its own
timeout, so one busy turn permanently withheld every version-gated flag,
the skill bridge and /btw from the opencode process. That is what made
test-side-question's native /btw spec flaky under load. Keep the
deadline, drop only the answer that described the machine, and make the
specs resolve their probes up front instead of racing them.

* Cap deadline re-probes at two per probe key
… safe boundary (#72)

opencode 2 reports a server it is still connecting to as `pending`, a
sixth status opencode 1 does not have, and the runtime-status overlay
read it as "not connected" and dropped the server from that turn's
spawn. A reused process keeps the config it was spawned with, so it
stayed invisible for the rest of the conversation.

`pending` now leaves the configured value alone, as a missing entry
already does, and the first turn waits up to `mcpConnectWaitMs` (3 s,
`0` disables) for the host to decide. opencode 1 never engages it: its
own status call blocks until every server resolves.

The hot reload that catches a genuinely later connect already worked
and is kept, but moved behind `decideMcpHotReload`: a safe boundary
(not interactive, no pending proxy call, no turn in flight, no
outstanding plan approval), server names rather than hashes in the
log, and one respawn per conversation per minute so a flapping server
cannot respawn every turn.
Sign up for free to subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.