Skip to content

Report the last API call's context as usage, not the turn sum - #63

Merged
khalilgharbaoui merged 1 commit into
khalilgharbaoui:masterfrom
broskees:fix/report-context-usage-not-turn-sum
Sep 29, 2026
Merged

khalilgharbaoui merged 1 commit into
khalilgharbaoui:masterfrom
broskees:fix/report-context-usage-not-turn-sum

Conversation

@broskees

Copy link
Copy Markdown

What this is

On the default headless transport, every finish reported the CLI result frame's usage, and that usage is the sum over every API call in the turn. opencode reads a finish's usage as how full the context is, so any tool-heavy turn looked several times its real size and tripped auto-compaction far below the window.

Measured on a real session (Opus 5.5, 1M window, plugin 0.27.1, opencode 1.18.29):

Turn Real context (Claude transcript, last call) What opencode recorded Next message
12 calls 109K 1.13M compaction
14 calls 180K 2.11M compaction
1 call 182K 182K normal
6 calls 195K 1.14M compaction
7 calls 204K 1.41M compaction

The 14-call turn's recorded cache read, 2,050,806, is exactly the sum of its 14 per-call cache reads in the transcript. Four compactions in 32 minutes, real context never above 204K, and the compactions did not shrink anything, because the plugin resumed the same Claude session afterwards.

Why toUsage's iterations preference never fired

toUsage already meant to avoid this (h #g109: "summing iterations inflates the window and triggers premature compaction"), by preferring usage.iterations[last]. But iterations is not a per-tool-loop breakdown. Read out of the CLI 2.1.281 binary: the cumulative accumulator sums input/output/cache across API responses and sets iterations: n.iterations, i.e. it copies the last response's server-side iterations field (per-request sampling iterations, normally []). So iter?.length was 0 and the turn sum was used every time. The CLI's own context-size helper works per assistant message.

The fix

Report what Anthropic documents for the Agent SDK (cost tracking): input and cache counts per API call from the assistant messages, output from the result, because "per-step output_tokens is a placeholder" and several assistant messages of one call share an id and identical usage. This is also the convention the interactive transport already uses (416bef0).

  • stream-parser.ts keeps the newest assistant frame's message.usage whose input + cache counts are non-zero (TurnState.lastCallUsage); a <synthetic> all-zero frame never overwrites it.
  • lastCallContextUsage (src/usage.ts) combines it with the result: the last call's input side (its last iterations entry if it carries one), the result's turn-total output. No call seen means the result's usage exactly as before.
  • All four finish sites use it. A tool-call finish with no result yet (the ordinary mid-turn proxied boundary) still reports zeros: real context there would let opencode compact while the CLI is parked in a proxied call, which nobody has verified is safe.
  • providerMetadata.anthropic.cacheCreationInputTokens now equals the reported cache write. opencode's getUsage does usage.cacheWriteInputTokens ?? metadata.anthropic.cacheCreationInputTokens, and toUsage sends a 0 cache write as undefined, so leaving the metadata at the turn total would put the sum straight back.
  • Unchanged: providerMetadata["claude-code"] (turn usage, costUsd) and turnStats, which must match total_cost_usd.

Trade-off, stated plainly: opencode computes its cost figure from these same tokens, so for a multi-call turn it now counts only the last call's input and cache plus all output. The CLI's real turn cost remains in costUsd and the turnStats line. README and SKILL.md say so.

Relation to #62

#62 (@bangnh1) found the same bug today and goes the same direction. This differs in three places, each measurable:

  1. Output tokens. miscalculated the token count #62 sums output_tokens over every assistant frame. One API call arrives as one frame per content block with identical usage, and per-frame output is a placeholder (the doc above). In the same session's transcript, 53 API calls were 134 records, and summing per record gave 174,557 output tokens against 62,722 counted once per call. Here output comes from result.usage.output_tokens, the CLI's own count.
  2. The metadata fallback. miscalculated the token count #62 leaves anthropic.cacheCreationInputTokens at the turn total, which opencode reads whenever the last call's cache write is 0.
  3. Synthetic frames. In miscalculated the token count #62 a trailing all-zero <synthetic> frame replaces the last real call, and the context reads 0.

#62 also carries the fix through a synthetic iterations entry, which keeps the misreading of that field alive; here iterations keeps its real meaning and is honoured per call. Either PR landing is fine by me; if you prefer #62, the three points above apply to it.

Not changed here, and noted in the history entry: the interactive transport's tailTurn also sums output_tokens per transcript record (416bef0), so it over-counts output the same way. It needs its own per-message-id fix and a Bun run to verify.

Tests

test-context-usage.ts, a fake CLI through a real doStream (registered in package.json):

  • three calls at 100K/110K/120K cache read and a summed result with iterations: []: the finish reports the last call's context, the result's output, and a matching anthropic.cacheCreationInputTokens;
  • a synthetic zero-usage frame after a real call does not reset it;
  • no assistant usage in the stream: the result's usage as before;
  • a mid-turn proxied boundary still reports zeros; a tool-call finish written after a result uses the last call;
  • a pure case for a call carrying non-empty iterations.

With src/stream-parser.ts, src/claude-code-language-model.ts and src/turn-controller.ts reverted to master, three fail (e.g. actual: 334315, expected: 120807, the turn sum instead of the last call) and the two guards pass either way.

npm run typecheck clean, npm test 903/903, npm run build clean. AGENTS.md rule plus history entry #g169 (next to g109, since it corrects g109's mechanism), README and one SKILL.md troubleshooting row.

Not live-verified: I did not run a paid probe. The evidence is the fake-CLI suite plus the real session's transcript and opencode DB above.

opencode reads a finish part's usage as how full the context is and
auto-compacts when input + output + cache read + cache write reaches
context - maxOutputTokens. On the headless transport every finish
reported the terminal result's usage, which the CLI sums over every API
call in the turn. A 14-call turn whose real context grew 121K -> 180K
reported 2,050,806 cache-read tokens (the sum of its per-call reads),
and opencode compacted four times in 32 minutes with the real context
never above 204K.

toUsage meant to avoid this by preferring the last entry of
usage.iterations, but that field is not a per-tool-loop breakdown: the
CLI (2.1.281) copies it from the last API response's server-side
sampling iterations, normally [], so the turn sum was used every time.

Track the newest assistant frame's message.usage whose input + cache
counts are non-zero (a <synthetic> frame never overwrites it) and report
that call's input side with the result's turn-total output, the same
convention the interactive transport already uses (416bef0). With no
call seen the result's usage is reported as before, and a tool-call
finish with no result yet still reports zeros, so opencode cannot
compact with a proxied call parked. anthropic.cacheCreationInputTokens
now matches the reported cache write, because opencode falls back to it
when the usage carries none.

providerMetadata["claude-code"] and turnStats keep the turn totals, so
the CLI's true cost stays available; opencode's own cost figure for a
multi-call turn now counts only the last call's input and cache.
@khalilgharbaoui
khalilgharbaoui merged commit 802ded5 into khalilgharbaoui:master Sep 29, 2026
@khalilgharbaoui

Copy link
Copy Markdown
Owner

Merged, thank you @broskees. I checked the three points against the real CLI before merging, and all three hold.

  • One API call arrives as several assistant frames with identical usage. A Haiku turn with four tool calls on Claude Code 2.1.28x produced 11 frames for 5 calls. The frames' output_tokens read 8, 3, 1, 1, 1 while the calls' real outputs (from message_delta) were 180, 93, 95, 92, 30, which sum to the result's 490.
  • The fallback is real. opencode 1.18.33's binary reads usage.cacheWriteInputTokens ?? metadata.anthropic.cacheCreationInputTokens.
  • End to end A/B through opencode 1.18.33, same prompt, scratch dirs: master stored 193,659 tokens (167,227 cache read, five calls summed) for the message; this branch stored 39,681 (38,982 cache read, the last call).

You are credited in the README. The interactive transport's per-record output sum you flagged is recorded as a follow-up.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants