Report the last API call's context as usage, not the turn sum - #63
Merged
khalilgharbaoui merged 1 commit intoSep 29, 2026
Conversation
opencode reads a finish part's usage as how full the context is and auto-compacts when input + output + cache read + cache write reaches context - maxOutputTokens. On the headless transport every finish reported the terminal result's usage, which the CLI sums over every API call in the turn. A 14-call turn whose real context grew 121K -> 180K reported 2,050,806 cache-read tokens (the sum of its per-call reads), and opencode compacted four times in 32 minutes with the real context never above 204K. toUsage meant to avoid this by preferring the last entry of usage.iterations, but that field is not a per-tool-loop breakdown: the CLI (2.1.281) copies it from the last API response's server-side sampling iterations, normally [], so the turn sum was used every time. Track the newest assistant frame's message.usage whose input + cache counts are non-zero (a <synthetic> frame never overwrites it) and report that call's input side with the result's turn-total output, the same convention the interactive transport already uses (416bef0). With no call seen the result's usage is reported as before, and a tool-call finish with no result yet still reports zeros, so opencode cannot compact with a proxied call parked. anthropic.cacheCreationInputTokens now matches the reported cache write, because opencode falls back to it when the usage carries none. providerMetadata["claude-code"] and turnStats keep the turn totals, so the CLI's true cost stays available; opencode's own cost figure for a multi-call turn now counts only the last call's input and cache.
Owner
|
Merged, thank you @broskees. I checked the three points against the real CLI before merging, and all three hold.
You are credited in the README. The interactive transport's per-record output sum you flagged is recorded as a follow-up. |
This was referenced Sep 29, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What this is
On the default headless transport, every finish reported the CLI
resultframe's usage, and that usage is the sum over every API call in the turn. opencode reads a finish's usage as how full the context is, so any tool-heavy turn looked several times its real size and tripped auto-compaction far below the window.Measured on a real session (Opus 5.5, 1M window, plugin 0.27.1, opencode 1.18.29):
The 14-call turn's recorded cache read, 2,050,806, is exactly the sum of its 14 per-call cache reads in the transcript. Four compactions in 32 minutes, real context never above 204K, and the compactions did not shrink anything, because the plugin resumed the same Claude session afterwards.
Why
toUsage'siterationspreference never firedtoUsagealready meant to avoid this (h #g109: "summing iterations inflates the window and triggers premature compaction"), by preferringusage.iterations[last]. Butiterationsis not a per-tool-loop breakdown. Read out of the CLI 2.1.281 binary: the cumulative accumulator sums input/output/cache across API responses and setsiterations: n.iterations, i.e. it copies the last response's server-sideiterationsfield (per-request sampling iterations, normally[]). Soiter?.lengthwas 0 and the turn sum was used every time. The CLI's own context-size helper works per assistant message.The fix
Report what Anthropic documents for the Agent SDK (cost tracking): input and cache counts per API call from the assistant messages, output from the result, because "per-step
output_tokensis a placeholder" and several assistant messages of one call share an id and identical usage. This is also the convention the interactive transport already uses (416bef0).stream-parser.tskeeps the newestassistantframe'smessage.usagewhose input + cache counts are non-zero (TurnState.lastCallUsage); a<synthetic>all-zero frame never overwrites it.lastCallContextUsage(src/usage.ts) combines it with the result: the last call's input side (its lastiterationsentry if it carries one), the result's turn-total output. No call seen means the result's usage exactly as before.resultyet (the ordinary mid-turn proxied boundary) still reports zeros: real context there would let opencode compact while the CLI is parked in a proxied call, which nobody has verified is safe.providerMetadata.anthropic.cacheCreationInputTokensnow equals the reported cache write. opencode'sgetUsagedoesusage.cacheWriteInputTokens ?? metadata.anthropic.cacheCreationInputTokens, andtoUsagesends a 0 cache write asundefined, so leaving the metadata at the turn total would put the sum straight back.providerMetadata["claude-code"](turnusage,costUsd) andturnStats, which must matchtotal_cost_usd.Trade-off, stated plainly: opencode computes its cost figure from these same tokens, so for a multi-call turn it now counts only the last call's input and cache plus all output. The CLI's real turn cost remains in
costUsdand theturnStatsline. README and SKILL.md say so.Relation to #62
#62 (@bangnh1) found the same bug today and goes the same direction. This differs in three places, each measurable:
output_tokensover everyassistantframe. One API call arrives as one frame per content block with identical usage, and per-frame output is a placeholder (the doc above). In the same session's transcript, 53 API calls were 134 records, and summing per record gave 174,557 output tokens against 62,722 counted once per call. Here output comes fromresult.usage.output_tokens, the CLI's own count.anthropic.cacheCreationInputTokensat the turn total, which opencode reads whenever the last call's cache write is 0.<synthetic>frame replaces the last real call, and the context reads 0.#62 also carries the fix through a synthetic
iterationsentry, which keeps the misreading of that field alive; hereiterationskeeps its real meaning and is honoured per call. Either PR landing is fine by me; if you prefer #62, the three points above apply to it.Not changed here, and noted in the history entry: the interactive transport's
tailTurnalso sumsoutput_tokensper transcript record (416bef0), so it over-counts output the same way. It needs its own per-message-id fix and a Bun run to verify.Tests
test-context-usage.ts, a fake CLI through a realdoStream(registered inpackage.json):resultwithiterations: []: the finish reports the last call's context, the result's output, and a matchinganthropic.cacheCreationInputTokens;resultuses the last call;iterations.With
src/stream-parser.ts,src/claude-code-language-model.tsandsrc/turn-controller.tsreverted tomaster, three fail (e.g.actual: 334315, expected: 120807, the turn sum instead of the last call) and the two guards pass either way.npm run typecheckclean,npm test903/903,npm run buildclean. AGENTS.md rule plus history entry#g169(next to g109, since it corrects g109's mechanism), README and one SKILL.md troubleshooting row.Not live-verified: I did not run a paid probe. The evidence is the fake-CLI suite plus the real session's transcript and opencode DB above.