Description
This is the new report requested in the closing comment of #43249 ("If you're seeing repeated cold compactions on
current V2, please open a new report with the exact version, provider/model, and input/cache-read counts").
On V2, summary compactions read almost nothing from the prompt cache. This also happens when the
previous model call in the same session ran less than a minute earlier and read more than 80% of its input from the
cache. On Anthropic Messages routes, about 96% of each compaction's input is billed as a new cache write.
Selection: compactions in the local opencode.db that completed on 2.0.12 or later (version estimated from the
sessions created just before), with at least 50k input tokens.
In each case, the previous model call in the session ran at most 5 minutes earlier and read at least 80% of its input
from the cache. Only 1 of these 132 compactions read at least 50% of its input from the cache.
| Route |
Compactions |
Versions |
Median input |
Median cache read |
Read share |
Write share |
Median gap since previous call |
Built-in openai provider (gpt-6-luna, gpt-6-sol, gpt-6-astra), auto |
113 |
2.0.16–2.0.21 |
253,321 |
20,992 |
7.4% |
0.0% |
0.68 min |
| Anthropic Messages (Anthropic API, Vertex AI, Azure AI Foundry; Claude models), mostly manual |
12 |
2.0.14–2.0.21 |
407,080 |
8,051 |
4.2% |
95.8% |
0.76 min |
Example rows (previous call input / previous call cache read → compaction input / cache read / cache write):
2.0.21 openai gpt-6-luna auto gap 0.54 min: 784,645 / 776,704 → 487,401 / 4,608 / 0
2.0.19 anthropic claude-sonnet-5-5 auto gap 0.31 min: 903,250 / 883,134 → 641,900 / 8,051 / 633,845
2.0.16 google-vertex claude-opus-5-5 manual gap 1.27 min: 837,966 / 836,224 → 672,165 / 0 / 672,161
2.0.22 google-vertex gemini-3.8-flash auto gap 0.57 min: 951,296 / 899,938 → 421,711 / 0 / 0
The constant read of about 21k tokens on OpenAI and about 8k tokens on Anthropic looks like the system prompt and tool
definitions only. The cache match appears to stop at the first conversation message.
Possible cause
In 2.0.22, summarize sends only the older part of the history, split.older, and keeps the recent part as text
outside the request:
For Anthropic, this would explain the result. The request's cache breakpoints sit at the end of the older part. The
cache entries that the main loop keeps fresh end at the end of the full history, so no live entry matches the shorter
prefix. For OpenAI, I would expect the shorter prefix to still match, so I could not explain the miss there. The
request may differ from the main-loop request in the first conversation message, or in transport or cache-key
routing. I did not confirm either of these.
The closing comment on #43249 cites a /compact test on 2.0.12 that read about 99% from the cache. The one hit in
our data (github-copilot, 2.0.20) read 113,792 of 114,744 input tokens. Its input was about the same size as the
previous call's input (113,815), which suggests that the older part covered almost the full history there.
Possible direction
One option is to send the main loop's full current request with the summary instruction appended, as
session.generate does, and apply the older/recent split to the stored result instead of to the request. #46369
proposes something in this direction as an opt-in. Related: #48542 (Anthropic compaction cache writes never read).
Plugins
A local plugin rewrites the cache_control TTL from 5m to 1h on Anthropic Messages requests. It does not change
OpenAI requests, so the OpenAI rows are unaffected by it.
OpenCode version
2.0.22 (data from 2.0.12 to 2.0.22)
Steps to reproduce
- Run a long session on the built-in
openai provider until auto compaction triggers, or run /compact on a
session with an Anthropic Messages model after more than one exchange.
- Compare
tokens.cache.read of the compaction message (session_message.type = 'compaction') with the previous
assistant message in the same session.
Operating System
Ubuntu 26.04.1 LTS
Terminal
Not applicable (OpenCode server)
Description
This is the new report requested in the closing comment of #43249 ("If you're seeing repeated cold compactions on
current V2, please open a new report with the exact version, provider/model, and input/cache-read counts").
On V2, summary compactions read almost nothing from the prompt cache. This also happens when the
previous model call in the same session ran less than a minute earlier and read more than 80% of its input from the
cache. On Anthropic Messages routes, about 96% of each compaction's input is billed as a new cache write.
Selection: compactions in the local
opencode.dbthat completed on 2.0.12 or later (version estimated from thesessions created just before), with at least 50k input tokens.
In each case, the previous model call in the session ran at most 5 minutes earlier and read at least 80% of its input
from the cache. Only 1 of these 132 compactions read at least 50% of its input from the cache.
openaiprovider (gpt-6-luna,gpt-6-sol,gpt-6-astra), autoExample rows (previous call input / previous call cache read → compaction input / cache read / cache write):
The constant read of about 21k tokens on OpenAI and about 8k tokens on Anthropic looks like the system prompt and tool
definitions only. The cache match appears to stop at the first conversation message.
Possible cause
In 2.0.22,
summarizesends only the older part of the history,split.older, and keeps the recent part as textoutside the request:
For Anthropic, this would explain the result. The request's cache breakpoints sit at the end of the older part. The
cache entries that the main loop keeps fresh end at the end of the full history, so no live entry matches the shorter
prefix. For OpenAI, I would expect the shorter prefix to still match, so I could not explain the miss there. The
request may differ from the main-loop request in the first conversation message, or in transport or cache-key
routing. I did not confirm either of these.
The closing comment on #43249 cites a
/compacttest on 2.0.12 that read about 99% from the cache. The one hit inour data (
github-copilot, 2.0.20) read 113,792 of 114,744 input tokens. Its input was about the same size as theprevious call's input (113,815), which suggests that the older part covered almost the full history there.
Possible direction
One option is to send the main loop's full current request with the summary instruction appended, as
session.generatedoes, and apply the older/recent split to the stored result instead of to the request. #46369proposes something in this direction as an opt-in. Related: #48542 (Anthropic compaction cache writes never read).
Plugins
A local plugin rewrites the
cache_controlTTL from 5m to 1h on Anthropic Messages requests. It does not changeOpenAI requests, so the OpenAI rows are unaffected by it.
OpenCode version
2.0.22 (data from 2.0.12 to 2.0.22)
Steps to reproduce
openaiprovider until auto compaction triggers, or run/compacton asession with an Anthropic Messages model after more than one exchange.
tokens.cache.readof the compaction message (session_message.type = 'compaction') with the previousassistant message in the same session.
Operating System
Ubuntu 26.04.1 LTS
Terminal
Not applicable (OpenCode server)