Skip to content

V2: summary compaction still reads almost nothing from the prompt cache, even right after a warm request #52761

Description

@sergiopadure

Description

This is the new report requested in the closing comment of #43249 ("If you're seeing repeated cold compactions on
current V2, please open a new report with the exact version, provider/model, and input/cache-read counts").

On V2, summary compactions read almost nothing from the prompt cache. This also happens when the
previous model call in the same session ran less than a minute earlier and read more than 80% of its input from the
cache. On Anthropic Messages routes, about 96% of each compaction's input is billed as a new cache write.

Selection: compactions in the local opencode.db that completed on 2.0.12 or later (version estimated from the
sessions created just before), with at least 50k input tokens.
In each case, the previous model call in the session ran at most 5 minutes earlier and read at least 80% of its input
from the cache. Only 1 of these 132 compactions read at least 50% of its input from the cache.

Route Compactions Versions Median input Median cache read Read share Write share Median gap since previous call
Built-in openai provider (gpt-6-luna, gpt-6-sol, gpt-6-astra), auto 113 2.0.16–2.0.21 253,321 20,992 7.4% 0.0% 0.68 min
Anthropic Messages (Anthropic API, Vertex AI, Azure AI Foundry; Claude models), mostly manual 12 2.0.14–2.0.21 407,080 8,051 4.2% 95.8% 0.76 min

Example rows (previous call input / previous call cache read → compaction input / cache read / cache write):

2.0.21 openai gpt-6-luna auto    gap 0.54 min: 784,645 / 776,704 → 487,401 / 4,608 / 0
2.0.19 anthropic claude-sonnet-5-5 auto gap 0.31 min: 903,250 / 883,134 → 641,900 / 8,051 / 633,845
2.0.16 google-vertex claude-opus-5-5 manual gap 1.27 min: 837,966 / 836,224 → 672,165 / 0 / 672,161
2.0.22 google-vertex gemini-3.8-flash auto gap 0.57 min: 951,296 / 899,938 → 421,711 / 0 / 0

The constant read of about 21k tokens on OpenAI and about 8k tokens on Anthropic looks like the system prompt and tool
definitions only. The cache match appears to stop at the first conversation message.

Possible cause

In 2.0.22, summarize sends only the older part of the history, split.older, and keeps the recent part as text
outside the request:

For Anthropic, this would explain the result. The request's cache breakpoints sit at the end of the older part. The
cache entries that the main loop keeps fresh end at the end of the full history, so no live entry matches the shorter
prefix. For OpenAI, I would expect the shorter prefix to still match, so I could not explain the miss there. The
request may differ from the main-loop request in the first conversation message, or in transport or cache-key
routing. I did not confirm either of these.

The closing comment on #43249 cites a /compact test on 2.0.12 that read about 99% from the cache. The one hit in
our data (github-copilot, 2.0.20) read 113,792 of 114,744 input tokens. Its input was about the same size as the
previous call's input (113,815), which suggests that the older part covered almost the full history there.

Possible direction

One option is to send the main loop's full current request with the summary instruction appended, as
session.generate does, and apply the older/recent split to the stored result instead of to the request. #46369
proposes something in this direction as an opt-in. Related: #48542 (Anthropic compaction cache writes never read).

Plugins

A local plugin rewrites the cache_control TTL from 5m to 1h on Anthropic Messages requests. It does not change
OpenAI requests, so the OpenAI rows are unaffected by it.

OpenCode version

2.0.22 (data from 2.0.12 to 2.0.22)

Steps to reproduce

  1. Run a long session on the built-in openai provider until auto compaction triggers, or run /compact on a
    session with an Anthropic Messages model after more than one exchange.
  2. Compare tokens.cache.read of the compaction message (session_message.type = 'compaction') with the previous
    assistant message in the same session.

Operating System

Ubuntu 26.04.1 LTS

Terminal

Not applicable (OpenCode server)

No activity

Activity on this issue will appear here.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

No labels
No labels

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions