Skip to content

compaction: auto-compaction loops when provider reports input tokens above the model's context limit #50474

Description

@thomasvdv

Summary

Auto-compaction can enter a runaway loop with the ollama-cloud provider. When the provider reports an implausible usage.input value (larger than the model's own context limit), OpenCode's compaction preflight treats it as real, compacts, and the next step re-reports the same inflated value — so it compacts again, indefinitely. In one session this fired 4 auto-compactions in 75 seconds ($23.41 in auto-compaction alone per the session's ledger, out of a $48.39 session total).

Environment

  • opencode version: 2.0.12
  • OS: Darwin 25.6.0 (macOS, arm64)
  • Terminal: xterm-256color (TERM_PROGRAM=ghostty)
  • Shell: /bin/zsh
  • Install/channel: latest (installed at ~/.opencode/bin/opencode)
  • Active plugins: None found in config (~/.config/opencode/opencode.jsonc has keys $schema, mcp, providers, skills — no plugins entry, no .opencode/plugin/ directory)
  • Relevant config: compaction is unset, so defaults apply (auto: true, buffer: 20000, keep.tokens: 15000)

Reproduction

Not yet reduced to a minimal repro — the trigger appears to require the specific provider and a session that has accumulated a large image payload. What was observed:

  1. Use a session on provider ollama-cloud, model deepseek-v4.1-flash (declared limit.context: 1048576, limit.output: 384000), variant max.
  2. Build up history including many full-size (1920x1080) screenshots read via the read tool, until the provider rejects a request with HTTP 413 Request Entity Too Large (this provider rejects a request once accumulated image data exceeds roughly 4 MB of base64).
  3. Run a manual compaction to recover: POST /api/session/{sessionID}/compact.
  4. Send any further prompt (here: the single word "continue").
  5. Observe: the next step reports an enormous usage.input, triggers an automatic compaction, and immediately repeats.

Observed sequence (session ses_f3977a8caffeCoXYUH4TxHDGxG, all times local):

seq message reason reported input tokens reported output tokens cost
1194 compaction manual 124,792,980 1,300,075 $19.4990
1211 assistant — 4,055,630 58,650 $0.6436
1226 compaction auto 25,391,744 2,041,207 $5.0335
1229 assistant — 3,406,996 30,625 $0.5295
1248 compaction auto 33,091,776 2,600,337 $6.5240
1251 assistant — 4,221,155 49,767 $0.6631
1268 compaction auto 33,401,627 2,802,525 $6.6918
1271 assistant — 19,573,066 982,098 $3.5252
1292 compaction auto 27,403,988 1,749,382 $5.1603

The four auto compactions at seq 1226/1248/1268/1292 occurred within 75 seconds (21:35:20, 21:35:40, 21:36:04, 21:36:35).

Expected Behavior

The compaction preflight should not schedule a checkpoint based on a usage report that is internally inconsistent with the model's own limits. A reported usage.input that is 24-32x the model's declared limit.context cannot be a real prompt size. Once a checkpoint is installed, the session should resume normally rather than immediately re-compacting.

Additionally, compaction.auto: true is documented to give one provider-overflow recovery attempt with a second overflow returned as an error — not an unbounded loop.

Actual Behavior

The loop repeats until the session is interrupted. Each cycle installs a fresh checkpoint, but the following model step again reports an inflated usage.input, which again exceeds the preflight ceiling, which schedules another auto-compaction.

The preflight ceiling with default settings and this model would be:

min(input limit - buffer, context limit - max(output reserve, buffer))

Even granting the full 1,048,576-token context as the input limit, the ceiling cannot exceed 1,048,576. A reported input of 25,391,744 exceeds it by ~24x, so every step re-triggers.

Why the reported usage cannot be genuine: the model's declared context limit is 1,048,576 tokens, and the entire stored session (all user + assistant messages, including every inline base64 screenshot) is ~5.1 MB — on the order of 1.3M tokens at the provider's own observed ratio of ~4 characters/token. There is no way for a request built from this session to contain 25-33M prompt tokens. The active context after the final compaction was 3 messages totaling 0.02 MB. The strongest form of this check, which does not depend on any estimate: a single model step reported usage.input = 25,391,744, which is 24x the model's context window and therefore physically impossible.

The derived cost is therefore unreliable too: because session_v2.cost is computed from these reported token counts, the $48.39 attributed to this session (89% of it from compaction) inherits the same problem. For balance: the provider's own usage endpoint reports a monthly figure of 0.187 for this account, which does not reconcile with OpenCode's $85.66 across 158 sessions. I could not resolve that discrepancy — a controlled request of ~100,000 prompt tokens incremented the provider's request_count (2795 to 2796) but left its cost figure unchanged, so the two measure different things or one is stale. I am not claiming the user was billed $48.39; I am claiming the token counts OpenCode acts on are provably impossible, and any cost derived from them should not be trusted or used to drive scheduling decisions.

Scope check: scanning every assistant step in this database, the only steps whose reported usage.input exceeds their own model's declared context limit are in this session (4 steps, worst 18.7x the limit). Other sessions on the same provider and model do not show it, so this is not a general per-provider accounting failure — it correlates with the screenshot-heavy history in this specific session. Note the reported values are also inconsistent with each other: compaction steps report 25-33M input tokens while the intervening plain assistant steps report ~3-4M, and the session's final active context was 0.02 MB.

Additional Context

  • The session involved reading many 1920x1080 PNG screenshots through the read tool, which the provider accepts inline as data:image/png;base64,... file parts. Screenshot data is the likely contributor to whatever produces the anomalous usage figure.
  • Manual compaction on the same session initially appeared to work (active context reduced to 2 messages / 0.006 MB) and a subsequent trivial prompt ("Reply with only the word READY") returned correctly with usage.input: 61198. The loop only began on the following prompt.
  • Workaround applied by the user: stopping the session. The natural mitigation is compaction: { auto: false } in config, which prevents the automatic path entirely.
  • The provider is entered via auth.json as provider id ollama-cloud, base URL https://ollama.com/v1, OpenAI-compatible (@opencode/ai/providers/openai-compatible). Ollama Cloud is not the standard Ollama local provider.
  • Related prior symptom in the same session, possibly a separate issue: this provider returns HTTP 413 Request Entity Too Large (ref: ffee2105-58a8-49fa-beff-d946d3f11b20) once accumulated image payload in the request exceeds roughly 4 MB of base64 data, and the session becomes permanently unusable because every retry re-sends the same history. OpenCode surfaced this as provider.invalid-request / AI.Error.InvalidRequest. Worth noting that overflow-classified errors are supposed to trigger one compaction-and-retry recovery; if this 413 were classified as context overflow rather than invalid-request, the session may have self-recovered instead of hard-failing.

Relevant log lines (from ~/.local/share/opencode/log/opencode.log), for the original 413 failure:

timestamp=2026-09-22T01:04:40.704Z level=ERROR run=c6164560 message="Failed to drain Session" cause="AI.Error: Request Entity Too Large (ref: ffee2105-58a8-49fa-beff-d946d3f11b20) ..." role=server sessionID=ses_f3977a8caffeCoXYUH4TxHDGxG

No secrets are included in this report; session and message ids are included for traceability.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

No labels
No labels

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions