Summary
Auto-compaction can enter a runaway loop with the ollama-cloud provider. When the provider reports an implausible usage.input value (larger than the model's own context limit), OpenCode's compaction preflight treats it as real, compacts, and the next step re-reports the same inflated value — so it compacts again, indefinitely. In one session this fired 4 auto-compactions in 75 seconds ($23.41 in auto-compaction alone per the session's ledger, out of a $48.39 session total).
Environment
- opencode version: 2.0.12
- OS: Darwin 25.6.0 (macOS, arm64)
- Terminal: xterm-256color (TERM_PROGRAM=ghostty)
- Shell: /bin/zsh
- Install/channel: latest (installed at
~/.opencode/bin/opencode)
- Active plugins: None found in config (
~/.config/opencode/opencode.jsonc has keys $schema, mcp, providers, skills — no plugins entry, no .opencode/plugin/ directory)
- Relevant config:
compaction is unset, so defaults apply (auto: true, buffer: 20000, keep.tokens: 15000)
Reproduction
Not yet reduced to a minimal repro — the trigger appears to require the specific provider and a session that has accumulated a large image payload. What was observed:
- Use a session on provider
ollama-cloud, model deepseek-v4.1-flash (declared limit.context: 1048576, limit.output: 384000), variant max.
- Build up history including many full-size (1920x1080) screenshots read via the
read tool, until the provider rejects a request with HTTP 413 Request Entity Too Large (this provider rejects a request once accumulated image data exceeds roughly 4 MB of base64).
- Run a manual compaction to recover:
POST /api/session/{sessionID}/compact.
- Send any further prompt (here: the single word "continue").
- Observe: the next step reports an enormous
usage.input, triggers an automatic compaction, and immediately repeats.
Observed sequence (session ses_f3977a8caffeCoXYUH4TxHDGxG, all times local):
| seq |
message |
reason |
reported input tokens |
reported output tokens |
cost |
| 1194 |
compaction |
manual |
124,792,980 |
1,300,075 |
$19.4990 |
| 1211 |
assistant |
— |
4,055,630 |
58,650 |
$0.6436 |
| 1226 |
compaction |
auto |
25,391,744 |
2,041,207 |
$5.0335 |
| 1229 |
assistant |
— |
3,406,996 |
30,625 |
$0.5295 |
| 1248 |
compaction |
auto |
33,091,776 |
2,600,337 |
$6.5240 |
| 1251 |
assistant |
— |
4,221,155 |
49,767 |
$0.6631 |
| 1268 |
compaction |
auto |
33,401,627 |
2,802,525 |
$6.6918 |
| 1271 |
assistant |
— |
19,573,066 |
982,098 |
$3.5252 |
| 1292 |
compaction |
auto |
27,403,988 |
1,749,382 |
$5.1603 |
The four auto compactions at seq 1226/1248/1268/1292 occurred within 75 seconds (21:35:20, 21:35:40, 21:36:04, 21:36:35).
Expected Behavior
The compaction preflight should not schedule a checkpoint based on a usage report that is internally inconsistent with the model's own limits. A reported usage.input that is 24-32x the model's declared limit.context cannot be a real prompt size. Once a checkpoint is installed, the session should resume normally rather than immediately re-compacting.
Additionally, compaction.auto: true is documented to give one provider-overflow recovery attempt with a second overflow returned as an error — not an unbounded loop.
Actual Behavior
The loop repeats until the session is interrupted. Each cycle installs a fresh checkpoint, but the following model step again reports an inflated usage.input, which again exceeds the preflight ceiling, which schedules another auto-compaction.
The preflight ceiling with default settings and this model would be:
min(input limit - buffer, context limit - max(output reserve, buffer))
Even granting the full 1,048,576-token context as the input limit, the ceiling cannot exceed 1,048,576. A reported input of 25,391,744 exceeds it by ~24x, so every step re-triggers.
Why the reported usage cannot be genuine: the model's declared context limit is 1,048,576 tokens, and the entire stored session (all user + assistant messages, including every inline base64 screenshot) is ~5.1 MB — on the order of 1.3M tokens at the provider's own observed ratio of ~4 characters/token. There is no way for a request built from this session to contain 25-33M prompt tokens. The active context after the final compaction was 3 messages totaling 0.02 MB. The strongest form of this check, which does not depend on any estimate: a single model step reported usage.input = 25,391,744, which is 24x the model's context window and therefore physically impossible.
The derived cost is therefore unreliable too: because session_v2.cost is computed from these reported token counts, the $48.39 attributed to this session (89% of it from compaction) inherits the same problem. For balance: the provider's own usage endpoint reports a monthly figure of 0.187 for this account, which does not reconcile with OpenCode's $85.66 across 158 sessions. I could not resolve that discrepancy — a controlled request of ~100,000 prompt tokens incremented the provider's request_count (2795 to 2796) but left its cost figure unchanged, so the two measure different things or one is stale. I am not claiming the user was billed $48.39; I am claiming the token counts OpenCode acts on are provably impossible, and any cost derived from them should not be trusted or used to drive scheduling decisions.
Scope check: scanning every assistant step in this database, the only steps whose reported usage.input exceeds their own model's declared context limit are in this session (4 steps, worst 18.7x the limit). Other sessions on the same provider and model do not show it, so this is not a general per-provider accounting failure — it correlates with the screenshot-heavy history in this specific session. Note the reported values are also inconsistent with each other: compaction steps report 25-33M input tokens while the intervening plain assistant steps report ~3-4M, and the session's final active context was 0.02 MB.
Additional Context
- The session involved reading many 1920x1080 PNG screenshots through the
read tool, which the provider accepts inline as data:image/png;base64,... file parts. Screenshot data is the likely contributor to whatever produces the anomalous usage figure.
- Manual compaction on the same session initially appeared to work (active context reduced to 2 messages / 0.006 MB) and a subsequent trivial prompt ("Reply with only the word READY") returned correctly with
usage.input: 61198. The loop only began on the following prompt.
- Workaround applied by the user: stopping the session. The natural mitigation is
compaction: { auto: false } in config, which prevents the automatic path entirely.
- The provider is entered via
auth.json as provider id ollama-cloud, base URL https://ollama.com/v1, OpenAI-compatible (@opencode/ai/providers/openai-compatible). Ollama Cloud is not the standard Ollama local provider.
- Related prior symptom in the same session, possibly a separate issue: this provider returns HTTP 413
Request Entity Too Large (ref: ffee2105-58a8-49fa-beff-d946d3f11b20) once accumulated image payload in the request exceeds roughly 4 MB of base64 data, and the session becomes permanently unusable because every retry re-sends the same history. OpenCode surfaced this as provider.invalid-request / AI.Error.InvalidRequest. Worth noting that overflow-classified errors are supposed to trigger one compaction-and-retry recovery; if this 413 were classified as context overflow rather than invalid-request, the session may have self-recovered instead of hard-failing.
Relevant log lines (from ~/.local/share/opencode/log/opencode.log), for the original 413 failure:
timestamp=2026-09-22T01:04:40.704Z level=ERROR run=c6164560 message="Failed to drain Session" cause="AI.Error: Request Entity Too Large (ref: ffee2105-58a8-49fa-beff-d946d3f11b20) ..." role=server sessionID=ses_f3977a8caffeCoXYUH4TxHDGxG
No secrets are included in this report; session and message ids are included for traceability.
Summary
Auto-compaction can enter a runaway loop with the
ollama-cloudprovider. When the provider reports an implausibleusage.inputvalue (larger than the model's own context limit), OpenCode's compaction preflight treats it as real, compacts, and the next step re-reports the same inflated value — so it compacts again, indefinitely. In one session this fired 4 auto-compactions in 75 seconds ($23.41 in auto-compaction alone per the session's ledger, out of a $48.39 session total).Environment
~/.opencode/bin/opencode)~/.config/opencode/opencode.jsonchas keys$schema,mcp,providers,skills— nopluginsentry, no.opencode/plugin/directory)compactionis unset, so defaults apply (auto: true,buffer: 20000,keep.tokens: 15000)Reproduction
Not yet reduced to a minimal repro — the trigger appears to require the specific provider and a session that has accumulated a large image payload. What was observed:
ollama-cloud, modeldeepseek-v4.1-flash(declaredlimit.context: 1048576,limit.output: 384000), variantmax.readtool, until the provider rejects a request with HTTP 413Request Entity Too Large(this provider rejects a request once accumulated image data exceeds roughly 4 MB of base64).POST /api/session/{sessionID}/compact.usage.input, triggers an automatic compaction, and immediately repeats.Observed sequence (session
ses_f3977a8caffeCoXYUH4TxHDGxG, all times local):The four
autocompactions at seq 1226/1248/1268/1292 occurred within 75 seconds (21:35:20, 21:35:40, 21:36:04, 21:36:35).Expected Behavior
The compaction preflight should not schedule a checkpoint based on a usage report that is internally inconsistent with the model's own limits. A reported
usage.inputthat is 24-32x the model's declaredlimit.contextcannot be a real prompt size. Once a checkpoint is installed, the session should resume normally rather than immediately re-compacting.Additionally,
compaction.auto: trueis documented to give one provider-overflow recovery attempt with a second overflow returned as an error — not an unbounded loop.Actual Behavior
The loop repeats until the session is interrupted. Each cycle installs a fresh checkpoint, but the following model step again reports an inflated
usage.input, which again exceeds the preflight ceiling, which schedules another auto-compaction.The preflight ceiling with default settings and this model would be:
Even granting the full 1,048,576-token context as the input limit, the ceiling cannot exceed 1,048,576. A reported input of 25,391,744 exceeds it by ~24x, so every step re-triggers.
Why the reported usage cannot be genuine: the model's declared context limit is 1,048,576 tokens, and the entire stored session (all user + assistant messages, including every inline base64 screenshot) is ~5.1 MB — on the order of 1.3M tokens at the provider's own observed ratio of ~4 characters/token. There is no way for a request built from this session to contain 25-33M prompt tokens. The active context after the final compaction was 3 messages totaling 0.02 MB. The strongest form of this check, which does not depend on any estimate: a single model step reported
usage.input = 25,391,744, which is 24x the model's context window and therefore physically impossible.The derived cost is therefore unreliable too: because
session_v2.costis computed from these reported token counts, the $48.39 attributed to this session (89% of it from compaction) inherits the same problem. For balance: the provider's own usage endpoint reports a monthly figure of0.187for this account, which does not reconcile with OpenCode's $85.66 across 158 sessions. I could not resolve that discrepancy — a controlled request of ~100,000 prompt tokens incremented the provider'srequest_count(2795 to 2796) but left its cost figure unchanged, so the two measure different things or one is stale. I am not claiming the user was billed $48.39; I am claiming the token counts OpenCode acts on are provably impossible, and any cost derived from them should not be trusted or used to drive scheduling decisions.Scope check: scanning every assistant step in this database, the only steps whose reported
usage.inputexceeds their own model's declared context limit are in this session (4 steps, worst 18.7x the limit). Other sessions on the same provider and model do not show it, so this is not a general per-provider accounting failure — it correlates with the screenshot-heavy history in this specific session. Note the reported values are also inconsistent with each other: compaction steps report 25-33M input tokens while the intervening plain assistant steps report ~3-4M, and the session's final active context was 0.02 MB.Additional Context
readtool, which the provider accepts inline asdata:image/png;base64,...file parts. Screenshot data is the likely contributor to whatever produces the anomalous usage figure.usage.input: 61198. The loop only began on the following prompt.compaction: { auto: false }in config, which prevents the automatic path entirely.auth.jsonas provider idollama-cloud, base URLhttps://ollama.com/v1, OpenAI-compatible (@opencode/ai/providers/openai-compatible). Ollama Cloud is not the standard Ollama local provider.Request Entity Too Large(ref:ffee2105-58a8-49fa-beff-d946d3f11b20) once accumulated image payload in the request exceeds roughly 4 MB of base64 data, and the session becomes permanently unusable because every retry re-sends the same history. OpenCode surfaced this asprovider.invalid-request/AI.Error.InvalidRequest. Worth noting that overflow-classified errors are supposed to trigger one compaction-and-retry recovery; if this 413 were classified as context overflow rather than invalid-request, the session may have self-recovered instead of hard-failing.Relevant log lines (from
~/.local/share/opencode/log/opencode.log), for the original 413 failure:No secrets are included in this report; session and message ids are included for traceability.