Skip to content

compaction: auto-compaction never triggers on hy3 — token count plateaus below the isOverflow threshold, context pinned at 196608 raw input tokens #45168

Description

@reisi007

Summary

When using model hy3 (provider opencode-go), the session context grows until it is pinned at exactly 196608 tokens, but auto-compaction never triggers. Every subsequent request re-sends the full ~196k context as raw (uncached) input tokens, causing a massive silent cost blowup. The user only notices via billing; there is no error and no warning.

Environment

  • opencode version: opencode2 v0.0.0-beta-18269
  • OS: Darwin 25.6.0 (darwin arm64) — macOS (Apple Silicon)
  • Terminal: Apple_Terminal (TERM=xterm-256color, COLORTERM=truecolor)
  • Shell: /bin/zsh
  • Install/channel: beta
  • Active plugins: none found in config

Reproduction

  1. Configure/use model hy3 (providerID: opencode-go, variant high) in a session.
  2. Work in one long session (orchestrator-style with many subagent/tool round-trips) so the context grows steadily.
  3. Observe token usage per assistant message: tokens.cache.read climbs normally (32k → 188k) until the context reaches 196608.
  4. From that point on, every assistant message reports tokens.input = 196608, cache.read = 0.

Evidence from an affected session (ses_fc64ce80cffeTu4Jbu2dtS3r4K, local SQLite storage):

  • 126 assistant messages on hy3, ~16.98M billed input tokens total
  • 24 consecutive-ish messages with input=196608 / cache_read=0 — i.e., full context re-billed raw each request:
    1787678464966  input=193668  cache_read=0
    1787678660441  input=196608  cache_read=0   <-- pinned at exactly 196608
    1787679843733  input=196608  cache_read=0
    1787681534348  input=196608  cache_read=0
    ... (24 such requests over ~2 hours)
    
  • No compaction event was recorded during that window. The only compaction in the session has "reason":"manual" and happened after manually switching back to another model (ox-alpha-free). Config has compaction.prune: true.

Expected Behavior

Auto-compaction should trigger as the context approaches the model's context limit (well before hitting it), summarizing/pruning the conversation so subsequent requests stay within budget. Requests should not silently resend the maximum context as raw uncached input indefinitely.

Actual Behavior

  • Auto-compaction never fires for hy3.
  • Once the context hits 196608 tokens (exactly 262144 − 65536, i.e., the provider's real input ceiling), every request sends all 196608 tokens as raw input with zero cache hits.
  • There is no error message; the failure mode is purely financial (silent cost blowup).

Root Cause (verified against source + models.dev)

Auto-compaction is gated by isOverflow() in packages/opencode/src/session/overflow.ts:

usable = model.limit.context − min(COMPACTION_BUFFER /* 20_000 */, maxOutputTokens)
count  = tokens.total || input + output + cache.read + cache.write
return count >= usable

For opencode-go/hy3, models.dev reports "limit": { "context": 256000, "output": 64000 }:

  • Compaction threshold = 256000 − min(20000, 64000) = 236,000 tokens
  • The provider's real constraint is binary-sized: context 262144 (256×1024), output 65,536 (64×1024), so input is capped at exactly 262144 − 65536 = 196,608 — matching the observed pin.
  • In agentic workloads the model generates small outputs (observed: 36–9,248 tokens/turn), so count plateaus at ≈ 196,608 + ~1,500 ≈ 198,100 tokens.

Result: count can never reach the 236,000 threshold (gap ≈ 38k tokens), so isOverflow() never returns true and auto-compaction never fires. The threshold assumes output tokens consume their full reserved budget toward the limit; when actual outputs are small relative to maxOutputTokens, any model whose real input cap + realistic output stays below advertised_context − buffer becomes permanently immune to auto-compaction.

Secondary effect: once pinned, prompt caching stopped working entirely (cache_read=0 on every request), so each request pays full raw-input price ($0.0175/M) instead of cached-read price ($0.004375/M).

Possible directions:

  • Base the overflow check on the effective input ceiling (limit.input when available, or a fraction of context that accounts for realistic output), rather than context − fixed_buffer.
  • React to observed ceilings: repeated requests pinned at a maxed-out input size indicate the true limit has been reached even if count < usable.

Additional Context

  • The value 196608 appears as an exact ceiling in 24 separate assistant messages spanning ~2 hours — a fixed server-side cap (262144 − 65536), not random truncation.
  • Frequency: consistently reproducible whenever a single session on hy3 grows past ~190k tokens.
  • Workaround: manually run /compact or switch models before reaching the limit.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions