⚠️ This project is currently unrecommended for use 😿
omp’s session transcript integrity is currently too unstable for cache warming to work reliably. Until omp supports transcript pinning, I don't recomend using it.
Keep your oh-my-pi sessions' prompt caches warm — resume hours later and pay cache-read prices, not cold re-reads.
Provider prompt caches are short-lived (Claude Code is 1h). Walk away from a 600k-token session for lunch and your next message re-reads the entire prefix at full price. This project fixes that with four cooperating pieces:
| Piece | Runs where | Job |
|---|---|---|
| Warmer daemon | launchd (macOS) / systemd user unit (Linux) | Pings idle sessions on a per-provider schedule so live caches never expire |
| prefix-pin extension | every omp process | Freezes each session's rendered system prompt so the prefix never drifts; upgrades Anthropic sessions to the 1 h cache tier |
| miss-guard extension | interactive omp sessions | Status-bar 🔥/❄ indicator; predicts cold and idle-flush misses and asks before you pay |
| live-warm extension | open omp sessions (armed via /livewarm-on) |
Self-warms the live process: real ping turn + rewind, dodging omp's 90 m idle-flush |
git clone https://github.com/blingdivinity/omp-cache-warmer
omp plugin link ./omp-cache-warmer # extensions: pinning + miss-guard + /warm-* commands
cd omp-cache-warmer && bun src/index.ts install # daemon: launchd on macOS, systemd --user on LinuxThe Linux installer also enables user lingering (or tells you the sudo loginctl enable-linger command) so the daemon runs without an SSH session.
The daemon scans ~/.omp/agent/sessions every minute for sessions whose last user message is within the warm window (default 24 h). When a session has been idle longer than its provider's interval, it re-sends the session's exact prompt prefix with a trivial Respond with only: OK message, refreshing the provider-side cache TTL.
Your session history never grows. Each warm resumes a temporary copy of the session file with a throwaway --session-dir — same system prompt, same tools, same history, same model → byte-identical prefix → cache refreshed, real .jsonl untouched.
Every warm is verified. The response's usage.cacheRead is checked: ~100 % confirms the prefix still matches. Outcomes land in an append-only ledger (history.jsonl); stats summarizes hits, drifts, and token totals.
| Provider | Interval | Why |
|---|---|---|
anthropic |
55 min | 1 h cache TTL, refreshed with 5 min safety margin |
openai-codex |
8 h 01 m | fires twice inside the 24 h window (8h01m, 16h02m) |
| everything else | 55 min (defaultIntervalMinutes) |
configurable per provider prefix |
There is no free way to ask a provider "does this cache still exist?" — cache status only comes back in the usage of a paid request. The warmer predicts expiry from TTL math (it knows exactly when each cache was last touched). With the default coldReprime: "never", a cache that is predicted expired is left cold: the session is disabled with a logged reason and zero spend. Cold prefixes are only re-primed deliberately —
- you send a real message (guarded by miss-guard if it's big), or
- you run
omp-cache-warmer warm <id>, which re-primes and re-enrolls the session.
Set coldReprime to a token count (re-prime automatically when the prefix is at most that big) or "always" if you want reviver behavior.
Warming is pointless if the prefix changes underneath you. omp re-renders the system prompt every turn — the date line, directory state, and extension changes all shift it, and every shift cold-misses the whole prefix.
extensions/prefix-pin.ts freezes it:
- A session's first turn captures the fully rendered system prompt to
~/.omp/agent/omp-cache-warmer/pins/<session-id>.json. - Every later turn replays it byte-for-byte — interactive resumes and warmer pings, which load the same extension. Date rollover, renamed directories, prompt tweaks: none of it moves the prefix anymore.
- Pins never expire. Even a cold-cached session resumes onto its original prefix, so the re-prime you pay is for the same stable lineage, not a freshly drifted one. Pins unused for 60 days are pruned (a later resume simply re-pins).
The extension also sets PI_CACHE_RETENTION=long in-process (unless you set it yourself), upgrading Anthropic sessions from 5-minute to 1-hour cache entries — without which the 55 min warm interval could never work.
In-session: /pin-status, /pin-refresh (drop the pin after intentionally changing your setup).
Limits: tool schemas are the one prefix component an extension cannot pin; they change on omp version bumps and cost one re-prime each. See docs/UPSTREAM-PIN-HANDOFF.md for the upstream prefix_snapshot proposal that closes this gap.
extensions/miss-guard.ts runs the same TTL math as the daemon on every message you submit. If the cache is predicted expired and the prefix is ≥ 40k tokens (missConfirmTokens), you get a yes/no dialog before anything is sent:
Predicted prompt-cache MISS — this session has been idle ~87m; sending will re-read ~85,541 tokens uncached. Send anyway?
Declining restores your typed message to the editor. Predicted-warm sends and small cold sends pass through silently.
omp's prune/supersede pass rewrites a live process's sent history once it has been idle >90 min (PRUNE_IDLE_FLUSH_MS), assuming the provider cache is cold — an assumption an external warmer breaks. Consequence: the daemon protects resumes, not open-idle processes. Typing into an omp window that sat idle >90 min misses regardless of warming (miss-guard flags this as (idle-flush) and points you to omp -c, which re-renders from the file and does hit the warmed prefix).
extensions/live-warm.ts closes the gap from inside: on the provider interval, an idle armed session sends a real Respond with only: OK turn through its own pipeline (byte-identical prefix by construction — measured 100 % cache hits) and then rewinds the session tree to the pre-ping leaf (measured ~99 % cached on the next real message). Each ping also resets the idle-flush clock, so the flush never fires.
- Enable with
"liveWarm": true, arm per session with/livewarm-on(a command context is required for rewind; upstream would need to widen the event context to make this automatic)./livewarm-pingruns one cycle manually. - Hard cutoff at 88 min idle: if the process slept past the window, pinging would trigger the flush, so it refuses and defers to miss-guard.
- Skips when the agent is busy or messages are pending; logs to
live-warm.log; coordinates with the daemon via shared state.
Warm pings rebuild the prompt through the same code path a real resume uses, so drift is detected and self-heals:
- One-off drift (omp updated, extension changed the prompt): one warm misses, and by missing re-primes the new prefix — the exact one your next resume sends. Logged as
prefix DRIFT, shown asdrifts=Ninstatus. - Constant drift (prompt embeds churning content): two consecutive drifts auto-disable the session — correct, because a real resume would miss anyway and money shouldn't be spent pretending otherwise.
- With pinning active, both cases essentially disappear.
omp-cache-warmer status sessions in the warm window, schedule, hit stats (default)
omp-cache-warmer once single sweep
omp-cache-warmer daemon run forever (sweeps every minute)
omp-cache-warmer warm <id> force-warm one session now (re-primes + re-enables)
omp-cache-warmer enable <id> re-enable a disabled session
omp-cache-warmer stats hit/miss ledger summary
omp-cache-warmer install install launchd agent (macOS) / systemd user unit (Linux)
omp-cache-warmer uninstall remove it
omp-cache-warmer config show config path + contents
(Run via bun src/index.ts <cmd> from a checkout, or the omp-cache-warmer bin if installed.)
The background service is named so future-you knows what it is: com.oh-my-pi.keep-ai-prompt-cache-warm (launchd) / omp-keep-ai-prompt-cache-warm.service (systemd).
~/.omp/agent/omp-cache-warmer/config.json (created on first run, re-read every sweep and every miss-guard check — edits apply live):
Sibling files: state.json (per-session warm state), history.jsonl (ledger), pins/ (frozen prompts), warmer.log / warmer.launchd.log (logs; journalctl --user -u omp-keep-ai-prompt-cache-warm on Linux).
In-session commands: /warm-status, /warm-off, /pin-status, /pin-refresh.
- A scheduled warm = one cache read of the prefix (0.1× input price on Anthropic) + a few output tokens. Warming a 24k-token session hourly for a day costs roughly one uncached request over it.
- A drift or forced re-prime = one cache write (1.25×/2× input price) — visible in
statsascacheWrite. - An unguarded cold resume of a big session = the thing this project exists to prevent.
- Active sessions warm themselves — the daemon watches file mtime and stays out of the way until you've been idle a full interval.
- Drift between the last warm and your resume (you rename files, then resume a minute later, without pinning) cannot be pre-warmed by anything: the new prefix doesn't exist until you send it.
- Sessions started before the plugin was installed are unprotected until restarted (
omp -c): they render unpinned prompts on the 5-minute tier while the warmer faithfully keeps a lineage they no longer send. Restarting aligns them with the pin.
MIT
{ "windowHours": 24, // keep warming this long after the last user message "windowHoursByProvider": {}, // per-provider window override (hours) "intervals": { // minutes, keyed by provider prefix of the model id "anthropic": 55, "openai-codex": 481 }, "defaultIntervalMinutes": 55, "message": "Respond with only: OK", "perProject": "latest", // "latest" = newest session per project dir, or "all" "maxWarmsPerSweep": 4, "warmTimeoutSeconds": 300, "coldReprime": "never", // "never" | "always" | max tokens to auto-revive a cold cache "missConfirmTokens": 40000, // miss-guard dialog threshold; false disables the dialog "pinPrefixes": true, // prefix pinning on/off "exclude": [], // session id prefixes to never warm "ompBin": "omp" }