Skip to content

perf: port upstream #505 #471 #367 and make KV checkpoint parking fast - #12

Merged
dejay2 merged 12 commits into
mtp-upstream-mergefrom
perf/park-fast-copy
Sep 23, 2026
Merged

dejay2 merged 12 commits into
mtp-upstream-mergefrom
perf/park-fast-copy

Conversation

@dejay2

@dejay2 dejay2 commented Sep 23, 2026

Copy link
Copy Markdown
Owner

Stacked on #11. Makes KV-parking RAM mode fast enough for the per-prompt checkpoint saves that FlashML-org#505 switched on.

Why

On the 5090 (2026-09-23) a 50k-token checkpoint (742 MiB) saved at 25-50 MB/s: 8-11 s inline, and up to ~30 s under Windows RAM pressure. Causes:

  • ~47k views and one tiny cudaMemcpyAsync per view. About 1 s of Python and dispatch per save before any copying.
  • A cudaHostAlloc of the full entry on every save. About 1.2 s, and an implicit device sync; eviction also paid cudaFreeHost syncs.
  • Continuation saves staged the whole borrowed prefix back to the host to compare it with the parent.

Restore had the same per-view cost.

What

  • QSAKVCache.page_byte_regions(): the pool's per-page byte tables, zero-copy, in page_byte_views order.
  • RAM save: gathers whole pages through a reusable 2×32 MiB device staging buffer, with one D2H per chunk into pageable host memory. No per-save pinning.
  • Prefix check: runs on the device.
  • Restore: H2D per chunk, then a scatter into the pool pages.
  • Unchanged: the byte format, so the SSD format, offsets and parent chains are identical. SSD mode is untouched. A pool whose regions don't match falls back to the old path.
  • Guard: out-of-range page ids raise before the device gather.

Test plan

  • Devbox (CPU): tests/scheduler tests/kvcache 1017 passed, 10 skipped (baseline 976/8; +41 new tests: byte identity with the old serialization, RAM equal to the SSD file, chunk boundaries, one-byte prefix mismatch, no pinned allocation, fallback)
  • CPU microbenchmark (scripts/bench_park_ram_copy.py, 782 pages, 743 MiB): root save 1,202 → 251 ms, continuation 1,017 → 289 ms, restore 739 → 286-453 ms. This is CPU only, not GPU/WSL speed.
  • Read-only review: no blocking bugs. Its two follow-ups (the out-of-range guard, and a test flake that also hit the parent commit) are fixed.
  • 5090: CUDA tests (test_cuda_ram_save_is_pageable_and_byte_identical and the park CUDA tests), then the live 50k-token A/B with RAM parking

🤖 Generated with Claude Code

dejay2 and others added 12 commits September 23, 2026 11:57
Port of upstream FlashML-org#505 (eamars). A continuation chunk
did not carry mamba_last_track_seqlen, so a final chunk of <= 64 tokens
(too short to write its own snapshot) dropped the previous one. The final
prefill commit then returned early in cache.py, skipping the donate and our
eager prompt-checkpoint save that KV parking relies on, so the next turn
recomputed the whole prompt.

The new parametrized test fails 10/16 cases without the prefill.py change
and passes 16/16 with it (devbox, CPU).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…sync

Port of upstream FlashML-org#500 (alvarorsouza-arch, with
chrisqianz's CPU fallback, bounds guard and parity test). The boolean-mask
index in _invalidate_prefill_buffer has a data-dependent shape and forced a
device-to-host sync twice per chunk per layer, stalling the prefill-overlap
pipeline (on by default here, engine/config.py moe_prefill_overlap).

Unmeasured on the RTX 5090 so far; upstream's synthetic run showed no
wall-time change, the production A/B reported shorter long-context turns.
The CUDA parity test skips on the GPU-less devbox.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Port of upstream FlashML-org#471 (cherry77-cloud). temperature=0
or top_k=1 is now argmax whatever top_p says, and a greedy row batched with
a sampled one (two chats at once) takes argmax instead of sampling at
T=1e-6.

Fork follow-through: spec_filter_params reads the new greedy_mask so the MTP
draft filter agrees with the server for a greedy row in a mixed batch, the
request_filter_params docstring drops the old "temperature 0 + top_p is
sampled" rule, and test_spec_draft_graph pins the new triple. Devbox
tests/engine: 7 failed / 749 passed after vs 8 / 748 before; the remaining
7 fail identically on the base commit (no flashinfer / spec_draft env).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Port of upstream FlashML-org#367 (cherry77-cloud), adapted to the
dynamic KV pool. allocate_paged gives each request whole 64-token pages, but
admission charged raw tokens, so with two running requests the reservation
could fall up to ~126 tokens short and overcommit a nearly full pool.

kv_reservation_tokens() in scheduler/cache.py is the one page-rounded cost,
used by PrefillAdder's never-fits gate, its admission_fits checks, the
reserved_size charge, the preparation carry-over reservation, and the
dynamic pool's probe_admission need_now, so the probe and the real
admission keep agreeing (upstream's patch predates the probe). Devbox:
tests/scheduler 558 passed.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The final-prefill checkpoint save ran inline, by design, so it held the
first answer token while it copied. It rarely ran before upstream FlashML-org#505:
the chunk continuation dropped mamba_last_track_seqlen and the commit
returned early. With FlashML-org#505 it runs on every long prompt. On the 5090
(2026-09-23, 50k-token prompt, RAM parking, 2 GiB budget) that took cold
TTFT from 18-25 s to 28-44 s, with an 8-11 s stall before the last chunk.
py-spy put ~3.4 s of scheduler time in the save: 1.2 s pinned allocation,
~2 s per-page copy slicing, ~1 s copy.

_save_prompt_checkpoint now queues the save with ParkStore.offer. The
checkpoint node takes one extra lock until the worker's D2H copy
completes, so its canonical pages and frozen tree slot cannot be evicted
or reissued. drain_pending_parks drops only that lock; nothing is freed
and the node stays in the tree. A full queue or a store without a worker
falls back to the synchronous save.

Tests: checkpoint tests settle the worker before inspecting the store.
The slot-pressure test accepts "saved or held by a copy lock". A new
gated-worker test shows the commit returns before the copy and holds the
node until it lands; it fails on the inline version. Devbox:
tests/scheduler + tests/kvcache 975 passed.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
… kernel

Measured on the 5090 on 2026-09-23 (same 50k-token prompt A/B, both with
FlashML-org#505): cold TTFT was 29-32 s with the kernel and 28-30 s without it, and
prefill chunks took ~3.2 s with it and ~2.8 s without. The hidden sync it
removes is mostly the host waiting for prefill GPU work that has to finish
anyway (py-spy: that time is the GPU, not the sync). Upstream's synthetic
run also showed no wall-time change, so the extra Triton kernel is not
worth carrying.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Live on the 5090 on 2026-09-23 (with the background checkpoint save): a
governor MoE-only cache step re-captured the decode graphs while the
park worker was copying a prompt checkpoint. The global-mode capture was
invalidated (cudaErrorStreamCaptureInvalidated), the worker's save failed
("operation not permitted when stream is capturing"), and the scheduler
latched failed with the server at 503.

- engine/graph.py captures in thread_local mode, as the three MTP
  captures already do for background threads, so the worker's pinned
  allocation and D2H copies on its own stream cannot invalidate it.
- rebuild_cache waits for queued park copies before a MoE-only step too.
  Only KV and mamba resizes went through prepare_rebuild's flush, so
  teardown and re-capture never overlap the worker.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…waits

Review of the background checkpoint save found two follow-ups:
- The MTP shadow verifier borrows an inactive GDN slot under full slot
  pressure, and a queued checkpoint's tree slot looked inactive, so a
  concurrent copy could park a corrupted state. It now skips
  CacheManager.pending_park_slots(). The shadow is off on the box
  (FREETOKEN_MTP_SHADOW=0).
- ensure_mamba_slots and _allocate blocked on any pending park, but a
  checkpoint hold releases only a lock. They now wait only when a pending
  park returns pages or slots.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…AM mode

A 782-page (50k-token, 742 MiB) prompt checkpoint parked to RAM at 25-50 MB/s on
the 5090 under WSL2: ~47k per-view D2H copies, a cudaHostAlloc of the whole entry
on every save (~1.2 s, up to ~30 s under memory pressure), and a second full D2H
pass through a fresh pinned window for the continuation prefix check.

- QSAKVCache.page_byte_regions exposes the storage behind page_byte_views as
  [num_pages, bytes] uint8 views in the same order, so one index_select per region
  per chunk rebuilds the exact page-major bytes (SSD format and offsets unchanged).
- RAM saves gather <= 32 MiB of whole pages into a reusable 2 x 32 MiB device
  staging buffer and move each chunk with one copy into ordinary pageable memory;
  nothing is pinned per save and eviction no longer pays cudaFreeHost.
- The continuation prefix check brings the parent's bytes H2D chunk by chunk and
  compares on the device; restore stages H2D and scatters with index_copy_.
- Pools whose regions do not line up with page_byte_views keep the per-view path.
  SSD mode is untouched.

CPU microbench (scripts/bench_park_ram_copy.py, 782 pages): root save 1,202 ->
251 ms, continuation 1,017 -> 289 ms, restore 739 -> 453 ms.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
- Byte identity of the region gather against the per-view serialization for
  shuffled pages (BF16 and FP8), and RAM buffer == SSD file regions.
- Chunking: sub-page, fractional, exact and larger-than-entry chunks stay on page
  boundaries; chained restores across chunks with page offsets.
- Continuation prefix check: a match links, one flipped byte in a K/V, index or
  scale region of a borrowed page falls back to a root.
- No pinned allocation in RAM mode; mismatched regions fall back to per-view.
- test_idle_threshold_waits_until_the_leaf_is_old_enough held no copy open while
  asserting the copy was in flight: 1/30 failures before, 3/30 with the faster
  copy, 0/30 now that the copy is held until the check runs.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
… a racy test

Review follow-ups for the page-major RAM copy:
- _page_source raises IndexError for a page id outside the pool before
  index_select/index_copy_. On CUDA an out-of-range index is a device-side
  assert that poisons the context; the old per-view path raised IndexError
  and only disabled parking.
- The twelve-slot checkpoint test asserted "saved or queued" in the window
  after copy_done (hold drained) and before _publish. That state is safe
  because the copy has finished. The test now flushes the store before
  failing. The flake also hit the parent commit (3/64 under load); 20/20
  now.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
On the 5090 the test's FakeStream reached the new page-major source
(torch.cuda.stream needs .device) before the patched _copy_to_ram, so the
recorded error was the fake's AttributeError. Inject the failure at
_page_source too; the fence-before-release assertion is unchanged.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@dejay2

dejay2 commented Sep 23, 2026

Copy link
Copy Markdown
Owner Author

5090 results (2026-09-23), live on the box with RAM parking:

  • GPU tests: park tests 202/202 and fast-copy tests 43/43, CUDA path included. One test double (FakeStream) needed updating (769b7f1).

  • A/B, 4 conversations of 50k tokens each, no restart:

    This branch Base
    Cold TTFT 16.0-16.2 s 18-25 s
    Follow-up 2.3-2.4 s 3.6-3.8 s

    The first rep's cold TTFT (26 s) included a KV pool grow.

  • Memory: Windows free RAM stayed at 15 GB or more, with zero page-outs. No errors.

  • The PC freezes seen during earlier testing came from repeated server restarts, not from parking. Each boot takes WSL from 20 to 74 GB and leaves Windows about 1.1 GB free for around 30 s.

🤖 Generated with Claude Code

@dejay2 dejay2 changed the title perf: copy parked KV pages to RAM in a few large transfers perf: port upstream #505 #471 #367 and make KV checkpoint parking fast Sep 23, 2026
@dejay2
dejay2 changed the base branch from perf/upstream-picks-2026-09-23 to mtp-upstream-merge September 23, 2026 17:10
@dejay2
dejay2 merged commit 18c8b78 into mtp-upstream-merge Sep 23, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant