perf: port upstream #505 #471 #367 and make KV checkpoint parking fast - #12
Merged
Merged
Conversation
Port of upstream FlashML-org#505 (eamars). A continuation chunk did not carry mamba_last_track_seqlen, so a final chunk of <= 64 tokens (too short to write its own snapshot) dropped the previous one. The final prefill commit then returned early in cache.py, skipping the donate and our eager prompt-checkpoint save that KV parking relies on, so the next turn recomputed the whole prompt. The new parametrized test fails 10/16 cases without the prefill.py change and passes 16/16 with it (devbox, CPU). Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…sync Port of upstream FlashML-org#500 (alvarorsouza-arch, with chrisqianz's CPU fallback, bounds guard and parity test). The boolean-mask index in _invalidate_prefill_buffer has a data-dependent shape and forced a device-to-host sync twice per chunk per layer, stalling the prefill-overlap pipeline (on by default here, engine/config.py moe_prefill_overlap). Unmeasured on the RTX 5090 so far; upstream's synthetic run showed no wall-time change, the production A/B reported shorter long-context turns. The CUDA parity test skips on the GPU-less devbox. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Port of upstream FlashML-org#471 (cherry77-cloud). temperature=0 or top_k=1 is now argmax whatever top_p says, and a greedy row batched with a sampled one (two chats at once) takes argmax instead of sampling at T=1e-6. Fork follow-through: spec_filter_params reads the new greedy_mask so the MTP draft filter agrees with the server for a greedy row in a mixed batch, the request_filter_params docstring drops the old "temperature 0 + top_p is sampled" rule, and test_spec_draft_graph pins the new triple. Devbox tests/engine: 7 failed / 749 passed after vs 8 / 748 before; the remaining 7 fail identically on the base commit (no flashinfer / spec_draft env). Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Port of upstream FlashML-org#367 (cherry77-cloud), adapted to the dynamic KV pool. allocate_paged gives each request whole 64-token pages, but admission charged raw tokens, so with two running requests the reservation could fall up to ~126 tokens short and overcommit a nearly full pool. kv_reservation_tokens() in scheduler/cache.py is the one page-rounded cost, used by PrefillAdder's never-fits gate, its admission_fits checks, the reserved_size charge, the preparation carry-over reservation, and the dynamic pool's probe_admission need_now, so the probe and the real admission keep agreeing (upstream's patch predates the probe). Devbox: tests/scheduler 558 passed. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The final-prefill checkpoint save ran inline, by design, so it held the first answer token while it copied. It rarely ran before upstream FlashML-org#505: the chunk continuation dropped mamba_last_track_seqlen and the commit returned early. With FlashML-org#505 it runs on every long prompt. On the 5090 (2026-09-23, 50k-token prompt, RAM parking, 2 GiB budget) that took cold TTFT from 18-25 s to 28-44 s, with an 8-11 s stall before the last chunk. py-spy put ~3.4 s of scheduler time in the save: 1.2 s pinned allocation, ~2 s per-page copy slicing, ~1 s copy. _save_prompt_checkpoint now queues the save with ParkStore.offer. The checkpoint node takes one extra lock until the worker's D2H copy completes, so its canonical pages and frozen tree slot cannot be evicted or reissued. drain_pending_parks drops only that lock; nothing is freed and the node stays in the tree. A full queue or a store without a worker falls back to the synchronous save. Tests: checkpoint tests settle the worker before inspecting the store. The slot-pressure test accepts "saved or held by a copy lock". A new gated-worker test shows the commit returns before the copy and holds the node until it lands; it fails on the inline version. Devbox: tests/scheduler + tests/kvcache 975 passed. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
… kernel Measured on the 5090 on 2026-09-23 (same 50k-token prompt A/B, both with FlashML-org#505): cold TTFT was 29-32 s with the kernel and 28-30 s without it, and prefill chunks took ~3.2 s with it and ~2.8 s without. The hidden sync it removes is mostly the host waiting for prefill GPU work that has to finish anyway (py-spy: that time is the GPU, not the sync). Upstream's synthetic run also showed no wall-time change, so the extra Triton kernel is not worth carrying. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Live on the 5090 on 2026-09-23 (with the background checkpoint save): a
governor MoE-only cache step re-captured the decode graphs while the
park worker was copying a prompt checkpoint. The global-mode capture was
invalidated (cudaErrorStreamCaptureInvalidated), the worker's save failed
("operation not permitted when stream is capturing"), and the scheduler
latched failed with the server at 503.
- engine/graph.py captures in thread_local mode, as the three MTP
captures already do for background threads, so the worker's pinned
allocation and D2H copies on its own stream cannot invalidate it.
- rebuild_cache waits for queued park copies before a MoE-only step too.
Only KV and mamba resizes went through prepare_rebuild's flush, so
teardown and re-capture never overlap the worker.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…waits Review of the background checkpoint save found two follow-ups: - The MTP shadow verifier borrows an inactive GDN slot under full slot pressure, and a queued checkpoint's tree slot looked inactive, so a concurrent copy could park a corrupted state. It now skips CacheManager.pending_park_slots(). The shadow is off on the box (FREETOKEN_MTP_SHADOW=0). - ensure_mamba_slots and _allocate blocked on any pending park, but a checkpoint hold releases only a lock. They now wait only when a pending park returns pages or slots. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…AM mode A 782-page (50k-token, 742 MiB) prompt checkpoint parked to RAM at 25-50 MB/s on the 5090 under WSL2: ~47k per-view D2H copies, a cudaHostAlloc of the whole entry on every save (~1.2 s, up to ~30 s under memory pressure), and a second full D2H pass through a fresh pinned window for the continuation prefix check. - QSAKVCache.page_byte_regions exposes the storage behind page_byte_views as [num_pages, bytes] uint8 views in the same order, so one index_select per region per chunk rebuilds the exact page-major bytes (SSD format and offsets unchanged). - RAM saves gather <= 32 MiB of whole pages into a reusable 2 x 32 MiB device staging buffer and move each chunk with one copy into ordinary pageable memory; nothing is pinned per save and eviction no longer pays cudaFreeHost. - The continuation prefix check brings the parent's bytes H2D chunk by chunk and compares on the device; restore stages H2D and scatters with index_copy_. - Pools whose regions do not line up with page_byte_views keep the per-view path. SSD mode is untouched. CPU microbench (scripts/bench_park_ram_copy.py, 782 pages): root save 1,202 -> 251 ms, continuation 1,017 -> 289 ms, restore 739 -> 453 ms. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
- Byte identity of the region gather against the per-view serialization for shuffled pages (BF16 and FP8), and RAM buffer == SSD file regions. - Chunking: sub-page, fractional, exact and larger-than-entry chunks stay on page boundaries; chained restores across chunks with page offsets. - Continuation prefix check: a match links, one flipped byte in a K/V, index or scale region of a borrowed page falls back to a root. - No pinned allocation in RAM mode; mismatched regions fall back to per-view. - test_idle_threshold_waits_until_the_leaf_is_old_enough held no copy open while asserting the copy was in flight: 1/30 failures before, 3/30 with the faster copy, 0/30 now that the copy is held until the check runs. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
… a racy test Review follow-ups for the page-major RAM copy: - _page_source raises IndexError for a page id outside the pool before index_select/index_copy_. On CUDA an out-of-range index is a device-side assert that poisons the context; the old per-view path raised IndexError and only disabled parking. - The twelve-slot checkpoint test asserted "saved or queued" in the window after copy_done (hold drained) and before _publish. That state is safe because the copy has finished. The test now flushes the store before failing. The flake also hit the parent commit (3/64 under load); 20/20 now. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
On the 5090 the test's FakeStream reached the new page-major source (torch.cuda.stream needs .device) before the patched _copy_to_ram, so the recorded error was the fake's AttributeError. Inject the failure at _page_source too; the fence-before-release assertion is unchanged. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Owner
Author
|
5090 results (2026-09-23), live on the box with RAM parking:
🤖 Generated with Claude Code |
dejay2
changed the base branch from
perf/upstream-picks-2026-09-23
to
mtp-upstream-merge
September 23, 2026 17:10
2 of 3 tasks
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Stacked on #11. Makes KV-parking RAM mode fast enough for the per-prompt checkpoint saves that FlashML-org#505 switched on.
Why
On the 5090 (2026-09-23) a 50k-token checkpoint (742 MiB) saved at 25-50 MB/s: 8-11 s inline, and up to ~30 s under Windows RAM pressure. Causes:
cudaMemcpyAsyncper view. About 1 s of Python and dispatch per save before any copying.cudaHostAllocof the full entry on every save. About 1.2 s, and an implicit device sync; eviction also paidcudaFreeHostsyncs.Restore had the same per-view cost.
What
QSAKVCache.page_byte_regions(): the pool's per-page byte tables, zero-copy, inpage_byte_viewsorder.Test plan
tests/scheduler tests/kvcache1017 passed, 10 skipped (baseline 976/8; +41 new tests: byte identity with the old serialization, RAM equal to the SSD file, chunk boundaries, one-byte prefix mismatch, no pinned allocation, fallback)scripts/bench_park_ram_copy.py, 782 pages, 743 MiB): root save 1,202 → 251 ms, continuation 1,017 → 289 ms, restore 739 → 286-453 ms. This is CPU only, not GPU/WSL speed.test_cuda_ram_save_is_pageable_and_byte_identicaland the park CUDA tests), then the live 50k-token A/B with RAM parking🤖 Generated with Claude Code