fix(scheduler): reserve paged KV at allocation granularity - #367
Merged
jason-fxz merged 1 commit intoSep 10, 2026
Merged
Conversation
Collaborator
|
LGTM, thanks for the fix. |
danielamadori
pushed a commit
to danielamadori/DFlash-FreeToken
that referenced
this pull request
Sep 15, 2026
danielamadori
added a commit
to danielamadori/DFlash-FreeToken
that referenced
this pull request
Sep 15, 2026
3ff80c5 made /props report min(ceiling, pool), which is the right fix: the engine enforces the lower of the two and reporting the other leaves the node the only party believing a figure nothing enforces. The two numbers are not in the same unit. total_pages is the manager's num_pages; a page holds page_size tokens. Everything token-valued beside it multiplies -- CacheManager.available_size is evictable + len(free_slots) * page_size and the prefill adder's _kv_reservation_size, just cherry-picked from upstream FlashML-org#367, returns in its own words 'the token-equivalent cost of the additional KV pages'. Only this comparison did not convert. It is invisible on the fleet as it stands, because thething and the Dell both run --page-size 1, where pages and tokens coincide: the 121899 measurement behind 3ff80c5 stands. On a node started with --page-size 16 the node would advertise a SIXTEENTH of the context it holds, and a router would turn away prompts it answers -- the same defect that commit removes, pointing the other way. page_size comes from the config build_props already has. A config whose attribute is None or 0 reads as 1: zero pages of context would be reported as no context at all, and the second test pins that. 17 passed; the two new tests fail against the old two-argument form.
nomanoma121
pushed a commit
to nomanoma121/My-FreeToken
that referenced
this pull request
Sep 16, 2026
nomanoma121
pushed a commit
to nomanoma121/My-FreeToken
that referenced
this pull request
Sep 16, 2026
…r's shape FlashML-org#438 folds qwen3_5_moe's four dense readers into one _DenseReader that asks the checkpoint's QuantConfig what each Linear stores, and deletes _iter_weights_attn_fp8 -- which is where --spec-mtp's weight reading lived. The head is read again on the new shape: _rename keeps mtp.* when the engine asks for it, the head's two pre_fc norms and its own final norm join the (1+w) list, and its per-expert bf16 experts are gathered on the host and yielded last as the two stacked tensors the engine quantizes into a bank layer. Two things the old path needed code for are now free. mtp* is off every quantizer's list, so scheme_for_name returns None for the head and it reads as stored -- the table that said "the head's q|k|v are bf16 where the decoder's are fp8" is gone. And the up-front refusal ("MIXED_PRECISION only") went with attn_quant, which FlashML-org#438 stops parsing for this family: the reader is layout-general now, so what is left to refuse is a head that is not in the checkpoint (the stacker's assert) or one the exporter quantized (a loud NotImplementedError -- the engine quantizes these itself, so a pre-quantized head would be stacked into the wrong format). FlashML-org#428 reads qwen4_exp's dense projections through the same QuantConfig. Its _DenseFuser replaces the fork's fuse_buf; the progress bar keeps asking the pipeline engine's rank rather than TP's. FlashML-org#367 reserves paged KV at allocation granularity -- the same page-span accounting this fork already does for the SWA pool, on the other currency. Admission is stricter than the token math by up to one page per request, which is the point of it. FlashML-org#411 honors the server's max_output_tokens in both APIs' defaults; it lands beside the image parts in chat_request_to_genspec without touching them. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
MT-z
added a commit
to MT-z/FreeToken
that referenced
this pull request
Sep 16, 2026
…L-org#367) Upstream FlashML-org#367 by taking-lying-flat, taken unmodified. `prefill.py` charged raw token counts against `cache_manager.available_size` (lines 105/108) while `allocate_paged()` hands each request whole pages independently, so several short requests could clear admission and then need more physical pages than the pool holds -- failing in eviction or on the "Eviction did not free enough space" assertion. Reservations are now counted in whole pages, using the same page-ceiling span as the allocator. Cheap insurance rather than an observed failure here: this box serves 4-6 concurrent requests on a KV pool that shares its budget with the MoE slot cache (--moe-cache-auto), which is the regime where over-admission by page rounding is easiest to hit. One file, +18/-4. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
MT-z
added a commit
to MT-z/FreeToken
that referenced
this pull request
Sep 16, 2026
…reservation and FlashML-org#428 FlashML-org#367 changes PrefillAdder's KV reservation from raw tokens to whole pages, in the file this box's admission behaviour hangs on (prefill.py reserves input_len + max_tokens up front, which is what caps concurrency here rather than --max-running-requests). On this box it is numerically the identity. page_size is 1 both by default (engine/config.py:66) and in the running serve's own ServerArgs line, and at page_size 1 the new div_ceil(total, 1) - div_ceil(cached, 1) is exactly the old extend_len + output_len -- remain_len is defined as input_len - cached_len two hunks down. It bites where page_size is forced above 1: engine.py:308 (kpool indexer layout, 64) and engine.py:1282 (DSV4, 128). FlashML-org#428 is qwen4_exp, which is not the model this box serves; it rides along with main. Suite: 8 failed, 1,922 passed, 92 skipped. The same eight by name as freetoken-systest results/20260910-suite-instr-merge.txt, not merely the same count -- test_e4m3_compat appears in that file's warnings summary, not its failures. Assisted-by: Claude Opus 5
trcwebdesign
pushed a commit
to trcwebdesign/FreeToken
that referenced
this pull request
Sep 21, 2026
2 of 3 tasks
dejay2
added a commit
to dejay2/FreeToken
that referenced
this pull request
Sep 23, 2026
…nd make KV checkpoint parking fast (#12) Ports three upstream FlashML-org/FreeToken fixes and makes KV-parking prompt checkpoints cheap enough to run on every long prompt. Upstream ports: - FlashML-org#505: carry mamba_last_track_seqlen across prefill chunks, so multi-chunk prompts keep their snapshot and actually save a prompt checkpoint (before this, they almost never did). - FlashML-org#471: greedy sampling in mixed batches; spec_filter_params follows greedy_mask. - FlashML-org#367: page-rounded KV reservations (kv_reservation_tokens), shared with the dynamic pool's probe_admission. - FlashML-org#500 was ported and reverted: no gain on the 5090. Parking: - Prompt checkpoints save on the background worker, with a node lock until the copy completes. - Decode graph capture uses thread_local mode, and MoE-only steps flush parks first. Without this, the server latched live on 2026-09-23. - The MTP shadow skips slots with a pending park. - RAM-mode save, prefix check and restore gather whole pages through a 2x32 MiB device staging buffer into pageable host memory. This replaces ~47k per-view copies plus a per-save cudaHostAlloc. The byte format is unchanged. 5090, 50k-token prompts, RAM parking: | | Before | After | |---|---|---| | Cold TTFT | 18-25 s | 16.0-16.2 s | | Follow-up | 3.6-3.8 s | 2.3-2.4 s | Windows free RAM stayed at 15 GB or more, with no page-outs. On the box, the park tests pass 202/202 and the fast-copy tests 43/43. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
lucaspirola
pushed a commit
to lucaspirola/FreeToken
that referenced
this pull request
Sep 28, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Problem
CacheManager.available_size reports token-equivalent capacity, but the allocator assigns whole pages independently to each request. PrefillAdder previously charged raw token counts, so multiple short requests could pass admission even when they required more physical pages than the pool contained. The subsequent allocate_paged() call could then fail during eviction or hit the "Eviction did not free enough space" assertion.
For example, with a 128-token page and one free page, two requests that each need two tokens were charged as four tokens total even though they require two separate pages.
Validation