Skip to content

fix(scheduler): reserve paged KV at allocation granularity - #367

Merged
jason-fxz merged 1 commit into
FlashML-org:mainfrom
taking-lying-flat:fix/prefill-page-admission
Sep 10, 2026
Merged

jason-fxz merged 1 commit into
FlashML-org:mainfrom
taking-lying-flat:fix/prefill-page-admission

Conversation

@taking-lying-flat

Copy link
Copy Markdown
Contributor

Summary

  • Account prefill KV reservations in whole pages per request.
  • Use the same page-ceiling span as CacheManager.allocate_paged() for both admission checks and cumulative reservations.
  • Keep requests queued when the remaining KV pool cannot provide a distinct page for each request.

Problem

CacheManager.available_size reports token-equivalent capacity, but the allocator assigns whole pages independently to each request. PrefillAdder previously charged raw token counts, so multiple short requests could pass admission even when they required more physical pages than the pool contained. The subsequent allocate_paged() call could then fail during eviction or hit the "Eviction did not free enough space" assertion.

For example, with a 128-token page and one free page, two requests that each need two tokens were charged as four tokens total even though they require two separate pages.

Validation

  • 51 relevant scheduler tests passed across radix, chunked prefill, hybrid, SWA page-size, DSV4, and abort paths.
  • Ruff check passed for python/freetoken/scheduler/prefill.py.
  • git diff --check passed.

@jason-fxz

Copy link
Copy Markdown
Collaborator

LGTM, thanks for the fix.

@jason-fxz
jason-fxz merged commit 46d2743 into FlashML-org:main Sep 10, 2026
gdevenyi pushed a commit to gdevenyi/FreeToken that referenced this pull request Sep 11, 2026
danielamadori pushed a commit to danielamadori/DFlash-FreeToken that referenced this pull request Sep 15, 2026
danielamadori added a commit to danielamadori/DFlash-FreeToken that referenced this pull request Sep 15, 2026
3ff80c5 made /props report min(ceiling, pool), which is the right fix: the
engine enforces the lower of the two and reporting the other leaves the node
the only party believing a figure nothing enforces.

The two numbers are not in the same unit. total_pages is the manager's
num_pages; a page holds page_size tokens. Everything token-valued beside it
multiplies -- CacheManager.available_size is

    evictable + len(free_slots) * page_size

and the prefill adder's _kv_reservation_size, just cherry-picked from
upstream FlashML-org#367, returns in its own words 'the token-equivalent cost of the
additional KV pages'. Only this comparison did not convert.

It is invisible on the fleet as it stands, because thething and the Dell both
run --page-size 1, where pages and tokens coincide: the 121899 measurement
behind 3ff80c5 stands. On a node started with --page-size 16 the node would
advertise a SIXTEENTH of the context it holds, and a router would turn away
prompts it answers -- the same defect that commit removes, pointing the other
way.

page_size comes from the config build_props already has. A config whose
attribute is None or 0 reads as 1: zero pages of context would be reported as
no context at all, and the second test pins that.

17 passed; the two new tests fail against the old two-argument form.
nomanoma121 pushed a commit to nomanoma121/My-FreeToken that referenced this pull request Sep 16, 2026
nomanoma121 pushed a commit to nomanoma121/My-FreeToken that referenced this pull request Sep 16, 2026
…r's shape

FlashML-org#438 folds qwen3_5_moe's four dense readers into one _DenseReader that asks the checkpoint's
QuantConfig what each Linear stores, and deletes _iter_weights_attn_fp8 -- which is where
--spec-mtp's weight reading lived. The head is read again on the new shape: _rename keeps mtp.*
when the engine asks for it, the head's two pre_fc norms and its own final norm join the (1+w)
list, and its per-expert bf16 experts are gathered on the host and yielded last as the two
stacked tensors the engine quantizes into a bank layer.

Two things the old path needed code for are now free. mtp* is off every quantizer's list, so
scheme_for_name returns None for the head and it reads as stored -- the table that said "the
head's q|k|v are bf16 where the decoder's are fp8" is gone. And the up-front refusal
("MIXED_PRECISION only") went with attn_quant, which FlashML-org#438 stops parsing for this family: the
reader is layout-general now, so what is left to refuse is a head that is not in the checkpoint
(the stacker's assert) or one the exporter quantized (a loud NotImplementedError -- the engine
quantizes these itself, so a pre-quantized head would be stacked into the wrong format).

FlashML-org#428 reads qwen4_exp's dense projections through the same QuantConfig. Its _DenseFuser replaces
the fork's fuse_buf; the progress bar keeps asking the pipeline engine's rank rather than TP's.

FlashML-org#367 reserves paged KV at allocation granularity -- the same page-span accounting this fork
already does for the SWA pool, on the other currency. Admission is stricter than the token math
by up to one page per request, which is the point of it.

FlashML-org#411 honors the server's max_output_tokens in both APIs' defaults; it lands beside the image
parts in chat_request_to_genspec without touching them.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Rhonstin pushed a commit to Rhonstin/FreeToken that referenced this pull request Sep 16, 2026
MT-z added a commit to MT-z/FreeToken that referenced this pull request Sep 16, 2026
…L-org#367)

Upstream FlashML-org#367 by taking-lying-flat, taken unmodified.
`prefill.py` charged raw token counts against `cache_manager.available_size`
(lines 105/108) while `allocate_paged()` hands each request whole pages
independently, so several short requests could clear admission and then need
more physical pages than the pool holds -- failing in eviction or on the
"Eviction did not free enough space" assertion. Reservations are now counted in
whole pages, using the same page-ceiling span as the allocator.

Cheap insurance rather than an observed failure here: this box serves 4-6
concurrent requests on a KV pool that shares its budget with the MoE slot cache
(--moe-cache-auto), which is the regime where over-admission by page rounding
is easiest to hit. One file, +18/-4.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
MT-z added a commit to MT-z/FreeToken that referenced this pull request Sep 16, 2026
…reservation and FlashML-org#428

FlashML-org#367 changes PrefillAdder's KV reservation from raw tokens to whole pages, in the file this
box's admission behaviour hangs on (prefill.py reserves input_len + max_tokens up front, which
is what caps concurrency here rather than --max-running-requests).

On this box it is numerically the identity. page_size is 1 both by default (engine/config.py:66)
and in the running serve's own ServerArgs line, and at page_size 1 the new
div_ceil(total, 1) - div_ceil(cached, 1) is exactly the old extend_len + output_len --
remain_len is defined as input_len - cached_len two hunks down. It bites where page_size is
forced above 1: engine.py:308 (kpool indexer layout, 64) and engine.py:1282 (DSV4, 128).

FlashML-org#428 is qwen4_exp, which is not the model this box serves; it rides along with main.

Suite: 8 failed, 1,922 passed, 92 skipped. The same eight by name as
freetoken-systest results/20260910-suite-instr-merge.txt, not merely the same count --
test_e4m3_compat appears in that file's warnings summary, not its failures.

Assisted-by: Claude Opus 5
trcwebdesign pushed a commit to trcwebdesign/FreeToken that referenced this pull request Sep 21, 2026
dejay2 added a commit to dejay2/FreeToken that referenced this pull request Sep 23, 2026
…nd make KV checkpoint parking fast (#12)

Ports three upstream FlashML-org/FreeToken fixes and makes KV-parking prompt checkpoints cheap enough to run on every long prompt.

Upstream ports:
- FlashML-org#505: carry mamba_last_track_seqlen across prefill chunks, so multi-chunk prompts keep their snapshot and actually save a prompt checkpoint (before this, they almost never did).
- FlashML-org#471: greedy sampling in mixed batches; spec_filter_params follows greedy_mask.
- FlashML-org#367: page-rounded KV reservations (kv_reservation_tokens), shared with the dynamic pool's probe_admission.
- FlashML-org#500 was ported and reverted: no gain on the 5090.

Parking:
- Prompt checkpoints save on the background worker, with a node lock until the copy completes.
- Decode graph capture uses thread_local mode, and MoE-only steps flush parks first. Without this, the server latched live on 2026-09-23.
- The MTP shadow skips slots with a pending park.
- RAM-mode save, prefix check and restore gather whole pages through a 2x32 MiB device staging buffer into pageable host memory. This replaces ~47k per-view copies plus a per-save cudaHostAlloc. The byte format is unchanged.

5090, 50k-token prompts, RAM parking:

| | Before | After |
|---|---|---|
| Cold TTFT | 18-25 s | 16.0-16.2 s |
| Follow-up | 3.6-3.8 s | 2.3-2.4 s |

Windows free RAM stayed at 15 GB or more, with no page-outs. On the box, the park tests pass 202/202 and the fast-copy tests 43/43.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
lucaspirola pushed a commit to lucaspirola/FreeToken that referenced this pull request Sep 28, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants