forked from FlashML-org/FreeToken
-
Notifications
You must be signed in to change notification settings - Fork 0
Pull requests: gdevenyi/FreeToken
Author
Label
Projects
Milestones
Reviews
Assignee
Sort
Pull requests list
5080deploy: merge upstream #525, host-resident GDN snapshot / prefix tier for hybrid GDN models
#59
opened Sep 19, 2026 by
gdevenyi
Owner
Loading…
5080deploy: merge upstream #499, host-RAM KV tier for sparse-attention models (--kv-host-pages)
#58
opened Sep 19, 2026 by
gdevenyi
Owner
Loading…
5080deploy: merge upstream #222, deliver AbortMsg when the client disconnects (streaming and not)
#57
opened Sep 18, 2026 by
gdevenyi
Owner
Loading…
5080deploy: merge upstream #487, hoist system messages for system-first chat templates
#56
opened Sep 18, 2026 by
gdevenyi
Owner
Loading…
5080deploy: merge upstream #459, validate FTW v1 indexes eagerly
#55
opened Sep 18, 2026 by
gdevenyi
Owner
Loading…
5080deploy: merge upstream #198, reserve explicit KV pages during MoE auto-sizing
#54
opened Sep 18, 2026 by
gdevenyi
Owner
Loading…
5080deploy: merge upstream #340, reserve the pool's dummy page in the --moe-cache-auto KV floor
#53
opened Sep 18, 2026 by
gdevenyi
Owner
Loading…
5080deploy: merge upstream #476, --default-thinking-mode server flag
#52
opened Sep 18, 2026 by
gdevenyi
Owner
Loading…
5080deploy: merge upstream #118, clamp admission to the actual KV pool
#51
opened Sep 18, 2026 by
gdevenyi
Owner
Loading…
5080deploy: merge upstream #414, order the batch-memcpy probe against the current stream
#50
opened Sep 18, 2026 by
gdevenyi
Owner
Loading…
5080deploy: merge upstream #466, tolerate coalesced msgpack frames in the zmq pull queues
#49
opened Sep 18, 2026 by
gdevenyi
Owner
Loading…
5080deploy: merge upstream #505, preserve the pending hybrid checkpoint across prefill chunks
#48
opened Sep 18, 2026 by
gdevenyi
Owner
Loading…
feat(server): OpenAI protocol coverage -- sampling extras, n, echo/suffix, usage details, tokenize/detokenize/metrics
#47
opened Sep 15, 2026 by
gdevenyi
Owner
Loading…
4 tasks done
fix(tokenizer): accept the OpenAI developer role (map to system for templates without it)
#46
opened Sep 15, 2026 by
gdevenyi
Owner
Loading…
feat(qwen4_exp): load-time per-tensor FP8 dense projections (W8A8 via _scaled_mm)
#45
opened Sep 15, 2026 by
gdevenyi
Owner
Loading…
test(qwen4_exp): compare chunked prefill to one-shot at bf16 tolerance
#44
opened Sep 15, 2026 by
gdevenyi
Owner
Loading…
feat(qwen4_exp): tensor parallelism for Qwen3.8-Flash-Next (offload backend)
#43
opened Sep 15, 2026 by
gdevenyi
Owner
Loading…
feat(kvcache): 8-bit DSV4 window/compressed KV behind --kv-cache-dtype
#42
opened Sep 15, 2026 by
gdevenyi
Owner
Loading…
perf(dsv4): compute the hyper-connection pre-norm in one kernel
#41
opened Sep 15, 2026 by
gdevenyi
Owner
Loading…
perf(dsfp4): pick the grouped-prefill tile from route density on sm_89
#40
opened Sep 15, 2026 by
gdevenyi
Owner
Loading…
fix(fp8): round onto the e4m3 grid before the native float8e4nv downcast
#39
opened Sep 15, 2026 by
gdevenyi
Owner
Loading…
bench(bw): sweep decode batch size to measure cross-token expert dedup
#38
opened Sep 15, 2026 by
gdevenyi
Owner
Loading…
perf(cpu-moe): opt-in AMX tile GEMM for the deduped bf16 pass 1
#37
opened Sep 15, 2026 by
gdevenyi
Owner
Loading…
perf(cpu-moe): extend expert dedup to ds_fp4 and mxfp4
#36
opened Sep 15, 2026 by
gdevenyi
Owner
Loading…
Previous Next
ProTip!
Add no:assignee to see everything that’s not assigned.