fix(server): guarantee AbortMsg delivery on stream cancellation - #222
Artemowka22 wants to merge 4 commits into
Conversation
stream_with_cancellation reacted to a client disconnect with an unowned asyncio.create_task(abort_user(uid)): request teardown could outrun delivery and a failure inside the task degraded to a never-retrieved-exception warning, so the scheduler kept decoding for a client that was gone. Await the abort inline behind asyncio.shield (a second cancellation cannot kill the delivery task), and make abort_user claim the uid first — exactly one AbortMsg even if cancellation runs twice, and none at all when the stream already finished normally. Found via the freetoken-mlx downstream audit (docs/AUDIT.md, defect 2).
|
Supporting evidence, and a case this PR may not cover: the non-streaming path also keeps generating after the client is gone. Environment: FreeToken git Repro: a non-stream |
Follow-up to the review evidence on FlashML-org#222 (benwilson): the non-streaming path kept generating after the client was gone -- with --max-running-requests 1 an abandoned request is a full outage for its remaining max_tokens (measured repro: ~70 s of dead decode, the next client's first token 61 s late). stream_with_cancellation was the only place a disconnect was observed. Give the plain handlers its non-streaming twin: _await_watching_disconnect() runs the generation drain as a task and polls request.is_disconnected() once a second; when the client goes away it delivers the same shielded abort_user (claim + AbortMsg first, so the drain task's own cleanup cannot swallow the claim), then winds the drain down and answers 499 (client closed request -- for the access log; the wire is dead). Handler cancellation (server shutdown) delivers the abort too, mirroring the streaming path. Covers /v1/chat/completions and each prompt of a non-streaming /v1/completions batch. Requests with request=None (adapter-internal callers) are unaffected. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
|
Thanks for the measured repro — that is the other half of the bug, and with The PR now covers it: the non-streaming handlers run the generation drain as a task and poll Tests cover the disconnect-abort for both endpoints and the no-abort happy path. |
…gprobs for chat and legacy completions Upstream FlashML-org#224 at 855650d, merged onto deploy/chatdnp for the PR sweep. Conflicts: engine.py keeps FlashML-org#231's stats readout before the logprobs-aware return; openai_api.py keeps the vision `images` argument and FlashML-org#222's disconnect-watching drain with the logprobs entries added; generation.py keeps FlashML-org#266's marker filter and routes every content delta through FlashML-org#224's _content_delta so the logprobs entries ride the filtered text. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0173pf9k9fSVtwbm3f898HDt
|
Tried on 2 x RTX 6000 Ada (sm_89) serving Qwen3.8-Flash-Next (RadixArk NVFP4) at TP=2, offload backend, merged onto my deploy branch with ten other open PRs, tests run on the box, then put in production. Merged clean, the three unit tests pass, and a Two things this PR's path cannot see on the installed stack (uvicorn 0.52.4, Starlette 1.6.0):
Those two turned out to be red herrings on this stack (a minimal uvicorn + Starlette app delivers the disconnect fine, with or without a What fixed it on my branch (three commits, regression tests added): key the abort's idempotency on a set of aborted uids instead of the maps (the scheduler acks an abort for a uid it no longer has), deliver the abort from |
…inally On a disconnect the CancelledError unwinds from the response task's innermost await -- wait_for_ack -- whose finally pops ack_map and event_map on the way out. stream_with_cancellation then called abort_user, whose claim was `event_map.pop(uid) is not None`: always False by that point, so the AbortMsg was never sent and the scheduler kept decoding to max_tokens, with no log line either (the cancellation interrupts the async for, not the is_disconnected branch). A body closed after the response started arrives as GeneratorExit, which the `except CancelledError` never saw either. Diagnosed with a live repro by gdevenyi on the PR thread. Key the claim on an insertion-ordered aborted_uids dict, FIFO-bounded like the scheduler's abort tombstones (a duplicate AbortMsg is safe: the scheduler acks aborts for uids it no longer has), and move delivery into the stream wrapper's finally, gated on the stream not finishing. One path now covers server-side cancellation, late closes, and generator errors; the accounting drain's abort barrier also stops missing requests whose maps were already emptied. Assisted-by: Claude
|
Confirmed, and thank you — the unwind analysis is exactly right and your probe caught what our disconnect tests could not: the CancelledError lands on Folded into the PR (new commit):
Regression tests added: cancellation delivered while the drain generator is suspended in The Validation on our side is CPU-only (macOS: the server suite, 544 passed, plus the full-tree run diffed against the clean base — identical failure lists, all environmental). If you get a chance to point |
One conflict, in python/freetoken/server/openai_api.py: upstream made the default output budget configurable (default_max_tokens, FlashML-org#411) on the same lines where this branch wraps the non-streaming completion drain for disconnect delivery. Kept the drain wrapper; its sampling resolution now passes default_max_tokens through, as upstream does elsewhere. Assisted-by: Claude
Follow-up to the review evidence on FlashML-org#222 (benwilson): the non-streaming path kept generating after the client was gone -- with --max-running-requests 1 an abandoned request is a full outage for its remaining max_tokens (measured repro: ~70 s of dead decode, the next client's first token 61 s late). stream_with_cancellation was the only place a disconnect was observed. Give the plain handlers its non-streaming twin: _await_watching_disconnect() runs the generation drain as a task and polls request.is_disconnected() once a second; when the client goes away it delivers the same shielded abort_user (claim + AbortMsg first, so the drain task's own cleanup cannot swallow the claim), then winds the drain down and answers 499 (client closed request -- for the access log; the wire is dead). Handler cancellation (server shutdown) delivers the abort too, mirroring the streaming path. Covers /v1/chat/completions and each prompt of a non-streaming /v1/completions batch. Requests with request=None (adapter-internal callers) are unaffected. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> (cherry picked from commit 274a2ce)
…ream cancellation # Conflicts: # python/freetoken/server/api_server.py # python/freetoken/server/openai_api.py
…isconnect Guards the FlashML-org#222 resolution against FlashML-org#393's fan-out: the disconnect watcher takes the whole uid list, not the first sample's uid. Assisted-by: Claude Fable 5.1 (cherry picked from commit 93ff96349235ea3ca1504733a2abbd6e6bcccf94)
…can see a client disconnect
Starlette's @app.middleware("http") (BaseHTTPMiddleware) wraps the handler's receive channel
and never yields the client's http.disconnect while a non-streaming request is still
computing, so request.is_disconnected() stays False and the disconnect watcher from
FlashML-org#222 could not stop an abandoned request. Streaming requests only
worked because the socket write failed. The recorder is the one such middleware; as pure
ASGI it records the same row at response start.
Assisted-by: Claude Fable 5.1
…ream cancellation # Conflicts: # python/freetoken/server/api_server.py # python/freetoken/server/openai_api.py # Conflicts: # python/freetoken/server/api_server.py # python/freetoken/server/openai_api.py
…isconnect Guards the FlashML-org#222 resolution against FlashML-org#393's fan-out: the disconnect watcher takes the whole uid list, not the first sample's uid. Assisted-by: Claude Fable 5.1 (cherry picked from commit 93ff96349235ea3ca1504733a2abbd6e6bcccf94) (cherry picked from commit 69a5efa)
…can see a client disconnect
Starlette's @app.middleware("http") (BaseHTTPMiddleware) wraps the handler's receive channel
and never yields the client's http.disconnect while a non-streaming request is still
computing, so request.is_disconnected() stays False and the disconnect watcher from
FlashML-org#222 could not stop an abandoned request. Streaming requests only
worked because the socket write failed. The recorder is the one such middleware; as pure
ASGI it records the same row at response start.
Assisted-by: Claude Fable 5.1
(cherry picked from commit 8adde91)
…ream cancellation # Conflicts: # python/freetoken/server/api_server.py # python/freetoken/server/openai_api.py # Conflicts: # python/freetoken/server/api_server.py # python/freetoken/server/openai_api.py # Conflicts: # python/freetoken/server/api_server.py # python/freetoken/server/openai_api.py
…isconnect Guards the FlashML-org#222 resolution against FlashML-org#393's fan-out: the disconnect watcher takes the whole uid list, not the first sample's uid. Assisted-by: Claude Fable 5.1 (cherry picked from commit 93ff96349235ea3ca1504733a2abbd6e6bcccf94) (cherry picked from commit 69a5efa) (cherry picked from commit 40e9941)
…can see a client disconnect
Starlette's @app.middleware("http") (BaseHTTPMiddleware) wraps the handler's receive channel
and never yields the client's http.disconnect while a non-streaming request is still
computing, so request.is_disconnected() stays False and the disconnect watcher from
FlashML-org#222 could not stop an abandoned request. Streaming requests only
worked because the socket write failed. The recorder is the one such middleware; as pure
ASGI it records the same row at response start.
Assisted-by: Claude Fable 5.1
(cherry picked from commit 8adde91)
(cherry picked from commit 202d7fe)
…etrics timing reads Upstream FlashML-org#504 times the prefill span in Scheduler._prefill_start; the stubs from FlashML-org#222 and FlashML-org#505's tests build the scheduler without it. Assisted-by: Claude Opus 5.5
Summary
FrontendManager.stream_with_cancellationreacts to a client disconnect with a bareasyncio.create_task(self.abort_user(uid))and re-raises. The task has no owner: request teardown can complete before the abort coroutine ever runs, and an exception inside it degrades to a "Task exception was never retrieved" warning. The scheduler then keeps decoding for a client that is gone — under a longmax_tokensthis wastes the GPU for minutes and pins KV pages.Fix (ported from a downstream audit — agisota/freetoken-mlx, docs/AUDIT.md, defect 2):
abort_useridempotent — exactly oneAbortMsgper uid even if the handler runs twice (disconnect + server-side cancel).Test plan
pytest tests/server/test_stream_cancellation.py(new): cancellation delivers exactly one AbortMsg before teardown finishes; double cancellation stays single-shot; normal completion sends no aborttests/server/suite passes