You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
TestBackpressure's WebSocket oracle counts clientCloseFail, prints it in its summary line, and never asserts on it. One run of the io_uring variant recorded:
Sixty-six of ninety-six connections never reached the close handshake at all. The test reported success.
closeTimeout is asserted; clientCloseFail is not. So the oracle fails only when a connection reaches the close handshake and then stalls, and stays silent when the connection never gets that far. That is the larger of the two populations here, and it is the same underlying stall: the probe in #607 caught these connections mid-write rather than mid-close-wait, with the same empty middleware reader, the same starved handler, and the same unread bytes in the server's kernel receive queue.
This is not a theoretical hole. It is why #607 looked like a 12-in-24 problem rather than what the runs actually contain, and why the oracle's own headline number moved in the wrong direction under conditions that should have made it worse.
What to change
Assert on clientCloseFail, or fold it into the same assertion as closeTimeout with a message that distinguishes them. A connection that never reaches the close handshake is not a pass.
Keep both counters printed separately. They are different stall points and the distinction is diagnostic.
Sibling of #611, found while root-causing #607.
TestBackpressure's WebSocket oracle countsclientCloseFail, prints it in its summary line, and never asserts on it. One run of the io_uring variant recorded:Sixty-six of ninety-six connections never reached the close handshake at all. The test reported success.
closeTimeoutis asserted;clientCloseFailis not. So the oracle fails only when a connection reaches the close handshake and then stalls, and stays silent when the connection never gets that far. That is the larger of the two populations here, and it is the same underlying stall: the probe in #607 caught these connections mid-write rather than mid-close-wait, with the same empty middleware reader, the same starved handler, and the same unread bytes in the server's kernel receive queue.This is not a theoretical hole. It is why #607 looked like a 12-in-24 problem rather than what the runs actually contain, and why the oracle's own headline number moved in the wrong direction under conditions that should have made it worse.
What to change
clientCloseFail, or fold it into the same assertion ascloseTimeoutwith a message that distinguishes them. A connection that never reaches the close handshake is not a pass.Related