Repository navigation
fix(reliability): coalesce BusStateMonitor's error-frame rechecks - #87
Conversation
BusStateMonitor subscribes ErrorFrameReceived/FaultOccurred as low-latency hints and posted one recheck per hint. A bus-off or error-passive storm raises those thousands of times per second, far faster than the loop can drain them, so a transient bus fault turned into an unbounded mailbox backlog — starving exactly the protocol work the state change exists to abort (FR-RAW-020..022). At most one un-run hint recheck is now outstanding: hints arriving while one is queued are dropped instead of posted. That is lossless for what the monitor reports, because a recheck is a sample of a level (ICanBus.BusState is a plain getter), not the delivery of a queued event — N back-to-back samples of an unchanged level say what one says. The gate is released before the sample is taken, so a hint racing an in-flight recheck posts a follow-up and the last hint of a storm is always succeeded by a sample taken after it; releasing afterwards would push that hint's state change out to the next poll tick. What is deliberately not preserved is a per-error-frame count: an error frame is not a state transition, and the hints were never a transition log. Every edge the monitor samples is still raised individually and in order, with Previous chained to the last reported state. Sampling granularity is unchanged and still tuned the way it always was, via pollInterval. Public surface is untouched: no new type, member or parameter. Tests drive the monitor through a queue-only IProtocolActor double whose Schedule never becomes due, so the poll is out of the picture and post counts are exact without any wall-clock waiting. Closes #22 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
PR SummaryMedium Risk Overview Coalescing caps hint-driven traffic at one outstanding recheck via an Tests add Reviewed by Cursor Bugbot for commit a68994e. Bugbot is set up for automated code reviews on this repo. Configure here. |
Closes #22.
BusStateMonitorsubscribesErrorFrameReceived/FaultOccurredas low-latency hints and postedone recheck per hint. A bus-off or error-passive storm raises those thousands of times per second —
far faster than the loop drains them — so a transient bus fault became an unbounded mailbox backlog
that starves exactly the protocol work a
BusOffexists to abort (FR-RAW-020..022), and fedstraight into the unbounded drain of #21.
What is coalesced
Hint posts, not state transitions. At most one un-run hint recheck is outstanding in the
mailbox at any time; hints arriving while one is queued are dropped rather than posted
(
Interlocked.CompareExchangegate, released by the recheck itself).This is lossless with respect to what the monitor reports, because a recheck is a sample of a
level —
ICanBus.BusStateis a plain getter with no change event — and not the delivery of aqueued event. N back-to-back samples of an unchanged level yield exactly what one sample yields.
What is dropped is redundant mailbox traffic.
The gate is released before the sample is taken. That ordering is the correctness argument: a
hint raised while the recheck is in flight then claims the gate again and posts a follow-up, so the
last hint of a storm is always succeeded by a sample taken after it. Releasing afterwards would let
that hint be dropped and push the state change it announced out to the next poll tick — turning the
hints' latency guarantee back into a poll-interval one at the exact moment it matters most. There
is a dedicated test for that ordering (see below).
What is deliberately preserved, and what is not
Preserved:
StateChanged, in order, withPreviouschained to the last reported state. A subscriber neversees a gap, a re-ordering, or a wrong
Previous.CurrentStateconvergence. Because the gate is released before sampling, the final state ofa storm is always observed by a hint-driven recheck, not merely by the next poll.
optimisation, and an adapter that refuses the subscriptions still degrades to poll-only.
Not preserved (and never was):
exposed a count of them.
BusStateMonitor's only observable isStateChanged.through
ErrWarningandErrPassiveon its way toBusOffbetween two samples, oneErrActive → BusOffedge is reported. This was already true before this change — a hint says"something happened", not "this transition happened", so whether the intermediate levels were
caught depended on whether a redundant sample happened to land on them. It is also exactly what
happens today on the poll-only path (
AllowErrorInfo=false). The honest statement is that themonitor is a sampler; the observable difference is that the storm no longer buys extra samples by
flooding the mailbox.
Why no opt-out was added: the existing public knob already covers the only legitimate need. A
consumer who wants finer sampling granularity shortens
pollInterval, which raises the sample ratedeterministically. An opt-out would instead be a switch labelled "please flood my mailbox", whose
sampling benefit is incidental and load-dependent, and it would grow the public surface of a
shipping package to preserve a side effect of a bug.
Public API
Unchanged. No new type, member, parameter or overload;
ApiApprovals/CanKit.Pro.Reliability.approved.txtis untouched, and both
netstandard2.0andnet10.0assets build as before. No new clock ortiming primitive was needed, so nothing here depends on
TickCount64or anything else missing onnetstandard2.0.Tests
Three tests in
BusStateMonitorTests, driven through a new privateMailboxActordouble — anIProtocolActorthat only queues onPost, drains when the test says so, and returns aSchedulehandle that never becomes due. That removes the self-rearming poll from the picture, so every post
counted is hint-driven and no assertion depends on machine speed. There is no
Task.Delayand nowall-clock synchronisation anywhere in them.
An_Error_Frame_Storm_Coalesces_Into_One_Pending_Actor_Post— 1000 error frames in, exactly onepost and a mailbox depth of one out.
The_Coalesced_Recheck_Reports_The_Transition_And_Reopens_The_Gate— the storm's transition isstill reported exactly once with the right
Previous/Current, and a second storm after therecheck has run posts again (the gate is a window, not a one-shot latch).
A_Hint_Arriving_While_The_Recheck_Is_In_Flight_Posts_A_Follow_Up— a hint raised from insidethe
StateChangedhandler, i.e. in the one window where release ordering decides whether atransition can be lost, must post a follow-up that observes it.
How the tests were verified to fail without the fix
BusStateMonitor.cstoorigin/mainand re-ran: tests 1 and 2 go red withfound 1000againstexpected 1(post count, and drained work items). Test 3 passes there, asit must — with no coalescing at all nothing can be swallowed.
swapping
HintRecheckOnLoopto release the gate afterRecheckOnLoop()makes it go red withExpected actor.MailboxDepth to be 1 ... but found 0, i.e. the in-flight hint was swallowed.The other two tests stay green under that mutant, so each of the three fails for its own reason.
Full suite:
dotnet test CanKit.Pro.sln -c Release→ 412 passed, 0 failed. Thenet48leg wascompile-checked with
dotnet build tests/CanKit.Pro.Tests -f net48 -p:CanKitProTestNetFrameworkLeg=true -c Release(0 warnings, 0 errors); it cannot be executed on macOS.
Test-infrastructure notes
ControllableBus.RaiseErrorFrame(...)added —ErrorFrameReceivedhad no way to be raised, andits
#pragma warning disable CS0067 // Never raised: nothing in CanKit.Pro subscribes to thesecomment was already stale, since
BusStateMonitordoes subscribe to it. Moved out of that group.StubErrorInfoadded: CanKit's ownICanErrorInfoimplementation is internal toCanKit.Core,so there is nothing public to construct, and passing
nullinto an event whose contract says itcarries error info would quietly excuse a subscriber that dereferences it.
Relationship to #21
Complementary, and no files overlap. This PR reduces what is produced into the mailbox; #21
bounds how it is drained. The produce side is the right place for this particular fix: the
information content of N hint posts is identical to that of one, so dropping them costs nothing
that a drain-side bound could give back — a bounded drain would still have to carry, order and
eventually discard 1000 identical work items per storm-millisecond.
Defect noticed but not fixed
dotnet format CanKit.Pro.sln --verify-no-changesfails onmain(whitespace inTestCases/Uds/UdsTransferTests.cs, import ordering inUdsTransferTests.cs,UdsClientTests.csand
Nfr006ErrorArchitectureTests.cs). Pre-existing, unrelated to this change, and left alone.🤖 Generated with Claude Code