What was observed
J1939TpTests.Bam_Roundtrip_ReceiverReassemblesIdenticalPayload failed once, on the macos-latest leg of #276 (run 37220399008, head d80756a): Failed: 1, Passed: 1531, Total: 1532, on net10.0. The other two legs passed on the same head.
CanKit.Pro.J1939Tp.J1939TpAbortException : J1939-TP Bam RX session timed out (T1) waiting for next TP.DT from 0x11.
at J1939TpChannel.UnwrapInboxItem (J1939TpChannel.cs:372)
at J1939TpChannel.ReceiveAsync (J1939TpChannel.cs:349)
at J1939TpTests.Bam_Roundtrip_ReceiverReassemblesIdenticalPayload (J1939TpTests.cs:166)
#276 changes ci.yml, GitVersion.yml, a shell script and Markdown; no .cs file is in the diff, so the change under test is not involved.
What the numbers say
The test sends a 100-byte BAM as TP.CM + 15 TP.DT with bamPacketSpacing = 5 ms (J1939TpTests.cs:148), each DT scheduled on the sender's actor (J1939TpChannel.cs:961, :1144). The receiver's T1 is the default 750 ms (J1939TpOptions.cs:14) and is re-armed per DT. For it to expire, two consecutive DT frames have to arrive more than 750 ms apart in a stream the sender spaces at 5 ms. The margin is 150× the spacing, so the test is not a tight-tolerance case like the three in #92; something stalled sender, virtual bus or receiver actor for three quarters of a second. The failure sits 4 min 28 s into the suite on a runner whose whole suite took 4 m 48 s against the Linux leg's 3 m.
The same morning the macos-latest leg of the main run for #274 (run 37195459987) failed on a different wall-clock assertion, J1939NodeTests.Disposing_A_Periodic_Schedule_Cancels_An_Emission_Waiting_For_Its_Confirmation, by 1 ms over a 500 ms bound. Two different timing tests red on the macOS runner on one day is consistent with that runner being slow today; it does not rule out a stall inside the actor.
What is not known
Whether the >750 ms gap was host perturbation or a scheduling defect in the BAM pacing or in the receiver's T1 re-arm. The log holds neither the DT index at which T1 fired nor how many DTs the receiver had, so the next occurrence will not say either unless the abort message carries them.
Suggested next step
Have the T1 expiry message name the packet index it was waiting for and the count received (ExpireRxSession, J1939TpChannel.cs:790), so a repeat says whether the gap was at a fixed point in the stream or random. Not reproduced locally in this issue's filing; if it recurs, run the test under forced actor starvation before deciding which side it is on.
Related: #92 (closed by #156), #240 (same shape, CANopen, ubuntu).
🤖 Generated with Claude Code
What was observed
J1939TpTests.Bam_Roundtrip_ReceiverReassemblesIdenticalPayloadfailed once, on themacos-latestleg of #276 (run 37220399008, headd80756a):Failed: 1, Passed: 1531, Total: 1532, on net10.0. The other two legs passed on the same head.#276 changes
ci.yml,GitVersion.yml, a shell script and Markdown; no.csfile is in the diff, so the change under test is not involved.What the numbers say
The test sends a 100-byte BAM as TP.CM + 15 TP.DT with
bamPacketSpacing= 5 ms (J1939TpTests.cs:148), each DT scheduled on the sender's actor (J1939TpChannel.cs:961, :1144). The receiver's T1 is the default 750 ms (J1939TpOptions.cs:14) and is re-armed per DT. For it to expire, two consecutive DT frames have to arrive more than 750 ms apart in a stream the sender spaces at 5 ms. The margin is 150× the spacing, so the test is not a tight-tolerance case like the three in #92; something stalled sender, virtual bus or receiver actor for three quarters of a second. The failure sits 4 min 28 s into the suite on a runner whose whole suite took 4 m 48 s against the Linux leg's 3 m.The same morning the
macos-latestleg of themainrun for #274 (run 37195459987) failed on a different wall-clock assertion,J1939NodeTests.Disposing_A_Periodic_Schedule_Cancels_An_Emission_Waiting_For_Its_Confirmation, by 1 ms over a 500 ms bound. Two different timing tests red on the macOS runner on one day is consistent with that runner being slow today; it does not rule out a stall inside the actor.What is not known
Whether the >750 ms gap was host perturbation or a scheduling defect in the BAM pacing or in the receiver's T1 re-arm. The log holds neither the DT index at which T1 fired nor how many DTs the receiver had, so the next occurrence will not say either unless the abort message carries them.
Suggested next step
Have the T1 expiry message name the packet index it was waiting for and the count received (
ExpireRxSession,J1939TpChannel.cs:790), so a repeat says whether the gap was at a fixed point in the stream or random. Not reproduced locally in this issue's filing; if it recurs, run the test under forced actor starvation before deciding which side it is on.Related: #92 (closed by #156), #240 (same shape, CANopen, ubuntu).
🤖 Generated with Claude Code