Skip to content

Stop the same-frame throttle from dropping commands that await an ACK #11

Description

@X-F-R

Context

Arm._tx_allowed() (src/litearm/arm.py:1382) is a same-frame throttle on the single write point. When a command code has a positive interval configured in _tx_repeat_min_interval, a repeat with a byte-identical payload inside that interval is dropped:

with self._tx_repeat_lock:
    if (self._tx_last_payload.get(cmd) == payload
            and now - self._tx_last_stamp.get(cmd, 0.0) < min_interval):
        self._tx_throttled_frames += 1
        return False

It sits on _raw_write (src/litearm/arm.py:1370):

if not self._tx_allowed(cmd, payload):
    return

The table is empty by default, so the mechanism is off unless a caller configures it.

Expected

Dropping a frame is a local decision, and a caller waiting on that frame should be told so locally.

Actual

The return is silent with respect to the caller. Every command that awaits an answer does so through expect(), which has no way to know the frame never left: it waits out its whole window and then reports that the link did not answer.

So a dropped movej produces a "no reply" error naming the command, and the reader is pushed towards the cable, the port and the firmware. The actual cause is a counter in this process, _tx_throttled_frames, which the caller never sees and the error message never mentions.

The drop is also not free of consequences further down. On _raw_write the throttle is deliberately ordered before the cart-queue clear (src/litearm/arm.py:1366), and the comment explains that this ordering only holds because a dropped frame leaves the firmware untouched. That reasoning is sound, but it means the throttle is load-bearing for the queue-clear invariant: anything that later moves it, or drops a frame after the clear, breaks test_a_dropped_frame_does_not_clear_the_cart_queue in a way that is hard to see.

Reproduce

arm.connect()
arm._set_tx_repeat_min_interval(P.CMD_MOVE_J, 1.0)
arm.movej([0.0] * 7)      # sent
arm.movej([0.0] * 7)      # frame dropped; raises a "no reply" error

The second call reports the link as unresponsive for a frame this process chose not to send.

I have not run this against the Python implementation. It is the reason the C++ port does not carry this mechanism at all: there, the single write point's invariant is that reaching it means the frame goes out, so the "queue cleared but the frame never sent" half-state cannot be constructed.

Suggested direction

Two options, in order of preference:

  1. Return the drop to the caller so _raw_write can raise a local error naming the throttle, instead of letting expect time out.
  2. Keep the drop but have the awaiting path consult _tx_throttled_frames and report the local cause when it is the reason the window expired.

Either keeps the feature while removing the misattribution.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugSomething isn't working

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions