Problem
A DHT PUT counts as successful when one peer acknowledges it, but a peer acknowledges records that it then drops.
- The node filters inbound records itself (
StoreInserts::FilterBoth, crates/net/src/net_interface.rs:744). It rejects a record, with only a debug log, when the sender already has 64 records in its store (DHT_MAX_RECORDS_PER_PEER), when the store holds 1024 records (DHT_MAX_RECORDS), or when the size, expiry, or key check fails (net_interface.rs:982-1030).
- libp2p-kad 0.48 sends
PutRecordRes to the sender whether or not the application keeps the record (libp2p-kad-0.48.0/src/behaviour.rs:1931-1945).
- The publisher uses
Quorum::One (net_interface.rs:1702), so one acknowledgement from a peer that dropped the record completes the PUT.
Scenario
A node publishes DKG documents (ThresholdShareCreated, EncryptionKeyCreated, DecryptionKeyShared) for one E3, or for several concurrent E3s. Its 65th distinct record to a given peer is acknowledged and dropped there. The publisher reports success and does not retry. Committee members can then fetch the document only from the publisher's own store. If the publisher is offline or not among the peers a GET reaches, the document is missing and the DKG waits for its deadline.
Constraints for a fix
- Kademlia also provides peer discovery. Changing its protocol or record format while v0.17.0 or v0.18.0 nodes run could split the network and stop gossip, including decryption shares, from reaching the aggregator.
- A new application-level storage acknowledgement is a protocol addition that old nodes would not answer. Add it only when every node understands it, and keep the current behaviour as the fallback.
- Documents are used only during DKG. E3-2 finished its DKG, so this does not affect E3-2's decryption, but keep Kademlia unchanged until E3-2 completes.
Suggested direction
Local changes that keep the protocol as it is:
- When the store is full, evict expired or oldest records instead of rejecting new ones.
- Size the per-peer quota to the number of documents a node publishes for its concurrent E3s.
- After a PUT, read the record back through a GET that excludes the local store, and publish it again if it is missing.
Found in an adversarial review of #2046 (finding 6) and verified against the code.
Problem
A DHT PUT counts as successful when one peer acknowledges it, but a peer acknowledges records that it then drops.
StoreInserts::FilterBoth,crates/net/src/net_interface.rs:744). It rejects a record, with only a debug log, when the sender already has 64 records in its store (DHT_MAX_RECORDS_PER_PEER), when the store holds 1024 records (DHT_MAX_RECORDS), or when the size, expiry, or key check fails (net_interface.rs:982-1030).PutRecordResto the sender whether or not the application keeps the record (libp2p-kad-0.48.0/src/behaviour.rs:1931-1945).Quorum::One(net_interface.rs:1702), so one acknowledgement from a peer that dropped the record completes the PUT.Scenario
A node publishes DKG documents (
ThresholdShareCreated,EncryptionKeyCreated,DecryptionKeyShared) for one E3, or for several concurrent E3s. Its 65th distinct record to a given peer is acknowledged and dropped there. The publisher reports success and does not retry. Committee members can then fetch the document only from the publisher's own store. If the publisher is offline or not among the peers a GET reaches, the document is missing and the DKG waits for its deadline.Constraints for a fix
Suggested direction
Local changes that keep the protocol as it is:
Found in an adversarial review of #2046 (finding 6) and verified against the code.