Skip to content

keyshare: expulsion, failover and error-handling gaps in DKG and aggregation #2065

Description

@hmzakhalid

Summary

This issue lists liveness and recovery defects in the keyshare and aggregator actors. Each item was checked against main at 29fbd7152 on 2026-09-30. Unless noted, the same code is in v0.17.0 and v0.18.0.

Expulsions

  • New collectors ignore expulsions that happened before they were created (medium).
    • What happens:
      • The share collector (crates/keyshare/src/threshold_keyshare/effects/coordinate_collectors.rs:34) and the encryption-key collector (:109) start with every party still to do.
      • Expulsions go only to collectors that already exist (:220). A collector created after an expulsion therefore waits for the expelled party until its cutoff.
      • Rebuilt share collectors are seeded since fix(keyshare): recover early and restarted DKG inputs [skip-line-limit] #2058 (effects/recovery.rs:306).
    • Fix: seed saved expulsions in both creation helpers, before any input is delivered.
  • An expelled dealer blocks growth of a saved share batch (medium).
    • What happens:
      • Expansion removes expelled parties from the available inputs (effects/recovery.rs:188). But it compares them with the saved batch, which still includes those parties (:194): available.len() <= current.len() || !current.is_subset(&available).
      • Once a dealer in the batch is expelled, no later batch can pass, and the node stays one or more verified dealers short.
      • The same unpruned comparison is at :110.
    • Fix: compare against the saved batch's non-expelled dealers.
  • Ready-extension checks still count expelled dealers (medium).
    • What happens:
      • Local Ready construction excludes expelled dealers (effects/coordinate_roster.rs:50), but the extension check compares against the whole previously published Ready (:87).
      • The check is candidate.len() > existing.len() && existing.iter().all(|d| candidate.contains(d)) (:499).
      • The same check applies to authenticated remote Ready updates (:194).
      • A Ready that replaces an expelled dealer with a newly verified one is rejected, so roster support can stall (roster.rs:95).

Aggregator failover

  • Public-key failover throws away work that has already started (medium).
    • What happens:
      • Failover deadlines come from assignment and phase times, and ignore worker progress. The deadline is 10 minutes (crates/sortition/src/ciphernode_selection/actor.rs:89; failover.rs:148, :170, :234).
      • A demoted public-key aggregator drops its verification, signing and compute results (crates/aggregator/src/public_key_aggregation/handlers.rs:202, :221, :259).
      • The EVM writer also rejects its result, including a submission that is already queued (crates/evm/src/ciphernode_registry/handlers.rs:513, :533).
      • A slow but healthy aggregator is thus replaced, and the next one starts from scratch.
    • Existing fix to copy: fix(node): address e3-2 review and audit findings [skip-line-limit] #2043 fixed this for plaintext aggregation (plaintext_aggregation/actor.rs:208). The public-key path needs the same treatment in the actor and in both writer gates.

Error handling

  • Six keyshare Results are discarded (medium).
    • What happens:
      • effects/route_events.rs:23 drops the result of saving the aggregated public key and the decryption domain.
      • :34, :71, :75 and :78 drop the share, encryption-key, C1-signing and DKG-proof handler results.
      • coordinate_collectors.rs:205 drops the result of saving the expelled and honest parties.
    • Why it matters: a rejected snapshot write (crates/data/src/persistable.rs:174-231) also skips the in-memory update. Decryption later needs the saved public key and domain (create_decryption_share.rs:140).
    • Fix: route these through trap with their event context, and make intentional no-ops return Ok explicitly.

Restart

  • A restart can re-arm decryption-share re-sends for ended E3s (low, main only).
    • What happens:
      • After a restart, shares returned by the history query go to remember_decryption_share (crates/net/src/network_sync/effects/rebroadcast.rs:192). That function does not check whether the E3 has ended (:38).
      • Each share gets a fresh lifetime (:77) of up to 8 hours (actor.rs:82).
      • A later canonical terminal event can cancel it again, so the node may gossip old shares for hours.
    • Fix: gate the immediate re-send and the scheduling on the E3's terminal state.

No activity

Activity on this issue will appear here.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugSomething isn't workingciphernodeRelated to the ciphernode package

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions