Skip to content

master: first-pass stats maintenance blocks every sweep and compaction for its whole duration #271

Description

@beinan

What happened

Rolling the six prod masters to be86f49 stalled all merge-wal and compaction sweeps for 36 minutes (19:19 -> 20:15 UTC).

The first master to take the stats-writer coordination lock ran the untimed first maintain_stats pass (scanner.rs: "running first stats maintenance pass without a timeout"). The _stats table had accumulated ~68k versions since the previous restart, so the pass deleted ~31k data files, ~8k deletion files and ~32k manifests at ~30 objects/s against ADLS. It held state.stats (tokio Mutex) and the stats-writer lock the whole time.

Meanwhile on the other five masters:

  • 11 Compact tasks sat in running for 26+ minutes inside update_stats_after_compaction, waiting on coordination_lock("stats-writer") (which spins forever).
  • Both sweeps stopped: they call state.stats.lock() to read candidates, and the lock was held by the maintenance pass.
  • Masters logged nothing for 40 minutes; CPU ~idle.

It ended only because ADLS returned a 500 on a bulk delete (stats maintenance failed error=... Bulk delete request failed ...), which released the lock. Sweeps resumed within 4 seconds. Since the pass did not complete, the next master restart will replay it.

Worker side was fine: pending generations grew (823 across 325 shards on one worker) but the merge memory budget kept RSS flat at ~3 GiB.

Why it matters

Restarting masters is routine (image bumps, node drains). Each restart silently pauses the control plane for as long as the stats table backlog takes to delete, and the pause scales with time since the last successful pass. On a bad ADLS day that is an hour of no merges, and the only signal is that master logs go quiet.

Proposed fix

  1. Do not hold state.stats across the object-store deletes. Dataset::cleanup_old_versions only needs the dataset handle. Read what is needed under the lock, drop it, run the cleanup, then re-lock to record the new version. Concurrent stats readers (sweeps, get_experiment) are unaffected by a cleanup of old versions.
  2. Bound the first pass too. A 36-minute untimed pass is not better than a timed one that reclaims in chunks. Use the same MAINTENANCE_TIMEOUT and let subsequent passes finish the backlog, or cap by object count per pass.
  3. update_stats_after_compaction should use try_coordination_lock with a short deadline, and on failure just skip the stats refresh: the task's real work (the compaction) is already committed, and the row will be refreshed by the next stats scan anyway. Today a stats-writer stall pins one TASK_CONCURRENCY slot per in-flight compaction on every master.
  4. Optional: reduce _stats version churn at the source, e.g. batch upserts per scan instead of one version per row, so the backlog does not reach 68k versions in 42 hours.

Observed on

lance-context be86f49, 6 masters, ADLS backend, _stats.rollout.lance with ~10.4k hot experiments.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions