What happened
Rolling the six prod masters to be86f49 stalled all merge-wal and compaction sweeps for 36 minutes (19:19 -> 20:15 UTC).
The first master to take the stats-writer coordination lock ran the untimed first maintain_stats pass (scanner.rs: "running first stats maintenance pass without a timeout"). The _stats table had accumulated ~68k versions since the previous restart, so the pass deleted ~31k data files, ~8k deletion files and ~32k manifests at ~30 objects/s against ADLS. It held state.stats (tokio Mutex) and the stats-writer lock the whole time.
Meanwhile on the other five masters:
- 11
Compact tasks sat in running for 26+ minutes inside update_stats_after_compaction, waiting on coordination_lock("stats-writer") (which spins forever).
- Both sweeps stopped: they call
state.stats.lock() to read candidates, and the lock was held by the maintenance pass.
- Masters logged nothing for 40 minutes; CPU ~idle.
It ended only because ADLS returned a 500 on a bulk delete (stats maintenance failed error=... Bulk delete request failed ...), which released the lock. Sweeps resumed within 4 seconds. Since the pass did not complete, the next master restart will replay it.
Worker side was fine: pending generations grew (823 across 325 shards on one worker) but the merge memory budget kept RSS flat at ~3 GiB.
Why it matters
Restarting masters is routine (image bumps, node drains). Each restart silently pauses the control plane for as long as the stats table backlog takes to delete, and the pause scales with time since the last successful pass. On a bad ADLS day that is an hour of no merges, and the only signal is that master logs go quiet.
Proposed fix
- Do not hold
state.stats across the object-store deletes. Dataset::cleanup_old_versions only needs the dataset handle. Read what is needed under the lock, drop it, run the cleanup, then re-lock to record the new version. Concurrent stats readers (sweeps, get_experiment) are unaffected by a cleanup of old versions.
- Bound the first pass too. A 36-minute untimed pass is not better than a timed one that reclaims in chunks. Use the same
MAINTENANCE_TIMEOUT and let subsequent passes finish the backlog, or cap by object count per pass.
update_stats_after_compaction should use try_coordination_lock with a short deadline, and on failure just skip the stats refresh: the task's real work (the compaction) is already committed, and the row will be refreshed by the next stats scan anyway. Today a stats-writer stall pins one TASK_CONCURRENCY slot per in-flight compaction on every master.
- Optional: reduce
_stats version churn at the source, e.g. batch upserts per scan instead of one version per row, so the backlog does not reach 68k versions in 42 hours.
Observed on
lance-context be86f49, 6 masters, ADLS backend, _stats.rollout.lance with ~10.4k hot experiments.
What happened
Rolling the six prod masters to
be86f49stalled all merge-wal and compaction sweeps for 36 minutes (19:19 -> 20:15 UTC).The first master to take the
stats-writercoordination lock ran the untimed firstmaintain_statspass (scanner.rs: "running first stats maintenance pass without a timeout"). The_statstable had accumulated ~68k versions since the previous restart, so the pass deleted ~31k data files, ~8k deletion files and ~32k manifests at ~30 objects/s against ADLS. It heldstate.stats(tokio Mutex) and thestats-writerlock the whole time.Meanwhile on the other five masters:
Compacttasks sat inrunningfor 26+ minutes insideupdate_stats_after_compaction, waiting oncoordination_lock("stats-writer")(which spins forever).state.stats.lock()to read candidates, and the lock was held by the maintenance pass.It ended only because ADLS returned a 500 on a bulk delete (
stats maintenance failed error=... Bulk delete request failed ...), which released the lock. Sweeps resumed within 4 seconds. Since the pass did not complete, the next master restart will replay it.Worker side was fine: pending generations grew (823 across 325 shards on one worker) but the merge memory budget kept RSS flat at ~3 GiB.
Why it matters
Restarting masters is routine (image bumps, node drains). Each restart silently pauses the control plane for as long as the stats table backlog takes to delete, and the pause scales with time since the last successful pass. On a bad ADLS day that is an hour of no merges, and the only signal is that master logs go quiet.
Proposed fix
state.statsacross the object-store deletes.Dataset::cleanup_old_versionsonly needs the dataset handle. Read what is needed under the lock, drop it, run the cleanup, then re-lock to record the new version. Concurrent stats readers (sweeps,get_experiment) are unaffected by a cleanup of old versions.MAINTENANCE_TIMEOUTand let subsequent passes finish the backlog, or cap by object count per pass.update_stats_after_compactionshould usetry_coordination_lockwith a short deadline, and on failure just skip the stats refresh: the task's real work (the compaction) is already committed, and the row will be refreshed by the next stats scan anyway. Today a stats-writer stall pins oneTASK_CONCURRENCYslot per in-flight compaction on every master._statsversion churn at the source, e.g. batch upserts per scan instead of one version per row, so the backlog does not reach 68k versions in 42 hours.Observed on
lance-context
be86f49, 6 masters, ADLS backend,_stats.rollout.lancewith ~10.4k hot experiments.