Labels: bug, training, dlio
Component: dlio_benchmark/utils/config.py → Config.validate() (per-node memory-budget guard)
Version: mlpstorage 3.0.33 (mlcommons/storage @ fe907a9)
Benchmark: training / retinanet, CLOSED, POSIX/file backend on Lustre
Related: follow-on to #448 (world-vs-local comm_size double-count)
Summary
The per-node memory-budget guard caps reader.read_threads using psutil.virtual_memory().available sampled at startup inside initialize(), before the run's first per-epoch drop_caches. On POSIX filesystems whose client pins a large reclaimable page cache (Lustre in particular), MemAvailable is understated once a prior run has warmed the cache. This causes false-positive startup failures on the 2nd..Nth of the 5+ required consecutive runs, even though (a) the cache is reclaimable so workers won't actually OOM, and (b) the benchmark already drops that cache every epoch by default. Only the first (cold) run passes; every later run in the submission crashes at startup → reportgen: INVALID: 0 runs.
Guard code (config.py, validate())
The per-epoch flush that makes the cache a non-constraint (main.py, default-on, only skipped if sudo -n fails):
Steps to reproduce
16 hosts × 16 local ranks, num_accelerators=250, read_threads=17, prefetch_size=2, 188 GB/node.
Dataset 50,535,754 files (≥ 5× aggregate host memory, so Rules.md §3.1.2 is satisfied by size).
mlpstorage closed training retinanet run file in a 6-run loop (1 warm-up + 5 measured), no cache drop between runs.
Observed
Run Startup MemAvailable Guard verdict Result
1 (cold) ~170 GB 17×16×0.5 ≈ 136 GB < budget PASS — AU 85.56%, 33,198 MB/s, 8 epochs
2–6 ~147–150 GB (page cache from run 1) demands read_threads ≤ 16 CRASH in validate()
Exact exception (runs 2–6):
(Prior run at 18 threads: ... = 288 worker processes, estimated ~144 GB (available RAM on this node: 159 GB; total: 188 GB). Reduce reader.read_threads to at most 17.)
reportgen → INVALID: 0 runs / "Training submission requires 5 runs". Running sync; echo 1 > /proc/sys/vm/drop_caches on all nodes restores MemAvailable to ~174 GB and the guard passes again — confirming the "used" RAM is reclaimable cache, not a real allocation.
Expected
A read_threads value that passes run 1 should pass all consecutive runs. The guard should not count reclaimable page cache as unavailable, especially since the benchmark itself drops that cache every epoch.
Why it's a bug
MemAvailable treats filesystem-pinned page cache (Lustre) as unavailable, but it's evictable — the guard's OOM premise doesn't hold.
The project already relies on dropping this exact cache every epoch, so counting it against the startup budget is internally inconsistent.
A legal config that passes run 1 fails runs 2..N, breaking the required 5-consecutive-run submission with no user error.
Suggested fixes (any one)
A. Best-effort drop_caches (same sudo -n as the epoch flush) once before the guard reads virtual_memory().available; degrade gracefully if unavailable.
B. Base the budget on reclaimable-inclusive free memory (MemFree + SReclaimable + page cache, or MemTotal − non-cache used) instead of raw MemAvailable.
C. Add a documented between-runs cache-drop step to mlpstorage's multi-run orchestration (the DLIO_DROP_CACHES_TIMEOUT plumbing already exists).
Workaround
Drop caches on all client nodes between consecutive mlpstorage run invocations, before each run's startup guard executes.
Environment
mlpstorage 3.0.33 @ fe907a9; 16 × ~188 GB RAM, Linux, Lustre POSIX mount, PyTorch loader; read_threads=17, ranks_per_node=16, num_accelerators=250.
I am going ahead with cache drop between run as work around
Labels: bug, training, dlio
Component: dlio_benchmark/utils/config.py → Config.validate() (per-node memory-budget guard)
Version: mlpstorage 3.0.33 (mlcommons/storage @ fe907a9)
Benchmark: training / retinanet, CLOSED, POSIX/file backend on Lustre
Related: follow-on to #448 (world-vs-local comm_size double-count)
Summary
The per-node memory-budget guard caps reader.read_threads using psutil.virtual_memory().available sampled at startup inside initialize(), before the run's first per-epoch drop_caches. On POSIX filesystems whose client pins a large reclaimable page cache (Lustre in particular), MemAvailable is understated once a prior run has warmed the cache. This causes false-positive startup failures on the 2nd..Nth of the 5+ required consecutive runs, even though (a) the cache is reclaimable so workers won't actually OOM, and (b) the benchmark already drops that cache every epoch by default. Only the first (cold) run passes; every later run in the submission crashes at startup → reportgen: INVALID: 0 runs.
Guard code (config.py, validate())
The per-epoch flush that makes the cache a non-constraint (main.py, default-on, only skipped if sudo -n fails):
Steps to reproduce
16 hosts × 16 local ranks, num_accelerators=250, read_threads=17, prefetch_size=2, 188 GB/node.
Dataset 50,535,754 files (≥ 5× aggregate host memory, so Rules.md §3.1.2 is satisfied by size).
mlpstorage closed training retinanet run file in a 6-run loop (1 warm-up + 5 measured), no cache drop between runs.
Observed
Run Startup MemAvailable Guard verdict Result
1 (cold) ~170 GB 17×16×0.5 ≈ 136 GB < budget PASS — AU 85.56%, 33,198 MB/s, 8 epochs
2–6 ~147–150 GB (page cache from run 1) demands read_threads ≤ 16 CRASH in validate()
Exact exception (runs 2–6):
(Prior run at 18 threads: ... = 288 worker processes, estimated ~144 GB (available RAM on this node: 159 GB; total: 188 GB). Reduce reader.read_threads to at most 17.)
reportgen → INVALID: 0 runs / "Training submission requires 5 runs". Running sync; echo 1 > /proc/sys/vm/drop_caches on all nodes restores MemAvailable to ~174 GB and the guard passes again — confirming the "used" RAM is reclaimable cache, not a real allocation.
Expected
A read_threads value that passes run 1 should pass all consecutive runs. The guard should not count reclaimable page cache as unavailable, especially since the benchmark itself drops that cache every epoch.
Why it's a bug
MemAvailable treats filesystem-pinned page cache (Lustre) as unavailable, but it's evictable — the guard's OOM premise doesn't hold.
The project already relies on dropping this exact cache every epoch, so counting it against the startup budget is internally inconsistent.
A legal config that passes run 1 fails runs 2..N, breaking the required 5-consecutive-run submission with no user error.
Suggested fixes (any one)
A. Best-effort drop_caches (same sudo -n as the epoch flush) once before the guard reads virtual_memory().available; degrade gracefully if unavailable.
B. Base the budget on reclaimable-inclusive free memory (MemFree + SReclaimable + page cache, or MemTotal − non-cache used) instead of raw MemAvailable.
C. Add a documented between-runs cache-drop step to mlpstorage's multi-run orchestration (the DLIO_DROP_CACHES_TIMEOUT plumbing already exists).
Workaround
Drop caches on all client nodes between consecutive mlpstorage run invocations, before each run's startup guard executes.
Environment
mlpstorage 3.0.33 @ fe907a9; 16 × ~188 GB RAM, Linux, Lustre POSIX mount, PyTorch loader; read_threads=17, ranks_per_node=16, num_accelerators=250.
I am going ahead with cache drop between run as work around