Repository navigation
fix(storage#741): drop page cache before read_threads memory guard - #48
Merged
FileSystemGuy merged 1 commit intoJul 9, 2026
Merged
Conversation
The guard in `ConfigArguments.validate` reads `psutil.virtual_memory().available` and caps `read_threads` against it. On POSIX filesystems whose client pins a large reclaimable page cache (Lustre in particular), `MemAvailable` understates once a prior run has warmed the cache — a legal config that passed the guard on run 1 will trip it on runs 2..N of a 5-run submission, aborting the second run in `initialize()` before the per-epoch flush ever gets to run. reportgen then sees `INVALID: 0 runs` for the submission even though every rejected run's memory pressure came from evictable cache the benchmark drops every epoch anyway. Fix: right before sampling `virtual_memory()`, drop the local host's page cache via the same `sudo -n sh -c 'echo 3 > /proc/sys/vm/drop_caches'` command already used per-epoch in `main.py`. Gate on `local_rank() == 0` so exactly one flush runs per host, then barrier so non-leaders read post-drop memory. Why `local_rank()` not `MPI.node()`: `MPI.node()` is unreliable under `--map-by node` (the same anti-pattern PR mlcommons/storage#675 fixed for the `host_memory_GB` collector array — two hosts' local-rank-0 can compute the same node index). `local_rank()` derives from `MPI.COMM_TYPE_SHARED` which is independent of rank assignment. Worst-case failure mode is redundant (idempotent) flushes, never a missed flush on a host that needed one. Why this doesn't need SSH: the added flush runs inside the DLIO process, which is already up on each host via whatever launcher mlpstorage used (OpenMPI, PALS palsd, Slurm slurmstepd). We piggyback on the transport that already worked — no new privilege or reachability requirement beyond the passwordless-sudo already assumed by the per-epoch flush. Fail-open: every subprocess or barrier failure is swallowed so the guard still runs. Matches the per-epoch flush's posture; the worst case is the guard trips on this host, which is current behavior. Timeout is the existing `DLIO_DROP_CACHES_TIMEOUT` env var (default 30s), so operators keep one knob for both the pre-guard and per-epoch flushes.
FileSystemGuy
added a commit
to mlcommons/storage
that referenced
this pull request
Jul 9, 2026
Picks up DLIO_local_changes PR #48 (`fix(storage#741): drop page cache before read_threads memory guard`), which was merged after the previous storage bump PR (#746) had already been opened against DLIO main HEAD 86945a7a. What #48 fixes (from the DLIO PR body / storage#741): The read_threads per-node memory guard in ConfigArguments.validate samples psutil.virtual_memory().available and caps read_threads against it. On POSIX filesystems whose client pins a large reclaimable page cache (Lustre in particular), MemAvailable understates by the size of that cache — so after run 1 warms the cache, runs 2..N of a 5-run submission see the cache as "unavailable" and DLIO refuses to launch in initialize(), before the per-epoch flush ever gets to run. reportgen then reports INVALID: 0 runs for a submission whose only real problem is that the benchmark's own per-epoch flush hadn't run yet. The fix drops the local host's page cache right before virtual_memory() is sampled, gated on local_rank() == 0 (not MPI.node(), which is unreliable under --map-by node) with a Barrier so non-leaders read post-drop memory. Reuses the existing DLIO_DROP_CACHES_TIMEOUT env var — no new operator knob. Not bumping mlpstorage version: this is a pure DLIO-side pin advance with no mlp-storage code change. The version bump already went out with #746 (3.0.39 -> 3.0.40); #746 was opened before #48 merged, so this follow-up captures #48 into the same 3.0.40 line before it ships. VERIFICATION ============ $ uv lock --upgrade-package dlio-benchmark Resolved 108 packages in 1.33s Updated dlio-benchmark v3.0.2 (86945a7a) -> v3.0.2 (95c6a9d4) $ uv sync --active && uv pip install -e ./vdb_benchmark -e ./kv_cache_benchmark $ uv run pytest tests/unit -q # 2689 passed, 1 skipped $ uv run pytest mlpstorage_py/tests -q # 840 passed $ uv run pytest vdb_benchmark/tests -q # 174 passed $ uv run pytest kv_cache_benchmark/tests -q # 238 passed Total: 3941 passed, 1 skipped across all four CI suites. Refs: #741 mlcommons/DLIO_local_changes#48
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Problem
ConfigArguments.validateinutils/config.pysamplespsutil.virtual_memory().availableand capsread_threadsagainst it(the mlcommons/storage#448 per-node budget fix). On POSIX filesystems
whose client pins a large reclaimable page cache (Lustre in
particular),
MemAvailableunderstates by the size of thatcache — so after run 1 warms the cache, runs 2..N of a 5-run
submission see the cache as "unavailable" and DLIO refuses to launch:
This aborts
initialize()before the per-epoch page-cache flushever gets to run. reportgen then reports
INVALID: 0 runsfor asubmission whose only real problem is that the benchmark's own
per-epoch flush (which drops that exact cache every epoch) hadn't run
yet. Reporter's evidence: run 1 passes at ~170 GB MemAvailable; runs
2–5 crash at ~147–150 GB;
sync; echo 1 > drop_cachesrestores ~174GB and the guard passes again — confirming the "used" RAM is
reclaimable cache.
Filed as mlcommons/storage#741 (
@mirajeev).Fix
Right before sampling
virtual_memory(), drop the local host's pagecache using the same
sudo -n sh -c 'echo 3 > /proc/sys/vm/drop_caches'command already invoked per-epoch in
main.py. Gated onlocal_rank() == 0so exactly one flush runs per host, followed by acomm().Barrier()so non-leader ranks read post-drop memory.Design choices worth flagging on review
local_rank()notMPI.node().MPI.node()is computed fromcumulative offsets of
mpi_ppn_list(utility.py:313–318) and assumescontiguous per-node rank blocks. Under
--map-by node(whichmlpstorage sets by default for multi-host) that assumption is
violated: rank 0 → host 0, rank 1 → host 1, …, rank N → host 0, and
two different physical hosts end up computing the same node index.
This is the same anti-pattern PR mlcommons/storage#675 fixed for the
host_memory_GBcollector array (the[376.18 × 7, 187.90, 0.0 × 7]shape).
local_rank()derives fromMPI.COMM_TYPE_SHARED— theMPI-3 primitive that groups ranks by shared memory — and is
independent of rank assignment. Worst-case failure mode is
redundant/idempotent flushes on a host, never a missed flush on a host
that needed one.
No SSH dependency. The flush runs inside the DLIO process, which
is already up on each host via whatever launcher was used (OpenMPI,
PALS
palsd, Slurmslurmstepd). We piggyback on the transport thatalready worked. In particular this is compatible with the
--skip-ssh-check/ PALS/Slurm auto-skip path inmlcommons/storage#740 — those users need this fix precisely because
their sites disallow passwordless SSH between compute nodes, so any
mlpstorage-side fan-out via
mpirunwould have been unusable forthem.
Fail-open matches per-epoch flush. Every failure mode —
sudo -nrefused, kernel timeout,
Barrier()error,subprocessOSError—is swallowed. The guard downstream still runs; the worst case is that
it still trips on this host, which is exactly current behavior. No
regression path.
Same env-var timeout knob. Uses
_resolve_drop_caches_timeout()(30s default,
DLIO_DROP_CACHES_TIMEOUToverride). Operators alreadytune this for the per-epoch flush; introducing a second knob would
double the operational surface for no gain.
Skipped when MPI not initialized. In child processes and test
harnesses
mpi_state != MPI_INITIALIZED; the flush is skipped socomm()isn't called and the guard still runs unchanged.Test plan
tests/test_drop_caches_before_memory_guard.py:MPI-state gating,
local_rank()gate (leader + non-leaderparametrized), fail-open across every subprocess exception class
+ non-zero exit +
Barrierfailure on leader/non-leader, timeoutplumbing (default + env override + bad-env-value).
tests/test_drop_caches_timeout.py(22 tests) stillpasses — no changes to
_resolve_drop_caches_timeout.uv run pytest tests/ --ignore=<benchmark end-to-end suites>—183 passed, 2 skipped, no regressions.
uv run pytest tests/test_fast_ci.py— 92 passed, 1 skipped.Includes
TestEndToEndSmoke::test_train_npy_smokewhichexercises the modified
validate()path withread_threads > 0and PyTorch data loader.
Related
@mirajeevalongside a between-runs cache-drop workaroundcomm_size-wide) worker count → false positives at scale storage#448 (per-node vs world-comm scoping fixthis guard was originally built for)
MPI.node()unreliability motivatinglocal_rank()gate:Bug: Submission checker rule 3.1.2 double-counts host memory, doubling the required dataset size storage#669 / PR fix(#669): aggregate host_memory_GB via sum(), not num_hosts × array[0] storage#675