Skip to content

Rules.md 4.6.2 "more than 4 accelerators per host": working-group disposition needed — 7 v3.0 orgs (and 4 v2.0 rows) ran exactly 4 per host #866

Description

@FileSystemGuy

Source: mlcommons/storage PR #862, validator run over the frozen v3.0 submissions tree with main @ 21f24c7.

The question

Rules.md §4.6.2 reads:

For CLOSED submissions, submitters may adjust the number of simulated accelerators per host, as long as each host uses more than 4 simulated accelerators and the total number of simulated accelerators (the total number of processes) matches the requirement. (see table 2)

The submission checker's rule function implements the text literally: num_accelerators / num_hosts <= 4 is an ERROR. That function was dormant from the day it was written (division compared lowercase-only, #842) and woke up in PR #862. On the v3.0 tree it fails eight CLOSED checkpoint workload directories from seven organizations, seven of them at exactly 4.00 per host. Those rows were reviewed and published in v3.0 without anyone applying the rule, because the tool never applied it.

The working group needs to decide which of the two readings is the rule:

  • (A) "at least 4" — 4 per host is legal. One word changes in Rules.md; the comparison in the checker flips from <= 4 to < 4. Seven of the eight findings disappear.
  • (B) "more than 4" — 4 per host is illegal. Rules.md and the checker stay as they are. Seven published v3.0 rows are retroactively non-compliant (v3.0 is frozen at tag v3.0, so nothing published changes, but any re-validation of that tree reports them), and the v4.0 rule needs to be communicated to those seven submitters.

What the v3.0 tree shows

Org System Model Hosts Accelerators Per host
Azure AMLFS-Durable-Premium-125-160TiB-2-Standard_E192is_v6 llama3-8b 2 8 4.00
Azure AMLFS-Durable-Premium-125-1280TiB-16-Standard_E192is_v6 llama3-70b 16 64 4.00
Azure AMLFS-Durable-Premium-125-4096TiB-128-Standard_E104is_v5 llama3-405b 128 512 4.00
HPE e2000 llama3-70b 16 64 4.00
UBIX UbiPower18000 … 16hosts256GBMem llama3-70b 16 64 4.00
TuringData TuringData_F9200_2_Clients llama3-8b 2 8 4.00
YanRongTech YanRongTech_F9000X_2_Clients llama3-8b 2 8 4.00
XSKY XSKY_AIMesh_3StorageNode_8Client llama3-8b 8 8 1.00

(The validator prints 15 lines rather than 8 because split write/read invocations are each checked; UBIX ran a single combined invocation.)

Distribution across all CLOSED checkpoint workload directories on the tree:

Round Workload dirs > 4 per host = 4 per host < 4 per host
v3.0 59 51 7 1 (XSKY, 1.00)
v2.0 47 43 4 0

The v2.0 rows at exactly 4 per host were IBM BlueVela (8B and 70B), UBIX (70B) and YanRongTech (8B). They were reviewed and published under the same "more than 4" wording, which first appeared in the v2.0 submission guidelines:

… they may adjust the number of GPUs per host, as long as each host uses more than 4 GPUs. This allows the use of nodes with higher GPU density and fewer total nodes.

So in two consecutive rounds, 11 CLOSED rows at 4 per host were accepted. Whatever the text says, the practised rule has been "at least 4".

Why the floor exists

The v2.0 guideline sentence gives the intent: the default is 8 simulated accelerators per host (8B = 1 host × 8, 70B = 8 × 8, 405B = 64 × 8, 1T = 128 × 8), and the rule lets submitters use denser nodes and fewer hosts. The floor stops a submitter from going the other way: spreading a fixed process count across many thin hosts to multiply client-side NICs and page cache. "More than 4" versus "at least 4" is the difference between allowing at most 2× the default host count (4 per host) and forbidding it. Nothing in either document argues that 4 per host is a materially different storage workload from 5 per host; the process count, checkpoint size and per-process I/O are fixed by Table 2 either way.

XSKY is a separate case

XSKY ran the 8B model as 8 hosts × 1 process. That is 8× the default host count and is excluded under either reading. It is the only row on the tree that the rule would still fail under option A. Whether v3.0's frozen row needs any annotation is a review-chairs question, not a tool question; for v4.0 it is simply a rule violation the tool now catches at validation time.

Recommendation

Adopt (A), "at least 4". It matches two rounds of accepted practice, it matches the rule's stated purpose (bound the host count from above, not the density from below at a value nobody has defended), and it turns the tool's newly-woken finding into a one-word Rules.md change plus a one-character checker change instead of seven retroactive non-compliances. The XSKY row stays flagged under either option.

If (A) is adopted, the tool change is:

  • Rules.md §4.6.2: "more than 4" → "at least 4".
  • mlpstorage_py/submission_checker/checks/checkpointing_checks.py, closed_accelerators_per_host: if accelerators_per_host <= 4: → if accelerators_per_host < 4: and the message must be > 4 → must be >= 4; the unit test's boundary case moves with it.
  • Validator diff over the v3.0 tree: 13 of the 15 4.6.2 lines vanish; XSKY's 2 (write and read invocations) remain with the new wording. No other rule is affected.

If (B) is adopted, no tool change; the seven submitters should be told before the v4.0 window opens, and §4.6.2 could gain a sentence stating the floor's purpose so it is not re-litigated.

Two companion items from the same PR

These are unrelated to the 4.6.2 question but were surfaced by the same rule wake-up and also need a working-group answer.

  1. checkpoint.checkpoint_folder is in the tool's CLOSED training allow-list but not in the Rules.md §3.6.2 table. It has been accepted at launch time since the rules engine landed (2026-01-15) and 12 published CLOSED unet3d runs set it. It is a path parameter of the same class as dataset.data_folder, which is in the table. Either §3.6.2's table gains a row for it, or the tool drops it from the allow-list (and those 12 runs become §3.6.2 failures on re-validation). Recommendation: add the row.

  2. Everpure FBEXA 32-host llama3-1t holds two complete CLOSED combined invocations (20260715_131314 and 20260723_093316, eight days apart, each 10 writes + 10 reads at 1024 processes). §4.7.1 / §2.1.23 allow one combined run or a write-then-read pair, so the tree now reports an ERROR there. The published v3.0 row used one of the two; the tree does not say which. A review-chairs note on which invocation the published number came from would close it; for v4.0 the validator rejects the shape at submission time.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions