Source: mlcommons/storage PR #862, validator run over the frozen v3.0 submissions tree with main @ 21f24c7.
The question
Rules.md §4.6.2 reads:
For CLOSED submissions, submitters may adjust the number of simulated accelerators per host, as long as each host uses more than 4 simulated accelerators and the total number of simulated accelerators (the total number of processes) matches the requirement. (see table 2)
The submission checker's rule function implements the text literally: num_accelerators / num_hosts <= 4 is an ERROR. That function was dormant from the day it was written (division compared lowercase-only, #842) and woke up in PR #862. On the v3.0 tree it fails eight CLOSED checkpoint workload directories from seven organizations, seven of them at exactly 4.00 per host. Those rows were reviewed and published in v3.0 without anyone applying the rule, because the tool never applied it.
The working group needs to decide which of the two readings is the rule:
- (A) "at least 4" — 4 per host is legal. One word changes in Rules.md; the comparison in the checker flips from
<= 4 to < 4. Seven of the eight findings disappear.
- (B) "more than 4" — 4 per host is illegal. Rules.md and the checker stay as they are. Seven published v3.0 rows are retroactively non-compliant (v3.0 is frozen at tag
v3.0, so nothing published changes, but any re-validation of that tree reports them), and the v4.0 rule needs to be communicated to those seven submitters.
What the v3.0 tree shows
| Org |
System |
Model |
Hosts |
Accelerators |
Per host |
| Azure |
AMLFS-Durable-Premium-125-160TiB-2-Standard_E192is_v6 |
llama3-8b |
2 |
8 |
4.00 |
| Azure |
AMLFS-Durable-Premium-125-1280TiB-16-Standard_E192is_v6 |
llama3-70b |
16 |
64 |
4.00 |
| Azure |
AMLFS-Durable-Premium-125-4096TiB-128-Standard_E104is_v5 |
llama3-405b |
128 |
512 |
4.00 |
| HPE |
e2000 |
llama3-70b |
16 |
64 |
4.00 |
| UBIX |
UbiPower18000 … 16hosts256GBMem |
llama3-70b |
16 |
64 |
4.00 |
| TuringData |
TuringData_F9200_2_Clients |
llama3-8b |
2 |
8 |
4.00 |
| YanRongTech |
YanRongTech_F9000X_2_Clients |
llama3-8b |
2 |
8 |
4.00 |
| XSKY |
XSKY_AIMesh_3StorageNode_8Client |
llama3-8b |
8 |
8 |
1.00 |
(The validator prints 15 lines rather than 8 because split write/read invocations are each checked; UBIX ran a single combined invocation.)
Distribution across all CLOSED checkpoint workload directories on the tree:
| Round |
Workload dirs |
> 4 per host |
= 4 per host |
< 4 per host |
| v3.0 |
59 |
51 |
7 |
1 (XSKY, 1.00) |
| v2.0 |
47 |
43 |
4 |
0 |
The v2.0 rows at exactly 4 per host were IBM BlueVela (8B and 70B), UBIX (70B) and YanRongTech (8B). They were reviewed and published under the same "more than 4" wording, which first appeared in the v2.0 submission guidelines:
… they may adjust the number of GPUs per host, as long as each host uses more than 4 GPUs. This allows the use of nodes with higher GPU density and fewer total nodes.
So in two consecutive rounds, 11 CLOSED rows at 4 per host were accepted. Whatever the text says, the practised rule has been "at least 4".
Why the floor exists
The v2.0 guideline sentence gives the intent: the default is 8 simulated accelerators per host (8B = 1 host × 8, 70B = 8 × 8, 405B = 64 × 8, 1T = 128 × 8), and the rule lets submitters use denser nodes and fewer hosts. The floor stops a submitter from going the other way: spreading a fixed process count across many thin hosts to multiply client-side NICs and page cache. "More than 4" versus "at least 4" is the difference between allowing at most 2× the default host count (4 per host) and forbidding it. Nothing in either document argues that 4 per host is a materially different storage workload from 5 per host; the process count, checkpoint size and per-process I/O are fixed by Table 2 either way.
XSKY is a separate case
XSKY ran the 8B model as 8 hosts × 1 process. That is 8× the default host count and is excluded under either reading. It is the only row on the tree that the rule would still fail under option A. Whether v3.0's frozen row needs any annotation is a review-chairs question, not a tool question; for v4.0 it is simply a rule violation the tool now catches at validation time.
Recommendation
Adopt (A), "at least 4". It matches two rounds of accepted practice, it matches the rule's stated purpose (bound the host count from above, not the density from below at a value nobody has defended), and it turns the tool's newly-woken finding into a one-word Rules.md change plus a one-character checker change instead of seven retroactive non-compliances. The XSKY row stays flagged under either option.
If (A) is adopted, the tool change is:
Rules.md §4.6.2: "more than 4" → "at least 4".
mlpstorage_py/submission_checker/checks/checkpointing_checks.py, closed_accelerators_per_host: if accelerators_per_host <= 4: → if accelerators_per_host < 4: and the message must be > 4 → must be >= 4; the unit test's boundary case moves with it.
- Validator diff over the v3.0 tree: 13 of the 15 4.6.2 lines vanish; XSKY's 2 (write and read invocations) remain with the new wording. No other rule is affected.
If (B) is adopted, no tool change; the seven submitters should be told before the v4.0 window opens, and §4.6.2 could gain a sentence stating the floor's purpose so it is not re-litigated.
Two companion items from the same PR
These are unrelated to the 4.6.2 question but were surfaced by the same rule wake-up and also need a working-group answer.
-
checkpoint.checkpoint_folder is in the tool's CLOSED training allow-list but not in the Rules.md §3.6.2 table. It has been accepted at launch time since the rules engine landed (2026-01-15) and 12 published CLOSED unet3d runs set it. It is a path parameter of the same class as dataset.data_folder, which is in the table. Either §3.6.2's table gains a row for it, or the tool drops it from the allow-list (and those 12 runs become §3.6.2 failures on re-validation). Recommendation: add the row.
-
Everpure FBEXA 32-host llama3-1t holds two complete CLOSED combined invocations (20260715_131314 and 20260723_093316, eight days apart, each 10 writes + 10 reads at 1024 processes). §4.7.1 / §2.1.23 allow one combined run or a write-then-read pair, so the tree now reports an ERROR there. The published v3.0 row used one of the two; the tree does not say which. A review-chairs note on which invocation the published number came from would close it; for v4.0 the validator rejects the shape at submission time.
Source: mlcommons/storage PR #862, validator run over the frozen v3.0 submissions tree with
main@21f24c7.The question
Rules.md §4.6.2 reads:
The submission checker's rule function implements the text literally:
num_accelerators / num_hosts <= 4is an ERROR. That function was dormant from the day it was written (division compared lowercase-only, #842) and woke up in PR #862. On the v3.0 tree it fails eight CLOSED checkpoint workload directories from seven organizations, seven of them at exactly 4.00 per host. Those rows were reviewed and published in v3.0 without anyone applying the rule, because the tool never applied it.The working group needs to decide which of the two readings is the rule:
<= 4to< 4. Seven of the eight findings disappear.v3.0, so nothing published changes, but any re-validation of that tree reports them), and the v4.0 rule needs to be communicated to those seven submitters.What the v3.0 tree shows
(The validator prints 15 lines rather than 8 because split write/read invocations are each checked; UBIX ran a single combined invocation.)
Distribution across all CLOSED checkpoint workload directories on the tree:
The v2.0 rows at exactly 4 per host were IBM BlueVela (8B and 70B), UBIX (70B) and YanRongTech (8B). They were reviewed and published under the same "more than 4" wording, which first appeared in the v2.0 submission guidelines:
So in two consecutive rounds, 11 CLOSED rows at 4 per host were accepted. Whatever the text says, the practised rule has been "at least 4".
Why the floor exists
The v2.0 guideline sentence gives the intent: the default is 8 simulated accelerators per host (8B = 1 host × 8, 70B = 8 × 8, 405B = 64 × 8, 1T = 128 × 8), and the rule lets submitters use denser nodes and fewer hosts. The floor stops a submitter from going the other way: spreading a fixed process count across many thin hosts to multiply client-side NICs and page cache. "More than 4" versus "at least 4" is the difference between allowing at most 2× the default host count (4 per host) and forbidding it. Nothing in either document argues that 4 per host is a materially different storage workload from 5 per host; the process count, checkpoint size and per-process I/O are fixed by Table 2 either way.
XSKY is a separate case
XSKY ran the 8B model as 8 hosts × 1 process. That is 8× the default host count and is excluded under either reading. It is the only row on the tree that the rule would still fail under option A. Whether v3.0's frozen row needs any annotation is a review-chairs question, not a tool question; for v4.0 it is simply a rule violation the tool now catches at validation time.
Recommendation
Adopt (A), "at least 4". It matches two rounds of accepted practice, it matches the rule's stated purpose (bound the host count from above, not the density from below at a value nobody has defended), and it turns the tool's newly-woken finding into a one-word Rules.md change plus a one-character checker change instead of seven retroactive non-compliances. The XSKY row stays flagged under either option.
If (A) is adopted, the tool change is:
Rules.md§4.6.2: "more than 4" → "at least 4".mlpstorage_py/submission_checker/checks/checkpointing_checks.py,closed_accelerators_per_host:if accelerators_per_host <= 4:→if accelerators_per_host < 4:and the messagemust be > 4→must be >= 4; the unit test's boundary case moves with it.If (B) is adopted, no tool change; the seven submitters should be told before the v4.0 window opens, and §4.6.2 could gain a sentence stating the floor's purpose so it is not re-litigated.
Two companion items from the same PR
These are unrelated to the 4.6.2 question but were surfaced by the same rule wake-up and also need a working-group answer.
checkpoint.checkpoint_folderis in the tool's CLOSED training allow-list but not in the Rules.md §3.6.2 table. It has been accepted at launch time since the rules engine landed (2026-01-15) and 12 published CLOSED unet3d runs set it. It is a path parameter of the same class asdataset.data_folder, which is in the table. Either §3.6.2's table gains a row for it, or the tool drops it from the allow-list (and those 12 runs become §3.6.2 failures on re-validation). Recommendation: add the row.Everpure FBEXA 32-host llama3-1t holds two complete CLOSED combined invocations (
20260715_131314and20260723_093316, eight days apart, each 10 writes + 10 reads at 1024 processes). §4.7.1 / §2.1.23 allow one combined run or a write-then-read pair, so the tree now reports an ERROR there. The published v3.0 row used one of the two; the tree does not say which. A review-chairs note on which invocation the published number came from would close it; for v4.0 the validator rejects the shape at submission time.