When s3dlio writes large objects to the S3 backend (NVIDIA AIStore) using its
MultipartUploadWriter (the default path for objects ≥ S3DLIO_MULTIPART_THRESHOLD_MB,
16 MiB), a small fraction of objects come back truncated: the object's recorded
size metadata is the full intended size, but the actual stored/retrievable data is
short by one or more whole 16-MiB parts. The completed multipart object is missing
parts, yet reports the full size.
Reads of these objects then fail — the S3 GET sends Content-Length = full size but
the body ends early:
RuntimeError: S3 GET for s3://.../img_09401_of_56000.npz failed:
dispatch failure (I/O error: error decoding response body: error reading a body from connection)
During an actual training run this surfaces as end of file before message length reached and aborts the run — a single unreadable sample is fatal.
I strongly belive there is a race somewhere in client code which is causing this issue.
Severity
Data-integrity / silent corruption. Rate is low (~0.016%) but the failure is
persistent (retries fail) and fatal to a run (any one bad sample kills the
benchmark). It only manifests at scale/concurrency, so it can pass small smoke tests
and then fail a full-size run.
Environment
- s3dlio
0.9.102 (MultipartUploadWriter, 16-MiB parts)
- mlpstorage
3.0.25, DLIO mlcommons/DLIO_local_changes rev 252a54b
- Backend: NVIDIA AIStore S3 endpoint (proxy
:51080, 307-redirect to targets :51081)
- 56000-object unet3d dataset, generated with NP=32 writer processes, single host
Reproduction
-
Generate a large dataset with default multipart (do not set
S3DLIO_MULTIPART_THRESHOLD_MB), high writer concurrency:
mlpstorage open training unet3d datagen object \
--num-processes 32 --data-dir data/mp-dg \
--params dataset.num_files_train=56000 \
--hosts <host> --dlio-bin-path <venv>/bin \
--systemname <name> --results-dir <init'd-dir> --skip-validation
-
Verify every object — compare s3dlio.stat (S3 HEAD) vs a full s3dlio.get:
head = s3dlio.stat(uri).size # full intended size
got = len(s3dlio.get(uri)) # raises on the broken objects
-
For each failure, cross-check the stored bytes natively (bypasses the S3 GET path):
ais get ais://<bucket>/<key> /tmp/o.bin # then stat -c %s /tmp/o.bin
ais ls ais://<bucket>/<key> --props size # AIStore's own recorded size
Observed (this run): 9 of 56000 objects truncated
For every failing object, AIStore's own metadata size > its own native ais get
byte count, and the stored byte count is an exact multiple of 16 MiB (the
multipart part size):
| object |
ais ls size (metadata) |
native ais get bytes |
= 16-MiB parts |
| img_09401 |
189.53 MiB |
150994944 (144 MiB) |
9 |
| img_11156 |
207.68 MiB |
83886080 (80 MiB) |
5 |
| img_16246 |
65.11 MiB |
33554432 (32 MiB) |
2 |
| img_16935 |
251.86 MiB |
50331648 (48 MiB) |
3 |
| img_20780 |
107.06 MiB |
67108864 (64 MiB) |
4 |
| img_34725 |
155.97 MiB |
33554432 (32 MiB) |
2 |
| img_35032 |
152.59 MiB |
134217728 (128 MiB) |
8 |
| img_37012 |
113.08 MiB |
33554432 (32 MiB) |
2 |
| img_39774 |
130.28 MiB |
33554432 (32 MiB) |
2 |
The stored data is always a whole number of 16-MiB parts, fewer than the object
needs → one or more multipart parts were lost, but CompleteMultipartUpload
recorded the full size.
Analysis — which layer?
The 16-MiB-part alignment is the decisive clue: this is multipart part loss, not
arbitrary byte truncation. Two narrowing facts:
- Not the client buffer handling, not low-scale transport. A standalone s3dlio
stress test — 5000 × 140-MiB objects via the exact
with MultipartUploadWriter.from_uri(uri) as w: w.write(buf.getbuffer()) pattern,
in both fresh-buffer and reused-BytesIO+zero-copy-memoryview modes (the
latter mirrors DLIO's obj_store_lib.py:630/641) — was 100% clean (0 mismatch,
0 errors). So it is scale/concurrency-dependent, not a deterministic write-path
defect.
- AIStore's own metadata is internally inconsistent with its own stored data
(ais ls size > ais get bytes), in whole-part increments.
That leaves two candidates, to be separated by capturing the S3 wire exchange
(UploadPart count + the CompleteMultipartUpload part list vs the bytes AIStore
actually persists):
- s3dlio issues
CompleteMultipartUpload listing a part whose UploadPart did
not durably succeed (e.g. a dropped/under-retried part under high concurrency).
- AIStore acks an
UploadPart or accepts CompleteMultipartUpload without
durably storing every part under concurrent load.
Either way the result is an object whose recorded size exceeds its stored data.
Workaround (verified clean)
Force single-PUT uploads by raising the multipart threshold above the largest
object, disabling multipart entirely:
S3DLIO_MULTIPART_THRESHOLD_MB=4096 # > max object size → s3dlio.put_bytes(), no multipart
Regenerating the same 56000-object dataset this way produced 0 truncated objects,
and the training run then read cleanly at NIC line rate.
Question for the WG — is single-PUT permitted for CLOSED?
S3DLIO_MULTIPART_THRESHOLD_MB is a storage-library environment variable, not a
DLIO workload parameter. The mlpstorage rules engine
(CLOSED_ALLOWED_PARAMS, mlpstorage_py/rules/run_checkers/training.py) only governs
DLIO --params overrides, so it neither tracks nor blocks this setting today.
- Training: the dataset is produced in a separate datagen step; the upload
method there does not affect the measured (read) benchmark, so single-PUT datagen
appears unobjectionable.
- Checkpointing: the write path is the measured metric, and multipart vs
single-PUT materially changes write behavior/throughput. Is forcing single-PUT
(non-multipart) writes permissible in a CLOSED checkpointing submission?
Related
When
s3dliowrites large objects to the S3 backend (NVIDIA AIStore) using itsMultipartUploadWriter(the default path for objects ≥S3DLIO_MULTIPART_THRESHOLD_MB,16 MiB), a small fraction of objects come back truncated: the object's recorded
size metadata is the full intended size, but the actual stored/retrievable data is
short by one or more whole 16-MiB parts. The completed multipart object is missing
parts, yet reports the full size.
Reads of these objects then fail — the S3 GET sends
Content-Length= full size butthe body ends early:
During an actual training run this surfaces as
end of file before message length reachedand aborts the run — a single unreadable sample is fatal.I strongly belive there is a race somewhere in client code which is causing this issue.
Severity
Data-integrity / silent corruption. Rate is low (~0.016%) but the failure is
persistent (retries fail) and fatal to a run (any one bad sample kills the
benchmark). It only manifests at scale/concurrency, so it can pass small smoke tests
and then fail a full-size run.
Environment
0.9.102(MultipartUploadWriter, 16-MiB parts)3.0.25, DLIOmlcommons/DLIO_local_changesrev252a54b:51080, 307-redirect to targets:51081)Reproduction
Generate a large dataset with default multipart (do not set
S3DLIO_MULTIPART_THRESHOLD_MB), high writer concurrency:Verify every object — compare
s3dlio.stat(S3 HEAD) vs a fulls3dlio.get:For each failure, cross-check the stored bytes natively (bypasses the S3 GET path):
Observed (this run): 9 of 56000 objects truncated
For every failing object, AIStore's own metadata size > its own native
ais getbyte count, and the stored byte count is an exact multiple of 16 MiB (the
multipart part size):
ais lssize (metadata)ais getbytesThe stored data is always a whole number of 16-MiB parts, fewer than the object
needs → one or more multipart parts were lost, but
CompleteMultipartUploadrecorded the full size.
Analysis — which layer?
The 16-MiB-part alignment is the decisive clue: this is multipart part loss, not
arbitrary byte truncation. Two narrowing facts:
stress test — 5000 × 140-MiB objects via the exact
with MultipartUploadWriter.from_uri(uri) as w: w.write(buf.getbuffer())pattern,in both fresh-buffer and reused-
BytesIO+zero-copy-memoryviewmodes (thelatter mirrors DLIO's
obj_store_lib.py:630/641) — was 100% clean (0 mismatch,0 errors). So it is scale/concurrency-dependent, not a deterministic write-path
defect.
(
ais lssize >ais getbytes), in whole-part increments.That leaves two candidates, to be separated by capturing the S3 wire exchange
(
UploadPartcount + theCompleteMultipartUploadpart list vs the bytes AIStoreactually persists):
CompleteMultipartUploadlisting a part whoseUploadPartdidnot durably succeed (e.g. a dropped/under-retried part under high concurrency).
UploadPartor acceptsCompleteMultipartUploadwithoutdurably storing every part under concurrent load.
Either way the result is an object whose recorded size exceeds its stored data.
Workaround (verified clean)
Force single-PUT uploads by raising the multipart threshold above the largest
object, disabling multipart entirely:
S3DLIO_MULTIPART_THRESHOLD_MB=4096 # > max object size → s3dlio.put_bytes(), no multipartRegenerating the same 56000-object dataset this way produced 0 truncated objects,
and the training run then read cleanly at NIC line rate.
Question for the WG — is single-PUT permitted for CLOSED?
S3DLIO_MULTIPART_THRESHOLD_MBis a storage-library environment variable, not aDLIO workload parameter. The mlpstorage rules engine
(
CLOSED_ALLOWED_PARAMS,mlpstorage_py/rules/run_checkers/training.py) only governsDLIO
--paramsoverrides, so it neither tracks nor blocks this setting today.method there does not affect the measured (read) benchmark, so single-PUT datagen
appears unobjectionable.
single-PUT materially changes write behavior/throughput. Is forcing single-PUT
(non-multipart) writes permissible in a CLOSED checkpointing submission?
Related
mlpstorage_py.../storage/obj_store_lib.py:621-644— theput_datamultipart path(
_MULTIPART_THRESHOLD,MultipartUploadWriter.from_uri(...).write(payload)).