Skip to content

Multipart S3 upload silently drops whole 16-MiB parts at scale — objects are stored short (metadata full), breaking training reads #593

Description

@gaikwadabhishek

When s3dlio writes large objects to the S3 backend (NVIDIA AIStore) using its
MultipartUploadWriter (the default path for objects ≥ S3DLIO_MULTIPART_THRESHOLD_MB,
16 MiB), a small fraction of objects come back truncated: the object's recorded
size metadata is the full intended size, but the actual stored/retrievable data is
short by one or more whole 16-MiB parts
. The completed multipart object is missing
parts, yet reports the full size.

Reads of these objects then fail — the S3 GET sends Content-Length = full size but
the body ends early:

RuntimeError: S3 GET for s3://.../img_09401_of_56000.npz failed:
dispatch failure (I/O error: error decoding response body: error reading a body from connection)

During an actual training run this surfaces as end of file before message length reached and aborts the run — a single unreadable sample is fatal.

I strongly belive there is a race somewhere in client code which is causing this issue.

Severity

Data-integrity / silent corruption. Rate is low (~0.016%) but the failure is
persistent (retries fail) and fatal to a run (any one bad sample kills the
benchmark). It only manifests at scale/concurrency, so it can pass small smoke tests
and then fail a full-size run.

Environment

  • s3dlio 0.9.102 (MultipartUploadWriter, 16-MiB parts)
  • mlpstorage 3.0.25, DLIO mlcommons/DLIO_local_changes rev 252a54b
  • Backend: NVIDIA AIStore S3 endpoint (proxy :51080, 307-redirect to targets :51081)
  • 56000-object unet3d dataset, generated with NP=32 writer processes, single host

Reproduction

  1. Generate a large dataset with default multipart (do not set
    S3DLIO_MULTIPART_THRESHOLD_MB), high writer concurrency:

    mlpstorage open training unet3d datagen object \
      --num-processes 32 --data-dir data/mp-dg \
      --params dataset.num_files_train=56000 \
      --hosts <host> --dlio-bin-path <venv>/bin \
      --systemname <name> --results-dir <init'd-dir> --skip-validation
  2. Verify every object — compare s3dlio.stat (S3 HEAD) vs a full s3dlio.get:

    head = s3dlio.stat(uri).size          # full intended size
    got  = len(s3dlio.get(uri))           # raises on the broken objects
  3. For each failure, cross-check the stored bytes natively (bypasses the S3 GET path):

    ais get ais://<bucket>/<key> /tmp/o.bin   # then stat -c %s /tmp/o.bin
    ais ls  ais://<bucket>/<key> --props size # AIStore's own recorded size

Observed (this run): 9 of 56000 objects truncated

For every failing object, AIStore's own metadata size > its own native ais get
byte count
, and the stored byte count is an exact multiple of 16 MiB (the
multipart part size):

object ais ls size (metadata) native ais get bytes = 16-MiB parts
img_09401 189.53 MiB 150994944 (144 MiB) 9
img_11156 207.68 MiB 83886080 (80 MiB) 5
img_16246 65.11 MiB 33554432 (32 MiB) 2
img_16935 251.86 MiB 50331648 (48 MiB) 3
img_20780 107.06 MiB 67108864 (64 MiB) 4
img_34725 155.97 MiB 33554432 (32 MiB) 2
img_35032 152.59 MiB 134217728 (128 MiB) 8
img_37012 113.08 MiB 33554432 (32 MiB) 2
img_39774 130.28 MiB 33554432 (32 MiB) 2

The stored data is always a whole number of 16-MiB parts, fewer than the object
needs → one or more multipart parts were lost, but CompleteMultipartUpload
recorded the full size.

Analysis — which layer?

The 16-MiB-part alignment is the decisive clue: this is multipart part loss, not
arbitrary byte truncation. Two narrowing facts:

  • Not the client buffer handling, not low-scale transport. A standalone s3dlio
    stress test — 5000 × 140-MiB objects via the exact
    with MultipartUploadWriter.from_uri(uri) as w: w.write(buf.getbuffer()) pattern,
    in both fresh-buffer and reused-BytesIO+zero-copy-memoryview modes (the
    latter mirrors DLIO's obj_store_lib.py:630/641) — was 100% clean (0 mismatch,
    0 errors). So it is scale/concurrency-dependent, not a deterministic write-path
    defect.
  • AIStore's own metadata is internally inconsistent with its own stored data
    (ais ls size > ais get bytes), in whole-part increments.

That leaves two candidates, to be separated by capturing the S3 wire exchange
(UploadPart count + the CompleteMultipartUpload part list vs the bytes AIStore
actually persists):

  1. s3dlio issues CompleteMultipartUpload listing a part whose UploadPart did
    not durably succeed (e.g. a dropped/under-retried part under high concurrency).
  2. AIStore acks an UploadPart or accepts CompleteMultipartUpload without
    durably storing every part under concurrent load.

Either way the result is an object whose recorded size exceeds its stored data.

Workaround (verified clean)

Force single-PUT uploads by raising the multipart threshold above the largest
object, disabling multipart entirely:

S3DLIO_MULTIPART_THRESHOLD_MB=4096   # > max object size → s3dlio.put_bytes(), no multipart

Regenerating the same 56000-object dataset this way produced 0 truncated objects,
and the training run then read cleanly at NIC line rate.

Question for the WG — is single-PUT permitted for CLOSED?

S3DLIO_MULTIPART_THRESHOLD_MB is a storage-library environment variable, not a
DLIO workload parameter. The mlpstorage rules engine
(CLOSED_ALLOWED_PARAMS, mlpstorage_py/rules/run_checkers/training.py) only governs
DLIO --params overrides, so it neither tracks nor blocks this setting today.

  • Training: the dataset is produced in a separate datagen step; the upload
    method there does not affect the measured (read) benchmark, so single-PUT datagen
    appears unobjectionable.
  • Checkpointing: the write path is the measured metric, and multipart vs
    single-PUT materially changes write behavior/throughput. Is forcing single-PUT
    (non-multipart) writes permissible in a CLOSED checkpointing submission?

Related

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

No labels
No labels

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions