Summary
The leader-only storage sweep (run_storage_sweep_tick → maybe_spawn_sweep, crates/buzz-relay/src/storage_sweep.rs) fails continuously on every bb-public pod — every 300s usage tick, no backoff, ~90–104s per attempt — because the bb-public media bucket organically outgrew the sweep's 1,000,000-object cap (BUZZ_STORAGE_SWEEP_MAX_OBJECTS, default 1_000_000, storage_sweep.rs:65-68) (projected crossing ≈ 2026-08-01T16:43Z; the success-only gauge froze before observing it). The cap is doing its documented job; the defects are the consequences: no headroom warning, a misleading static error hint, unbounded no-backoff retry of a permanent condition (1001 paged LIST calls per attempt, ~30% of leader wall-clock, forever), and storage gauges that freeze silently at the last successful snapshot.
Found during the Blob investigation (#4556) — ruled out as the Blob carrier on timing evidence (see below). Both sweep source files are byte-identical between prod c104eecfb and current main, so the analysis covers both.
Diagnosis history: this issue has carried three successive framings — (1) IAM/s3:ListBucket grant, (2) rollout regression from a 16:33Z config change, (3) non-advancing S3 paginator. All three are dead (receipts below). The bucket really is >1M objects. Do not act on any earlier framing.
Evidence (Datadog + source; credit: Sami, verified by Eva/Wren against source)
The error, uniform across all 578 occurrences:
custom.error: 'object cap exceeded: 1001000 listed objects > cap 1000000'
seen=1,001,000 is the cap's saturation constant, not a measurement. fold_bucket_listing (crates/buzz-media/src/bucket_index.rs:389-395) increments seen by the actual page length before testing seen > cap, so a bucket of 1,000,001–1,000,999 objects reports its true count at the trip (partial final page), while every bucket ≥1,001,000 objects fills the 1001st page and reports the identical seen=1,001,000. The observed value proves at least 1,001,000 listed objects and nothing above that. (Source-verified independently by Eva and Wren; range refinement: Wren.)
The split is cluster/bucket, not pod/image/RS/IAM:
kube_cluster_name: bb-block-infra-cell-01 12,767 -> 17,909 objects (4d) — all pods sweep OK
kube_cluster_name: bb-public-infra-cell-01 747,557 -> 996,752 objects (4d) — ALL pods FAIL
578/578 failure log lines are bb-public; 0 elsewhere. Earlier "same image both fails and succeeds" was a cross-cluster comparison — the same image succeeds where the bucket is 17k and fails where it's ~1M.
Organic crossing, not a deploy: bb-public grew 975,954 → 996,752 between 08-01T11:12Z and 15:58Z (observed ~+4,363 objects/hr); projected 1M crossing ≈ 08-01T16:43Z — the gauge only emits on success, so the crossing itself was never observed, only the last success at 996,752 and the first failure 08-01T19:48:18Z (~3.8h later), consistent with organic crossing in that gap. Failures then continuous at 12/hr (one per 300s tick), 578 occurrences. Once the sweep started failing, buzz_total_storage_objects{bb-public} froze at 996,752 — the gauge only emits from a cached successful snapshot (storage_sweep.rs:308), so bb-public storage metrics have been silently stale since 08-01, reading as a plausible small-ish number rather than as absent.
Per-attempt cost: a failing attempt pages 1001 × ~92ms ≈ 90–104s of a 120s timeout, retried unconditionally on the next 300s tick (should_spawn, storage_sweep.rs:149-170, no backoff). The docstring's premise — "a permission failure (the common persistent case) costs a single cheap LIST call per retry" — does not hold for a legitimately oversized bucket: this permanent condition costs 1001 LIST calls every 300s, forever, ~30% of the leader's wall-clock.
Dead theories — do not act on these
- IAM /
s3:ListBucket grant — DEAD. Error is CapExceeded on all 578 occurrences; nothing to grant. Do not run a kubectl SA/IRSA diff.
- Rollout regression (16:33Z config change) — DEAD. Failures predate it by ~45h and align with organic bucket growth instead.
- Non-advancing paginator /
NonAdvancingPage detector — DEAD. The 1,001,000 figure is the cap's own arithmetic for any >1M bucket; the "true bucket = 11.6k" comparator was bb-block's bucket (cross-cluster {*} query artifact). Wren source-reviewed the token path (storage.rs:245-268 → rust-s3 ListObjectsV2) and found no layer that could manufacture a stale token; his local paginator-guard patch was discarded before pushing. A repeated-token guard may be reasonable generic hardening someday but is not evidenced by this incident and must not be written as its fix.
- SDK retry chain / Datadog-UI one-click ask — DEAD/WITHDRAWN (per-page latency is normal; the
error field is reachable via the log-search API's extra_fields).
The misleading log hint (sub-defect, keep)
The harvest arm logs a static storage sweep failed; verify s3:ListBucket (or MinIO list) permission is granted on the bucket for every SweepError variant (CapExceeded, Storage(MediaError), Timeout, MalformedPage), with the real error only in the structured error = %err field. In this incident that hint cost two dead theories, one wrong diagnosis, and an unnecessary human ask.
Why this is NOT the Blob (#4556)
Bounded out on timing (this argument does not depend on bucket size and survives all corrections): on tms7j, the first sweep attempt completed 3–4+ minutes before carrier-ramp onset, and a 44-minute window containing four failing attempts showed +7.4 MB/min vs the 300–900 MB/min carrier ramp. Note the transient workload per failing attempt is larger than earlier stated (1001 pages, ~1M objects folded, not ~12), so stop describing it as small — timing, not magnitude, is the exculpatory evidence. Kept on the Blob board as a controlled confound.
Suggested fixes (ordered; canonical ordering: Wren)
⚠️ A cap-only raise is invalid. The cap is coupled to an independent whole-sweep timeout: BUZZ_STORAGE_SWEEP_TIMEOUT_SECS (default 120s) wraps the entire fold via tokio::time::timeout (storage_sweep.rs:250-252), including parsing and aggregation. At the observed ~92.9ms/page, the network-only ceiling is ~1.29M objects — practical capacity is lower. Raising the cap by 1M does not buy runway; it converts CapExceeded into Timeout around ~1.29M.
- Immediate mitigation: disable the sweep in bb-public (lowest risk) while a sustainable configuration is designed. Mitigation vehicle: squareup/bb-public#347 (the Blob canary control PR) sets
BUZZ_STORAGE_METRICS: "off" in bb-public alongside the pinned canary control — merging it delivers this fix. Note the value is correctly quoted: unquoted off is YAML boolean false, which would render as the string "false" and leave the sweep enabled (from_env disables only on exactly off, storage_sweep.rs:69-73). No numeric cap/timeout pair is endorsed by this issue — if operations require restoring the sweep instead, choose cap + timeout + interval together from measured current cardinality (~1.0M objects, +4,363/hr) with explicit runway, noting the constraints below. Either path is a values.yaml change; restoring unfreezes the gauges.
- Stop retry amplification: classify permanent/configuration failures (
CapExceeded) separately and back them off. The source comment confuses bounded per-attempt work with bounded repeated cost — 1001 LIST calls every 300s forever is not operationally bounded.
- Restore truthful health: expose cap utilization before failure and snapshot age/staleness after failure; frozen success gauges must not read as current.
- Correct diagnosis output: variant-specific hints —
CapExceeded must not recommend IAM.
- Redesign the scan — scheduled now, not "longer-term": a full-bucket LIST/fold has a computable structural ceiling at bb-public growth. Options: prefix-sharded or incremental listing, or a much longer sweep interval decoupled from the usage tick.
Sizing constraints (why no numbers are blessed)
- Measured all-inclusive page cost: ~92.9 ms/page (from
started.elapsed() around the timeout-wrapped sweep, storage_sweep.rs:248-256 — includes HTTP, XML parse, classify, fold). Observed pod range 90.2–98.1 ms/page; will drift with key mix and S3 latency.
- Timeout wall: at that rate the default 120s timeout supports ~1.29M objects; cap-only headroom from the current ~997k is ~295k objects ≈ 2.8 days, not the ~9.5 days/1M naive cap arithmetic suggests.
- Scaling both limits is a ~weeks runway, not years: e.g. cap 4M needs timeout ≥ ~372s and buys only ~28.7 days at +4,363 objects/hr from today's ~1.0M.
- One-tick structural ceiling: a scan of ~3.23M objects takes ≥300s — the failure-retry cadence — arriving in ~21 days at current growth. A 4M cap already exceeds it. (Cadence note: successful sweeps run at
BUZZ_STORAGE_SWEEP_INTERVAL_SECS, default 3600s; the 300s usage tick is the scheduler/harvest cadence and becomes the retry cadence only after failure — so a >300s successful scan can still serve an hourly metric, but harvest is delayed and every failure retry costs a full scan.)
Relationship to other issues
Evidence locker: buzz-the-blob war-room channel, thread 62082b57.
Summary
The leader-only storage sweep (
run_storage_sweep_tick→maybe_spawn_sweep,crates/buzz-relay/src/storage_sweep.rs) fails continuously on every bb-public pod — every 300s usage tick, no backoff, ~90–104s per attempt — because the bb-public media bucket organically outgrew the sweep's 1,000,000-object cap (BUZZ_STORAGE_SWEEP_MAX_OBJECTS, default1_000_000,storage_sweep.rs:65-68) (projected crossing ≈ 2026-08-01T16:43Z; the success-only gauge froze before observing it). The cap is doing its documented job; the defects are the consequences: no headroom warning, a misleading static error hint, unbounded no-backoff retry of a permanent condition (1001 paged LIST calls per attempt, ~30% of leader wall-clock, forever), and storage gauges that freeze silently at the last successful snapshot.Found during the Blob investigation (#4556) — ruled out as the Blob carrier on timing evidence (see below). Both sweep source files are byte-identical between prod
c104eecfband currentmain, so the analysis covers both.Evidence (Datadog + source; credit: Sami, verified by Eva/Wren against source)
The error, uniform across all 578 occurrences:
seen=1,001,000is the cap's saturation constant, not a measurement.fold_bucket_listing(crates/buzz-media/src/bucket_index.rs:389-395) incrementsseenby the actual page length before testingseen > cap, so a bucket of 1,000,001–1,000,999 objects reports its true count at the trip (partial final page), while every bucket ≥1,001,000 objects fills the 1001st page and reports the identicalseen=1,001,000. The observed value proves at least 1,001,000 listed objects and nothing above that. (Source-verified independently by Eva and Wren; range refinement: Wren.)The split is cluster/bucket, not pod/image/RS/IAM:
578/578 failure log lines are bb-public; 0 elsewhere. Earlier "same image both fails and succeeds" was a cross-cluster comparison — the same image succeeds where the bucket is 17k and fails where it's ~1M.
Organic crossing, not a deploy: bb-public grew 975,954 → 996,752 between 08-01T11:12Z and 15:58Z (observed ~+4,363 objects/hr); projected 1M crossing ≈ 08-01T16:43Z — the gauge only emits on success, so the crossing itself was never observed, only the last success at 996,752 and the first failure 08-01T19:48:18Z (~3.8h later), consistent with organic crossing in that gap. Failures then continuous at 12/hr (one per 300s tick), 578 occurrences. Once the sweep started failing,
buzz_total_storage_objects{bb-public}froze at 996,752 — the gauge only emits from a cached successful snapshot (storage_sweep.rs:308), so bb-public storage metrics have been silently stale since 08-01, reading as a plausible small-ish number rather than as absent.Per-attempt cost: a failing attempt pages 1001 × ~92ms ≈ 90–104s of a 120s timeout, retried unconditionally on the next 300s tick (
should_spawn,storage_sweep.rs:149-170, no backoff). The docstring's premise — "a permission failure (the common persistent case) costs a single cheap LIST call per retry" — does not hold for a legitimately oversized bucket: this permanent condition costs 1001 LIST calls every 300s, forever, ~30% of the leader's wall-clock.Dead theories — do not act on these
s3:ListBucketgrant — DEAD. Error isCapExceededon all 578 occurrences; nothing to grant. Do not run a kubectl SA/IRSA diff.NonAdvancingPagedetector — DEAD. The 1,001,000 figure is the cap's own arithmetic for any >1M bucket; the "true bucket = 11.6k" comparator was bb-block's bucket (cross-cluster{*}query artifact). Wren source-reviewed the token path (storage.rs:245-268→ rust-s3 ListObjectsV2) and found no layer that could manufacture a stale token; his local paginator-guard patch was discarded before pushing. A repeated-token guard may be reasonable generic hardening someday but is not evidenced by this incident and must not be written as its fix.errorfield is reachable via the log-search API'sextra_fields).The misleading log hint (sub-defect, keep)
The harvest arm logs a static
storage sweep failed; verify s3:ListBucket (or MinIO list) permission is granted on the bucketfor everySweepErrorvariant (CapExceeded,Storage(MediaError),Timeout,MalformedPage), with the real error only in the structurederror = %errfield. In this incident that hint cost two dead theories, one wrong diagnosis, and an unnecessary human ask.Why this is NOT the Blob (#4556)
Bounded out on timing (this argument does not depend on bucket size and survives all corrections): on
tms7j, the first sweep attempt completed 3–4+ minutes before carrier-ramp onset, and a 44-minute window containing four failing attempts showed +7.4 MB/min vs the 300–900 MB/min carrier ramp. Note the transient workload per failing attempt is larger than earlier stated (1001 pages, ~1M objects folded, not ~12), so stop describing it as small — timing, not magnitude, is the exculpatory evidence. Kept on the Blob board as a controlled confound.Suggested fixes (ordered; canonical ordering: Wren)
BUZZ_STORAGE_METRICS: "off"in bb-public alongside the pinned canary control — merging it delivers this fix. Note the value is correctly quoted: unquotedoffis YAML booleanfalse, which would render as the string"false"and leave the sweep enabled (from_envdisables only on exactlyoff,storage_sweep.rs:69-73). No numeric cap/timeout pair is endorsed by this issue — if operations require restoring the sweep instead, choose cap + timeout + interval together from measured current cardinality (~1.0M objects, +4,363/hr) with explicit runway, noting the constraints below. Either path is a values.yaml change; restoring unfreezes the gauges.CapExceeded) separately and back them off. The source comment confuses bounded per-attempt work with bounded repeated cost — 1001 LIST calls every 300s forever is not operationally bounded.CapExceededmust not recommend IAM.Sizing constraints (why no numbers are blessed)
started.elapsed()around the timeout-wrapped sweep,storage_sweep.rs:248-256— includes HTTP, XML parse, classify, fold). Observed pod range 90.2–98.1 ms/page; will drift with key mix and S3 latency.BUZZ_STORAGE_SWEEP_INTERVAL_SECS, default 3600s; the 300s usage tick is the scheduler/harvest cadence and becomes the retry cadence only after failure — so a >300s successful scan can still serve an hourly metric, but harvest is delayed and every failure retry costs a full scan.)Relationship to other issues
Evidence locker:
buzz-the-blobwar-room channel, thread62082b57.