Skip to content

fix(media): explain clock skew behind an opaque upload 401 - #6929

Open
gabmichels wants to merge 1 commit into
block:mainfrom
gabmichels:fix/media-upload-clock-skew-preflight
Open

gabmichels wants to merge 1 commit into
block:mainfrom
gabmichels:fix/media-upload-clock-skew-preflight

Conversation

@gabmichels

Copy link
Copy Markdown

Closes #4075.

Media uploads fail with 401 authentication failed when the client's system clock has drifted. The 401 is deliberately opaque — all thirteen media auth failures collapse into it (buzz-media/src/error.rs) — so a user-fixable, non-Buzz problem is presented as an authentication failure, and people go looking for auth bugs. On #4075 the drift was ~6.75s on a Windows box whose w32time service had never started; uploads had worked days earlier and then stopped, with no client change in between.

This is option (2) from that issue: compare the local clock against the relay's Date header and say so, instead of leaving the user with a bare 401. No relay change — crates/buzz-media is untouched.

What it does

When an upload is refused with a 401, the client makes one short HEAD request to the upload route, reads the Date header, and appends an explanation if — and only if — the observation proves the clock is at fault:

{"error":"authentication failed"}

note: this machine's clock is at least 5.8s ahead of the relay's, which is outside the 5s that media upload auth allows. A clock this far out makes media uploads fail exactly like this, so it is worth ruling out before anything else. Sync the system clock and retry (Windows: start the w32time service, then w32tm /resync; macOS and Linux: turn on automatic network time).

The relay's own message is kept. Fifteen distinct auth failures collapse into that 401 and only one of them is the clock, so replacing the body would destroy the evidence every time the guess is wrong. The error also stays a CliError::Relay { status: 401 }, so exit codes and the JSON error category are unchanged.

What the message claims, and what it does not

It reports a measurement — this clock is provably outside the window — and offers it as the first thing to rule out. It does not claim to have diagnosed the 401 in hand, because a client cannot: the relay compares whole seconds (created > now + 5), both sides are truncated, and the transit between them is unknown, so a drift of ~5.6s is accepted or rejected depending on where the two clocks fall inside their respective seconds. Asserting causation would be wrong in exactly the band this feature exists to serve. Appending rather than replacing means a wrong guess costs the user nothing.

Why the diagnosis runs after the failure, not before it

The obvious design is a pre-flight check before minting the token. I built that first and it was wrong twice over.

Reading the Date off the rejection does not work. crates/buzz-relay/src/api/media.rs:139 verifies Blossom auth in FromRequestParts, before the body is read — the code comments call this out as the "pre-body auth-rejection guarantee". So on a 200 MB upload the 401's Date was stamped at t≈0 while the client spent the next minute pushing bytes. Treating it as "the relay's clock now" would pin every slow-upload 401 on a perfectly healthy clock. Hence a dedicated round trip: it is the only one short enough to bound.

A pre-flight that can refuse is a regression risk. Any check that runs before the upload and returns an error can turn a diagnostic into a new failure mode. Running only after a refusal makes that structurally impossible: there is no path where this code is the reason an upload does not happen.

The measurement is bracketed, not estimated

The local clock is read on both sides of the probe. Writing Δ for the true offset, D for the parsed header, and S for the server's real clock at stamp time, D ≤ S < D + 1s and before − Δ ≤ S ≤ after − Δ, which gives

before − D − 1s  ≤  Δ  ≤  after − D

Both bounds hold whatever the latency was, so no fudge factor is needed and none is applied. Each direction uses only its proven side: the "ahead" verdict reads the lower bound, sampled before the request, so latency cannot inflate it. A test pins this — 600 seconds of simulated latency on a synced clock yields no verdict. The cost of that rigour is silence when the interval straddles a bound, which is the right way to be wrong.

The local readings are in milliseconds; only the server's stamp is truncated. That matters for the reported case: with whole-second readings a 6.75s drift straddles the 5s tolerance and stays silent, while millisecond readings prove ≥5.75s.

A probe response carrying an Age header is discarded rather than measured — only a cache sends Age, and a cached reply's Date is as stale as the cache entry.

Both directions, and the backward bound is the token's own lifetime

Skew goes both ways, and the verifier's branches do not share a bound:

if created > now + 5              { return Err(MediaError::TimestampOutOfWindow); }
if now > created + max_age_secs   { return Err(MediaError::TimestampOutOfWindow); }
if exp_value <= now               { return Err(MediaError::TokenExpired); }

The check uses the expiration the client stamped on its own token. That is the tightest of the backward bounds — a clock behind by more than the token's whole lifetime mints one already expired on arrival — and, unlike max_age_secs, the client knows it exactly. It cannot know max_age_secs: the relay selects the image (600s) or video (3600s) pipeline by sniffing the body bytes, not from Content-Type (upload_blob: "Content-Type is advisory only; a bounded body prefix selects the streaming video path from actual bytes"), so any client-side guess keyed on MIME can disagree with what the relay actually applied.

Not done deliberately

The client could read the relay's clock and backdate created_at so skewed uploads succeed. That is a worse fix: a drifted clock also corrupts event timestamps, so masking it at the media layer hides a problem that is broken elsewhere too. The only correct repair is to fix the clock, and this change says so.

TimestampOutOfWindow is still collapsed into the blind 401. That is option (1) from the issue and belongs in its own PR.

Testing

  • buzz-core — 11 unit tests: Date parsing (IMF-fixdate, RFC 2822 at zero and non-zero offsets, junk), the bracketing arithmetic, out-of-order readings, the reported ~6.75s drift, the backward bound tracking the token lifetime, and three that exist to fail if the design regresses — latency alone never produces a verdict; a drift inside the tolerance stays silent at every fractional alignment of the server's stamp; the message never asserts causation.
  • buzz-cli — 6 integration tests against an axum stand-in relay that dates the probe and the rejection independently, so a test can prove which one the diagnosis trusts. Includes the regression test for the pre-body rejection (a rejection dated 600s stale with a clean probe must not be read as drift), an unreadable-Date case, and a non-401 failure that must not be probed at all.
  • desktop — the auth event's tag shape and that its expiration derives from the lifetime the diagnosis measures against, the server-tag omission path, and the probe guard for non-401 and cancelled uploads.
  • cargo clippy --workspace --all-targets and the Tauri clippy pass clean; cargo fmt --all --check clean; the desktop file-size ratchet passes (the new desktop code lives in its own module, and commands/media.rs is smaller than before).
  • Verified live against a running relay: a valid PNG still uploads 200, and a non-401 failure (422 invalid image data) still surfaces exactly as it did.

No UI screenshot: the desktop change alters the text of an existing upload-failure error string rather than any layout or interaction. Happy to add one if you would rather see it in situ.

Media uploads fail with a bare `401 authentication failed` when the
client's system clock has drifted. Every media auth failure collapses
into that one message to defeat enumeration, so a user-fixable,
non-Buzz problem is indistinguishable from a signing, scope, or
authorization bug.

After an upload is refused with a 401, the client now makes one short
`HEAD` of the upload route and reads the relay's `Date` header. The
local clock is sampled either side of that round trip, so the offset is
bracketed into the interval the observation proves rather than
estimated — latency can only widen the interval, never manufacture a
verdict. When the interval lies wholly outside what media upload auth
allows, the drift is reported alongside the relay's own message.

The probe is deliberately a separate request. The relay verifies
Blossom auth before it reads the body, so the `Date` on a large
upload's rejection was stamped at the start of a transfer that may
have run for minutes; reading the clock off it would blame a healthy
clock for someone else's auth failure.

The message reports a measurement, not a diagnosis. The relay compares
whole seconds across an unknown transit delay, so a drift near the
tolerance may be accepted or rejected depending on fractional
alignment — claiming causation would be wrong in exactly the band this
serves. The relay's message is appended to, never replaced, and the
error keeps its variant, status, exit code and JSON category.

Both directions are covered. The backward bound uses the `expiration`
the client stamped on its own token, which it knows exactly, unlike
`max_age_secs` — the relay selects the 600s or 3600s window by
sniffing the body bytes rather than trusting `Content-Type`.

Signed-off-by: Gab Michels <gab.michels@hotmail.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Bug] All media uploads fail with "TauriInvokeError: relay returned 401 Unauthorized: authentication failed"

1 participant