Skip to content

adaptive: forced switches leave keep-alive connections behind — async conns miss the 1.2 s flap window in every run, and a one-worker io_uring strands sync conns until their requests fail #657

Description

@FumingPower3925

Measured on celeris 468ce53 for celeris#641: go test -race -count=1 -timeout=300s -json ./adaptive/... in a Docker Desktop linux/arm64 container (golang:1.27 = go1.27.1, --cpus 4, seccomp=unconfined). Three runs at memlock 8 MiB, the GitHub-hosted default, where io_uring is capped to one worker. Three more at 128 MiB, where it gets four. Every number below comes from the tests' own log lines.

TestBidirectionalFlapAsync fails in all 6 runs, at both memlock limits

The test runs 64 keep-alive clients with AsyncHandlers: true. Each flap calls ForceSwitch, sleeps a flat 1200 ms, then asserts at least 32 conns on the new active engine and at most 16 left on the standby.

run flap 1 → io_uring, active / standby flap 3 → io_uring, active / standby request errors async promotions, epoll / io_uring
8 MiB r1 18 / 46 19 / 45 0 82 / 37
8 MiB r2 18 / 46 18 / 46 0 82 / 36
8 MiB r3 31 / 33 21 / 43 0 95 / 52
128 MiB r1 14 / 50 22 / 42 0 78 / 36
128 MiB r2 22 / 42 21 / 43 0 86 / 43
128 MiB r3 18 / 46 22 / 42 0 82 / 40

Flap 2 (io_uring → epoll) reaches 64 / 0 in every run, so the loss is one direction only. epoll → io_uring moves about a third of the async connections within 1.2 s, whatever io_uring's worker count, and no request fails.

By design, a promoted async conn moves only when its dispatch goroutine is parked, idle and flushed (engine/epoll/transplant.go:82-133). A client that sends its next request at once keeps that goroutine busy. Two readings fit the data, and they need different fixes:

To tell them apart, poll the per-engine counts after ForceSwitch (as awaitPhase does) and record how long it takes to reach 60 on the active engine. Record both engines' transplant accounting from EngineMetrics at the end.

TestBidirectionalFlap (sync) and TestReverseTransplant fail only with a one-worker io_uring

run flap 3 → io_uring, active / standby missing at flap 3 (64 − active − standby) errors at flap 3 errors by test end
8 MiB r1 31 / 12 21 21 58
8 MiB r2 27 / 34 3 3 49
8 MiB r3 34 / 27 3 3 49
128 MiB r1–r3 passes 0 0 0

TestReverseTransplant at 8 MiB converged at 64/64 with 0 errors in runs 1 and 2. In run 3 it converged at 47/64 with 17 errors. At 128 MiB it reached 64/64 with 0 errors in all three runs.

The load clients (driveKeepAlive) stop at their first write or read error, with a 2 s read deadline, and never redial. Each error is one connection gone for good.

  • At flap 3, the number of missing connections equals the error count exactly, in all three 8 MiB runs.
  • Between that measurement and the end of the test, another 37–46 clients fail. Those connections were still counted as active on one engine. Their requests were not answered within 2 s.

Every run in both arms hit oscillation detected, locking switches right before flap 3. The 128 MiB runs still pass, so the lock alone does not explain the loss.

Not established: what the errors are. The test counts them and discards the values, so they could be resets, EOFs or deadline expiries. Also unknown: why a one-worker io_uring strands transplanted connections.

First steps:

  1. Record the error strings; the ramp rig's errSamples already does this.
  2. Record both engines' transplant counters.
  3. Reproduce at 8 MiB.

The CI shape is exactly where this shows, and it is where ./adaptive/... would run under #641.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions