Skip to content

io_uring: a worker whose ring setup or first submit fails leaves its SO_REUSEPORT listen socket open, and adaptive retries the build every tick (found by reading) #656

Description

@FumingPower3925

Found by reading 468ce53 during the celeris#639 fd audit. The leak is visible in the code: every step is cited below. Its rate under real load has not been measured.

The leak

Worker.run in engine/iouring/worker.go:

  1. Creates the worker's SO_REUSEPORT listen socket and stores it: createListenSocket (815), w.listenFD = listenFD (820).
  2. If ring setup fails, it sends ready <- fmt.Errorf("worker %d ring setup: …") (827) and returns without closing listenFD.
  3. If the first submit fails, it sends ready <- fmt.Errorf("worker %d initial submit: …") (913) and returns without closing listenFD or the ring.

The worker's run has already returned, so its shutdown never runs. Engine.Listen (engine.go:261-265) cancels the other workers, joins them, and returns the error. Nothing ever closes the failed worker's listen socket.

Why this is worse than a leaked descriptor

The socket is still listening, in the port's SO_REUSEPORT group. The kernel keeps hashing a share of new connections to it. Nobody accepts from it, so those clients complete the handshake into a backlog that is never drained, and they hang.

The adaptive engine is where that group is shared.

  • A lazy standby build (adaptive/engine.go:685 for io_uring) starts built.Listen on the same port as the active epoll engine, then waits 5 s for Addr().
  • When the io_uring Listen fails as above, the wait times out. performSwitch logs aborting switch: lazy standby build failed; staying on current active and returns (686-689).
  • Nothing records a switch attempt or starts a cooldown on that path, and e.secondary stays nil. The eval ticker is 1 s (controller.go:95). Unless ctrl.evaluate backs off by itself (not checked here), the next tick builds and fails again.
  • Each failed attempt leaves one more dead listener on the serving port. Each attempt also holds e.mu for the full 5 s bind wait.

The engine is left serving on epoll with a growing fraction of its incoming connections stranded. No error comes back to the caller; only a Warn is logged.

What would make a worker's ring setup fail after the tier probe passed

The tier probe (engine.go ~220-238) creates and closes a single test ring. Per-worker rings are created later, one per worker.

The plausible causes are:

  • memlock accounting across N rings;
  • rings leaked by earlier failed attempts (the submit-failure path above leaks one);
  • SQPOLL limits.

The CI shape (RLIMIT_MEMLOCK 8 MiB, see #533's history) is where ring creation is known to be tight. Not yet measured.

Fix

  • Close listenFD (and, on the submit path, the ring) on every early return after it is created. A defer that runs unless the worker reached ready <- nil covers all of them.
  • Add a test that injects a ring setup failure for one worker. It asserts that no listening socket for the port survives Listen's error: count listeners on the port via /proc/net/tcp, or try bind without SO_REUSEPORT. The test must fail on main.
  • Separately, decide whether an aborted lazy standby build should back off instead of retrying every tick.

No activity

Activity on this issue will appear here.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions