Problem
In packages/orchestrator/pkg/sandbox/nbd/pool.go:369, the ReleaseDevice retry loop calls time.Sleep(500 * time.Millisecond) which is not interruptible by context cancellation:
// pool.go:360-370 (simplified)
for {
select {
case <-ctx.Done(): // ← only checked at TOP of loop
return ctx.Err()
default:
}
attempt++
err := d.release(ctx, idx)
if err == nil { return nil }
// ...
time.Sleep(500 * time.Millisecond) // ← BUG: ctx cancellation not checked here
}
When the orchestrator shuts down and the shutdown context is cancelled, any goroutine blocked inside this sleep will not wake up for up to 500 ms per retry. With WithInfiniteRetry() enabled (used by both DirectPathMount.Close and DevicePool.Close), a single stuck device can stall the shutdown for an unbounded number of 500 ms sleeps before the context cancel is noticed.
Affected callers
path_direct.go:336: ReleaseDevice(ctx, idx, WithInfiniteRetry()) — called on every sandbox stop
pool.go:395: ReleaseDevice(ctx, slot, WithInfiniteRetry(), WithTimeout(devicePoolCloseReleaseTimeout)) — called on pool close
Fix
Replace time.Sleep with a context-aware select:
timer := time.NewTimer(500 * time.Millisecond)
select {
case <-ctx.Done():
timer.Stop()
return ctx.Err()
case <-timer.C:
}
This change exists in the local branch fix/nbd-pool-release-retry-ctx but has not been merged to main.
Impact
Delayed orchestrator shutdown when NBD devices fail to release. Under normal conditions (device releases cleanly on first attempt) there is no impact.
Problem
In
packages/orchestrator/pkg/sandbox/nbd/pool.go:369, theReleaseDeviceretry loop callstime.Sleep(500 * time.Millisecond)which is not interruptible by context cancellation:When the orchestrator shuts down and the shutdown context is cancelled, any goroutine blocked inside this sleep will not wake up for up to 500 ms per retry. With
WithInfiniteRetry()enabled (used by bothDirectPathMount.CloseandDevicePool.Close), a single stuck device can stall the shutdown for an unbounded number of 500 ms sleeps before the context cancel is noticed.Affected callers
path_direct.go:336:ReleaseDevice(ctx, idx, WithInfiniteRetry())— called on every sandbox stoppool.go:395:ReleaseDevice(ctx, slot, WithInfiniteRetry(), WithTimeout(devicePoolCloseReleaseTimeout))— called on pool closeFix
Replace
time.Sleepwith a context-aware select:This change exists in the local branch
fix/nbd-pool-release-retry-ctxbut has not been merged to main.Impact
Delayed orchestrator shutdown when NBD devices fail to release. Under normal conditions (device releases cleanly on first attempt) there is no impact.