fix(envd): fall back to lazy unmount when forced NFS umount fails - #2683
Conversation
The current setupNFS path on resume runs `umount --force` against the inherited NFS mount. For NFS, MNT_FORCE only aborts in-flight RPCs (fs/nfs/super.c::nfs_umount_begin); it does not detach the mount when the customer's resumed processes still hold open FDs into it. With the NFS proxy reachable on the same node there are no in-flight RPCs to abort, so the call fails with EBUSY and the orchestrator surfaces "failed to setup NFS volumes: failed to unmount stale NFS mount" as HTTP 400 to the customer. Now: try `umount --force` first (unchanged fast path), and on any failure fall back to `umount --lazy` (MNT_DETACH) before bubbling the error. Lazy detach unlinks the mount from the namespace immediately, the path becomes available for the fresh mount, and the kernel cleans up when the surviving FDs close. Old FDs keep reading from the previous mount — correct when the NFS source IP is unchanged across resume, and produces the same ESTALE outcome the customer would have seen from a successful forced unmount when the source changed. When the fallback triggers we log a warning so the rate is observable.
PR SummaryMedium Risk Overview Reviewed by Cursor Bugbot for commit ff2c941. Bugbot is set up for automated code reviews on this repo. Configure here. |
❌ 7 Tests Failed:
View the full list of 13 ❄️ flaky test(s)
To view more test analytics, go to the Test Analytics Dashboard |
There was a problem hiding this comment.
Code Review
The fallback to umount --lazy will not execute if the initial umount --force fails due to a context timeout because both commands share the same context. To ensure the mount is detached when the first command hangs, the fallback should be attempted with a fresh context or a separate timeout.
|
Superseded by #2700 (move envd into a dedicated network namespace), which addresses the root cause. |
There was a problem hiding this comment.
LGTM, small defensive fallback for NFS unmount with the gemini-suggested fresh context fix already applied.
Extended reasoning...
Overview
Small two-file change in packages/envd: unmountNFS now falls back to umount --lazy when umount --force fails (customer FDs holding the mount), and the envd version is bumped to 0.5.24. Total logic change is ~7 lines.
Security risks
None. This is an internal-to-envd unmount fallback running on resume. No new inputs, no privilege changes, no exposure surface.
Level of scrutiny
Low. The change is mechanical, self-contained, and only adds a fallback on an existing error path. The prior failure mode was a hard error; the new behavior tries one more thing before returning the same error. The fresh-context concern raised by gemini-code-assist was already addressed (commit 36f8f9d) using context.WithoutCancel with a 2s timeout, so the lazy attempt isn't starved by the parent budget.
Other factors
Bug hunting system found no bugs. Codecov failures are all pre-existing flakes (>50% flake rate in main). No outstanding reviewer comments. Author noted this is bifrost-track and will be superseded by #2700, but the fix is still useful in the interim.
🤖 I have created a release *beep* *boop* --- ## 0.0.1 (2026-07-29) ### Features * **envd:** add --no-cgroups flag to disable cgroup management ([#2811](#2811)) ([e10814c](e10814c)) * **envd:** add optional EntryInfo to watch FilesystemEvent ([#2930](#2930)) ([bbbc7c8](bbbc7c8)) * **envd:** allow opting into watching network mounts ([#2982](#2982)) ([9799dd0](9799dd0)) * **envd:** give envd realtime IO priority, reset for user processes ([#2681](#2681)) ([f4bd1b2](f4bd1b2)) * **envd:** split collapse stats into real migrations vs already-huge ([#3021](#3021)) ([0d77614](0d77614)) * **envd:** support user-defined file metadata via xattrs ([#2732](#2732)) ([da8fbe4](da8fbe4)) * freeze user cgroup across pause/resume to keep envd /init responsive ([#2688](#2688)) ([eceb741](eceb741)) * **orch:** collapse envd's heap into 2 MiB hugepages before pause to cut cold-resume faults ([#2997](#2997)) ([6677f73](6677f73)) * **orch:** distro-aware template base-image provisioning ([#3411](#3411)) ([f8c7b5b](f8c7b5b)) ### Bug Fixes * added envd to artifact repository ([#3432](#3432)) ([6c4f0e2](6c4f0e2)) * correct 3 CVES ([#3218](#3218)) ([076823b](076823b)) * **envd:** avoid Start deadlock after request cancellation ([#3256](#3256)) ([04317f8](04317f8)) * **envd:** bound the in-memory logs queue ([#2676](#2676)) ([05c9939](05c9939)) * **envd:** discard output when no subscriber is connected ([#2639](#2639)) ([8cf1795](8cf1795)) * **envd:** fall back to lazy unmount when forced NFS umount fails ([#2683](#2683)) ([5346a0d](5346a0d)) * **envd:** ignore closed pty read errors ([#2769](#2769)) ([6118672](6118672)) * **envd:** include suppressed count in exporter error logs ([#2680](#2680)) ([35c1141](35c1141)) * **envd:** make /init lock ctx-aware to prevent retry pile-up ([#2702](#2702)) ([173afd4](173afd4)) * **envd:** make CA install lock ctx-aware ([#2690](#2690)) ([83ee89f](83ee89f)) * **envd:** replace env vars in /init instead of merging ([#2706](#2706)) ([1b52e9a](1b52e9a)) * **envd:** replace time.Sleep with ticker in ScanAndBroadcast for prompt shutdown ([#3374](#3374)) ([002fd9f](002fd9f)) * **envd:** self-heal MMDS routing on /init lookup failure ([#2701](#2701)) ([90944d5](90944d5)) * **envd:** stop freezing socat cgroup across pause/resume ([#2923](#2923)) ([8b6f2b9](8b6f2b9)) * **envd:** stop misleading CA install cancel errors on rapid /init ([#3206](#3206)) ([91d09e4](91d09e4)) * **envd:** suppress repeat MMDS poll failures ([#2678](#2678)) ([73d691a](73d691a)) * **envd:** tolerate busy tmpfs cleanup in tests ([#2938](#2938)) ([a485834](a485834)) * **envd:** use constant-time comparison for signature validation ([#3145](#3145)) ([fcf92fa](fcf92fa)) * **envd:** use WithoutCancel for CA cleanup goroutine ctx ([#3207](#3207)) ([ee7bf84](ee7bf84)) ### Performance Improvements * **envd:** stop logging streamed payload content ([#2755](#2755)) ([db3868c](db3868c)) * **sandbox:** keep envd logging out of journald ([#2675](#2675)) ([f6943ca](f6943ca)) --- This PR was generated with [Release Please](https://github.com/googleapis/release-please). See [documentation](https://github.com/googleapis/release-please#release-please). Co-authored-by: e2b-release-please[bot] <298072688+e2b-release-please[bot]@users.noreply.github.com>
🤖 I have created a release *beep* *boop* --- ## 0.0.1 (2026-07-29) ### Features * **envd:** add --no-cgroups flag to disable cgroup management ([#2811](#2811)) ([e10814c](e10814c)) * **envd:** add optional EntryInfo to watch FilesystemEvent ([#2930](#2930)) ([bbbc7c8](bbbc7c8)) * **envd:** allow opting into watching network mounts ([#2982](#2982)) ([9799dd0](9799dd0)) * **envd:** give envd realtime IO priority, reset for user processes ([#2681](#2681)) ([f4bd1b2](f4bd1b2)) * **envd:** split collapse stats into real migrations vs already-huge ([#3021](#3021)) ([0d77614](0d77614)) * **envd:** support user-defined file metadata via xattrs ([#2732](#2732)) ([da8fbe4](da8fbe4)) * freeze user cgroup across pause/resume to keep envd /init responsive ([#2688](#2688)) ([eceb741](eceb741)) * **orch:** collapse envd's heap into 2 MiB hugepages before pause to cut cold-resume faults ([#2997](#2997)) ([6677f73](6677f73)) * **orch:** distro-aware template base-image provisioning ([#3411](#3411)) ([1abece1](1abece1)) ### Bug Fixes * added envd to artifact repository ([#3432](#3432)) ([b7024ba](b7024ba)) * correct 3 CVES ([#3218](#3218)) ([076823b](076823b)) * **envd:** avoid Start deadlock after request cancellation ([#3256](#3256)) ([04317f8](04317f8)) * **envd:** bound the in-memory logs queue ([#2676](#2676)) ([05c9939](05c9939)) * **envd:** discard output when no subscriber is connected ([#2639](#2639)) ([8cf1795](8cf1795)) * **envd:** fall back to lazy unmount when forced NFS umount fails ([#2683](#2683)) ([5346a0d](5346a0d)) * **envd:** ignore closed pty read errors ([#2769](#2769)) ([6118672](6118672)) * **envd:** include suppressed count in exporter error logs ([#2680](#2680)) ([35c1141](35c1141)) * **envd:** make /init lock ctx-aware to prevent retry pile-up ([#2702](#2702)) ([173afd4](173afd4)) * **envd:** make CA install lock ctx-aware ([#2690](#2690)) ([83ee89f](83ee89f)) * **envd:** replace env vars in /init instead of merging ([#2706](#2706)) ([1b52e9a](1b52e9a)) * **envd:** replace time.Sleep with ticker in ScanAndBroadcast for prompt shutdown ([#3374](#3374)) ([002fd9f](002fd9f)) * **envd:** self-heal MMDS routing on /init lookup failure ([#2701](#2701)) ([90944d5](90944d5)) * **envd:** stop freezing socat cgroup across pause/resume ([#2923](#2923)) ([8b6f2b9](8b6f2b9)) * **envd:** stop misleading CA install cancel errors on rapid /init ([#3206](#3206)) ([91d09e4](91d09e4)) * **envd:** suppress repeat MMDS poll failures ([#2678](#2678)) ([73d691a](73d691a)) * **envd:** tolerate busy tmpfs cleanup in tests ([#2938](#2938)) ([a485834](a485834)) * **envd:** use constant-time comparison for signature validation ([#3145](#3145)) ([fcf92fa](fcf92fa)) * **envd:** use WithoutCancel for CA cleanup goroutine ctx ([#3207](#3207)) ([ee7bf84](ee7bf84)) ### Performance Improvements * **envd:** stop logging streamed payload content ([#2755](#2755)) ([db3868c](db3868c)) * **sandbox:** keep envd logging out of journald ([#2675](#2675)) ([f6943ca](f6943ca)) --- This PR was generated with [Release Please](https://github.com/googleapis/release-please). See [documentation](https://github.com/googleapis/release-please#release-please). Co-authored-by: e2b-release-please[bot] <298072688+e2b-release-please[bot]@users.noreply.github.com>
Falls back to
umount --lazywhen forced unmount on resume fails with EBUSY (customer FDs hold the mount).