nvidia-modeset: ignore nested nvRevokeDevice() calls - #1395
Open
SammyTourani wants to merge 1 commit into
Open
SammyTourani wants to merge 1 commit into
SammyTourani wants to merge 1 commit into
Conversation
If the core channel cannot be reallocated on resume, nvResumeDevEvo() calls nvRevokeDevice(), which calls FreeDeviceReference() for every client. For the modeset owner, FreeDeviceReference() releases ownership, RestoreConsole() fails to reallocate the core channel again and calls nvRevokeDevice() a second time. The nested call runs FreeDeviceReference() again for the owner, so its device reference is dropped twice and the device is freed while other clients still point at it. Their teardown then reads the freed NVDevEvoRec, which oopses in nvEvoDisableVblankSemControl() with a NULL pDispEvo. Return early from a nested nvRevokeDevice(); the outer loop already frees every remaining reference.
|
|
tdortman
added a commit
to tdortman/dotfiles
that referenced
this pull request
Sep 29, 2026
A failed resume makes `nvidia-modeset` call `nvRevokeDevice` from inside itself. That frees the device twice and oopses in the PM notifier, which wedges the machine. Apply `NVIDIA/open-gpu-kernel-modules#1395` (fixes #1284) to the open module, pinned to its commit. Drop it once the PR is merged.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Fixes #1284
nvRevokeDevice()now returns early when it is called while a revoke is already running.A file-scope
static NvBool revokeInProgresstracks this, following thestatic int suspendCounterpattern already used for nested suspend/resume in the samefile. The flag is set before the loop over
perOpenIoctlListand cleared after it.The nested call can only come from
FreeDeviceReference()->ReleaseModesetOwnership()->
RestoreConsole()when the core channel cannot be reallocated. At that point the outerloop is already freeing every client's reference. The nested pass adds only one thing: a
second
FreeDeviceReference()for the pOpenDev that is currently being freed. With theearly return, each client reference is dropped exactly once and the device is freed after
the last one, as the outer revoke intends. Revokes that are not nested behave as before.
Examples are a resume failure with no modeset owner, and
RestoreConsole()failing afterNVKMS_IOCTL_RELEASE_OWNERSHIP.One file, +19 lines (
src/nvidia-modeset/src/nvkms.c).Scope:
unload failed"). The GPU still needs recovery afterwards, but suspend now fails without
a kernel oops and a use-after-free.
FreeDeviceReference()can still be reached with no outer revoke:nvKmsClose()/NVKMS_IOCTL_FREE_DEVICEof the modeset owner while the core channelcannot be reallocated (for example, Xorg exiting after the GPU died while other clients
still hold the device). I left that path alone. Fixing it needs
FreeDeviceReference()to keep its own device reference until after
ReleaseModesetOwnership(), which changesthe ordering of the normal close path.
Verification
The repository has no test suite, and the bug needs a failing GSP. So I verified with (1) a
userspace reproduction that runs the unmodified nvidia-modeset code, and (2) a compile of
the changed file with the Makefile's exact flags for x86_64 Linux.
1. Userspace reproduction. The harness compiles the real
nvkms.c,nvkms-evo.c,nvkms-rm.c,nvkms-utils.candnvkms-vblank-sem-control.cwith clang andAddressSanitizer. It provides:
nvkms_*OS interface on libc;NV_ERR_TIMEOUTfor every call;nvAllocDevEvo()/nvAllocCoreChannelEvo()leave it in;return 0stubs for the ~136 functions outside those files that these paths never depend on.Clients are set up through the real
NVKMS_IOCTL_ALLOC_DEVICE,GRAB_OWNERSHIPandENABLE_VBLANK_SEM_CONTROLioctls. Then the realnvKmsSuspend()runs, the GPU "dies",and the real
nvKmsResume()runs, which is whatnv_suspend_devices()does when RMsuspend fails. A second build ("stale") keeps freed memory mapped with its old contents,
like a vfree()d buffer read through a stale TLB entry, so the run goes all the way to the
field crash instead of stopping at the first ASan report.
Command:
for m in asan stale; do ./build.sh $BASE out-base-$m $m; ./build.sh $FIX out-fix-$m $m; done; ./run-matrix.sh(run in the harness directory;
$BASE= pristine 615.71.09 tree,$FIX= this branch)Result: 28 runs (7 scenarios x 2 trees x 2 builds). The repository has no test suite of
its own, so these are the only runs.
Scenarios:
field: kernel client owns modeset (nvidia-drm, fbdev=1), then a user client with a vblank semaphore control.twousers: the owner plus two such user clients.xorg: nvidia-drm without ownership, a user-space owner, and a user client.userfirst: the user client opened before the owner.noowner: no modeset owner.release:RELEASE_OWNERSHIPwhile the GPU is dead (a revoke that is not nested).closeowner: the owner closes its fd while the GPU is dead (the path this patch leavesalone, see above).
Unpatched
field, symbolized ASan report (the freed region is the 10592-byte NVDevEvoRec):Unpatched
field, stale build:SIGSEGV, fault address 0x2740, pc -> nvEvoDisableVblankSemControl (nvkms-vblank-sem-control.c:236).0x2740 =
offsetof(NVDispEvoRec, vblankApiHeadState[0].vblankCount)in 615.71.09, whichis the field crash.
2. Compile check with the project's flags.
Command: I took the
src/nvkms.ccompile line printed bymake -n -C src/nvidia-modeset TARGET_ARCH=x86_64 CC=clang(
-mcmodel=kernel -mno-red-zone -msoft-float -ffreestanding -Wall -Wextra ... -std=gnu11 -c src/nvkms.c)and ran it unchanged except for
clang --target=x86_64-linux-gnu. I ran it insrc/nvidia-modesetof the pristine tree and of this branch.Result:
Not done: I did not run the patched driver on real hardware or build it with kbuild
against a kernel tree (no NVIDIA GPU or Linux kernel headers here).