Skip to content

rmapi: keep kernel clients' user BAR1 mappings in one range - #1403

Open
SammyTourani wants to merge 1 commit into
NVIDIA:mainfrom
SammyTourani:fix/issue-1134
Open

SammyTourani wants to merge 1 commit into
NVIDIA:mainfrom
SammyTourani:fix/issue-1134

Conversation

@SammyTourani

Copy link
Copy Markdown

Fixes #1134

One condition in memMap_IMPL() — only set ALLOW_DISCONTIG when the mapping is
neither a kernel client's mapping nor a MEM_SPACE_USER mapping:

if (!pMapParams->bKernel &&
    (DRF_VAL(OS33, _FLAGS, _MEM_SPACE, pMapParams->flags) != NVOS33_FLAGS_MEM_SPACE_USER))
{
    busMapFbFlags |= BUS_MAP_FB_FLAGS_ALLOW_DISCONTIG;
}

Real userspace clients still get discontiguous mappings (they always have
MEM_SPACE == CLIENT and nv-mmap maps every range). nvkms's user mapping no longer
does, so on a fragmented/full BAR1 it fails cleanly with NV_ERR_NO_MEMORY
instead of composing a mapping nvidia-drm mis-treats as contiguous. nvidia-drm
already handles that failure ("Failed to map NvKmsKapiMemory", -ENOMEM) — that is
the survivable path several reporters saw before the fatal escalation.

The single-range attempt is tried first regardless of the flag, and BAR1 mapping
reuse does not depend on it, so the only behaviour removed is the multi-range
fallback for the nvkms path. Scope: 1 file, +11/-1 (the change is one if; the
rest is a comment).

Verification

No GPU or kernel build on this machine, so I verified in a userspace ASan harness
that compiles the real RM sources unchanged (mapping_cpu.c,
kern_bus_gm107.c, kern_bus_tu102.c, mapping_reuse.c, containers/map.c,
os.c) and drives the real serverMap_Prologue -> memMap_IMPL ->
kbusMapFbAperture -> reusemappingdbMap -> osMapPciMemoryAreaUser path, then
replays the tail of rm_kernel_rmapi_op(NV04_MAP_MEMORY) verbatim and checks,
page by page, what nvidia-drm (contiguous, first-range-only) vs nv-mmap
(every-range) would actually reach through the returned address. Scenario A uses
the reporter's exact BAR1 layout; F–I cover static BAR1 (Turing+ large-BAR path,
real kbusGetStaticFbAperture_TU102).

Command:

bash /Volumes/SammyDisk/Downloads/oss-1134/harness/verify.sh \
     /Volumes/SammyDisk/oss-contrib-work/nvidia__open-gpu-kernel-modules@1134 \
     61dcc937 eababfaa

Result (real output):

################ base: 61dcc937 615.71.09
=== A: nvkms USER mapping, BAR1 fragmented (layout matching the #1134 report)
  map #1: memMap_IMPL -> 0x00000000, address 0xfccfdd0000
    RM BAR1 mapping behind the returned address: 3 range(s)
    nvidia-drm: ioremap_wc(0xfccfdd0000, 0x330000)
  resource: resource sanity check: requesting [mem 0x000000fccfdd0000-0x000000fcd00fffff], which spans more than 0000:01:00.0 [mem 0xfcc0000000-0xfccfffffff 64bit pref]
  caller __nv_drm_gem_nvkms_map+0x99/0xf0 [nvidia_drm] mapping multiple BARs
        408 / 816 pages -> correct object page
        152 / 816 pages -> unmapped BAR1 VA
        256 / 816 pages -> outside BAR1 (next BAR)
  RESULT: FAIL
...
6/9 scenarios passed
exit status: 1

################ fixed: eababfaa rmapi: keep kernel clients' user BAR1 mappings in one range
=== A: ...
  map #1: memMap_IMPL -> 0x00000051, address 0x0
    nvkms MapMemory() fails -> nvidia-drm: Failed to map NvKmsKapiMemory, -ENOMEM
  RESULT: PASS
...
9/9 scenarios passed
exit status: 0

Base: 6/9 (A, B nvkms-fragmented and G static-discontiguous compose cross-range
mappings). Fixed: 9/9 — those three fail cleanly with -ENOMEM; the contiguous,
reuse, and real-userspace scenarios (C, D, E, F, H, I) are unchanged and still
map every page correctly. The harness reproduces the reporter's exact
resource sanity check range on base.

Also compiled the changed file for real x86_64 Linux with the exact flags from
make -n -C src/nvidia TARGET_OS=Linux TARGET_ARCH=x86_64 (which includes
-Werror-implicit-function-declaration), base and fixed: both 0 warnings / 0
errors, identical compiler output, valid ELF x86-64 object.

Since 565.57.01, memMap_IMPL() sets BUS_MAP_FB_FLAGS_ALLOW_DISCONTIG
for every mapping with bKernel == NV_FALSE. That includes a kernel
client asking for an NVOS33_FLAGS_MEM_SPACE_USER mapping, which is what
nvkms MapMemory(NVKMS_KAPI_MAPPING_TYPE_USER) does for nvidia-drm.

When BAR1 is too fragmented for one range, kbusMapFbAperture_GM107()
splits the mapping. osMapPciMemoryAreaUser() returns only the start of
the first range, and rm_kernel_rmapi_op() frees the range list.
nvidia-drm then ioremaps and inserts PFNs for the whole object from that
address. The pages past the first range belong to other BAR1 mappings,
unmapped BAR1 space or the next PCI BAR.

Only allow a discontiguous mapping when it is not a kernel client's
MEM_SPACE_USER mapping. Such a mapping then fails with NV_ERR_NO_MEMORY
when BAR1 has no single range free. User clients still get
discontiguous mappings, because nv-mmap maps every range.

With static BAR1, the same flag let kbusGetStaticFbAperture_TU102()
return one range per physical chunk, with the same result in nvidia-drm.
For physically discontiguous memory such a mapping now fails with
NV_ERR_INVALID_ARGUMENT, as a kernel mapping of it already does.
@CLAassistant

Copy link
Copy Markdown

CLA assistant check
Thank you for your submission! We really appreciate it. Like many open source projects, we ask that you sign our Contributor License Agreement before we can accept your contribution.
You have signed the CLA already but the status is still pending? Let us recheck it.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

2 participants