NVIDIA Open GPU Kernel Modules Version
NVIDIA-SMI 550.54.15 Driver Version: 550.54.15 CUDA Version: 12.4
Please confirm this issue does not happen with the proprietary driver (of the same version). This issue tracker is only for bugs specific to the open kernel driver.
Operating System and Version
Description: Ubuntu 22.04.1 LTS
Kernel Release
5.15.0-60-generic #66-Ubuntu SMP Fri Jan 20 14:29:49 UTC 2023 x86_64 x86_64 x86_64 GNU/Linux
Please confirm you are running a stable release kernel (e.g. not a -rc). We do not accept bug reports for unreleased kernels.
Hardware: GPU
NVIDIA GeForce RTX 4090
Describe the bug
Enabling P2P capability on 8 RTX 4090 GPUs results in significantly lower performance in NCCL alltoall_perf tests compared to when P2P capability is disabled.
To Reproduce
Enabling P2P capability on two RTX 4090 GPUs significantly improves performance in the NCCL alltoall_perf tests compared to when P2P is disabled. However, when testing with eight GPUs, the performance gap between enabling and disabling P2P is much larger, with a severe performance drop when P2P is enabled. The relevant test data is as follows:
- P2P capability is enabled.
root@moons:~# nvidia-smi topo -p2p rw
GPU0 GPU1 GPU2 GPU3 GPU4 GPU5 GPU6 GPU7
GPU0 X OK OK OK OK OK OK OK
GPU1 OK X OK OK OK OK OK OK
GPU2 OK OK X OK OK OK OK OK
GPU3 OK OK OK X OK OK OK OK
GPU4 OK OK OK OK X OK OK OK
GPU5 OK OK OK OK OK X OK OK
GPU6 OK OK OK OK OK OK X OK
GPU7 OK OK OK OK OK OK OK X
Legend:
X = Self
OK = Status Ok
CNS = Chipset not supported
GNS = GPU not supported
TNS = Topology not supported
NS = Not supported
U = Unknown
- The
simpleP2P test passes.
Enabling peer access between GPU0 and GPU1...
Allocating buffers (64MB on GPU0, GPU1 and CPU Host)...
Creating event handles...
cudaMemcpyPeer / cudaMemcpy between GPU0 and GPU1: 21.11GB/s
Preparing host buffer and memcpy to GPU0...
Run kernel on GPU1, taking source data from GPU0 and writing to GPU1...
Run kernel on GPU0, taking source data from GPU1 and writing to GPU0...
Copy data back to host from GPU0 and verify results...
Disabling peer access...
Shutting down...
Test passed
The alltoall_perf test data for two GPUs with P2P disabled:
root@moons:~/nccl-tests-master# ./build/alltoall_perf -b 8 -e 128M -f 2 -g 2
# nThread 1 nGpus 2 minBytes 8 maxBytes 134217728 step: 2(factor) warmup iters: 5 iters: 20 agg iters: 1 validation: 1 graph: 0
#
# Using devices
# Rank 0 Group 0 Pid 120986 on moons device 0 [0x41] NVIDIA GeForce RTX 4090
# Rank 1 Group 0 Pid 120986 on moons device 1 [0x42] NVIDIA GeForce RTX 4090
#
# out-of-place in-place
# size count type redop root time algbw busbw #wrong time algbw busbw #wrong
# (B) (elements) (us) (GB/s) (GB/s) (us) (GB/s) (GB/s)
8 1 float none -1 11.87 0.00 0.00 0 11.64 0.00 0.00 N/A
16 2 float none -1 11.79 0.00 0.00 0 11.47 0.00 0.00 N/A
32 4 float none -1 11.92 0.00 0.00 0 11.46 0.00 0.00 N/A
64 8 float none -1 11.27 0.01 0.00 0 11.59 0.01 0.00 N/A
128 16 float none -1 11.45 0.01 0.01 0 11.34 0.01 0.01 N/A
256 32 float none -1 11.39 0.02 0.01 0 11.24 0.02 0.01 N/A
512 64 float none -1 11.24 0.05 0.02 0 11.18 0.05 0.02 N/A
1024 128 float none -1 11.64 0.09 0.04 0 11.32 0.09 0.05 N/A
2048 256 float none -1 11.33 0.18 0.09 0 11.16 0.18 0.09 N/A
4096 512 float none -1 11.33 0.36 0.18 0 11.07 0.37 0.19 N/A
8192 1024 float none -1 11.70 0.70 0.35 0 10.89 0.75 0.38 N/A
16384 2048 float none -1 11.51 1.42 0.71 0 11.66 1.41 0.70 N/A
32768 4096 float none -1 12.96 2.53 1.26 0 12.88 2.54 1.27 N/A
65536 8192 float none -1 18.67 3.51 1.75 0 18.40 3.56 1.78 N/A
131072 16384 float none -1 18.12 7.23 3.62 0 17.79 7.37 3.68 N/A
262144 32768 float none -1 23.19 11.31 5.65 0 22.85 11.47 5.74 N/A
524288 65536 float none -1 34.97 14.99 7.50 0 34.77 15.08 7.54 N/A
1048576 131072 float none -1 56.78 18.47 9.23 0 56.60 18.52 9.26 N/A
2097152 262144 float none -1 101.1 20.74 10.37 0 100.6 20.85 10.42 N/A
4194304 524288 float none -1 188.1 22.29 11.15 0 186.7 22.47 11.23 N/A
8388608 1048576 float none -1 357.0 23.50 11.75 0 353.3 23.74 11.87 N/A
16777216 2097152 float none -1 619.9 27.07 13.53 0 576.4 29.11 14.55 N/A
33554432 4194304 float none -1 1214.1 27.64 13.82 0 1129.3 29.71 14.86 N/A
67108864 8388608 float none -1 2410.7 27.84 13.92 0 2221.7 30.21 15.10 N/A
134217728 16777216 float none -1 4813.9 27.88 13.94 0 4396.9 30.53 15.26 N/A
# Out of bounds values : 0 OK
# Avg bus bandwidth : 4.85887
#
The alltoall_perf test data for two GPUs with P2P enabled:
root@moons:~/nccl-tests-master# NCCL_P2P_LEVEL=SYS ./build/alltoall_perf -b 8 -e 128M -f 2 -g 2
# nThread 1 nGpus 2 minBytes 8 maxBytes 134217728 step: 2(factor) warmup iters: 5 iters: 20 agg iters: 1 validation: 1 graph: 0
#
# Using devices
# Rank 0 Group 0 Pid 121058 on moons device 0 [0x41] NVIDIA GeForce RTX 4090
# Rank 1 Group 0 Pid 121058 on moons device 1 [0x42] NVIDIA GeForce RTX 4090
#
# out-of-place in-place
# size count type redop root time algbw busbw #wrong time algbw busbw #wrong
# (B) (elements) (us) (GB/s) (GB/s) (us) (GB/s) (GB/s)
8 1 float none -1 12.28 0.00 0.00 0 11.80 0.00 0.00 N/A
16 2 float none -1 12.03 0.00 0.00 0 11.36 0.00 0.00 N/A
32 4 float none -1 11.90 0.00 0.00 0 11.57 0.00 0.00 N/A
64 8 float none -1 11.52 0.01 0.00 0 11.72 0.01 0.00 N/A
128 16 float none -1 11.59 0.01 0.01 0 11.47 0.01 0.01 N/A
256 32 float none -1 11.63 0.02 0.01 0 11.60 0.02 0.01 N/A
512 64 float none -1 11.78 0.04 0.02 0 11.50 0.04 0.02 N/A
1024 128 float none -1 11.69 0.09 0.04 0 11.29 0.09 0.05 N/A
2048 256 float none -1 11.80 0.17 0.09 0 11.52 0.18 0.09 N/A
4096 512 float none -1 11.83 0.35 0.17 0 12.07 0.34 0.17 N/A
8192 1024 float none -1 11.74 0.70 0.35 0 11.50 0.71 0.36 N/A
16384 2048 float none -1 11.64 1.41 0.70 0 12.08 1.36 0.68 N/A
32768 4096 float none -1 11.83 2.77 1.38 0 11.63 2.82 1.41 N/A
65536 8192 float none -1 12.23 5.36 2.68 0 11.91 5.50 2.75 N/A
131072 16384 float none -1 15.97 8.21 4.10 0 15.68 8.36 4.18 N/A
262144 32768 float none -1 20.26 12.94 6.47 0 20.13 13.03 6.51 N/A
524288 65536 float none -1 30.07 17.44 8.72 0 29.61 17.71 8.85 N/A
1048576 131072 float none -1 42.37 24.75 12.38 0 42.20 24.85 12.42 N/A
2097152 262144 float none -1 69.70 30.09 15.04 0 67.78 30.94 15.47 N/A
4194304 524288 float none -1 123.4 33.99 16.99 0 118.3 35.46 17.73 N/A
8388608 1048576 float none -1 223.1 37.59 18.80 0 222.1 37.77 18.88 N/A
16777216 2097152 float none -1 433.9 38.67 19.33 0 423.5 39.61 19.81 N/A
33554432 4194304 float none -1 849.8 39.49 19.74 0 828.4 40.51 20.25 N/A
67108864 8388608 float none -1 1686.1 39.80 19.90 0 1639.1 40.94 20.47 N/A
134217728 16777216 float none -1 3353.1 40.03 20.01 0 3261.5 41.15 20.58 N/A
# Out of bounds values : 0 OK
# Avg bus bandwidth : 6.75317
#
From the two-GPU test, it's evident that enabling P2P results in a significant performance boost.
The alltoall_perf test data for eight GPUs with P2P disabled:
root@moons:~/nccl-tests-master# ./build/alltoall_perf -b 8 -e 128M -f 2 -g 8
# nThread 1 nGpus 8 minBytes 8 maxBytes 134217728 step: 2(factor) warmup iters: 5 iters: 20 agg iters: 1 validation: 1 graph: 0
#
# Using devices
# Rank 0 Group 0 Pid 121126 on moons device 0 [0x41] NVIDIA GeForce RTX 4090
# Rank 1 Group 0 Pid 121126 on moons device 1 [0x42] NVIDIA GeForce RTX 4090
# Rank 2 Group 0 Pid 121126 on moons device 2 [0x43] NVIDIA GeForce RTX 4090
# Rank 3 Group 0 Pid 121126 on moons device 3 [0x44] NVIDIA GeForce RTX 4090
# Rank 4 Group 0 Pid 121126 on moons device 4 [0x61] NVIDIA GeForce RTX 4090
# Rank 5 Group 0 Pid 121126 on moons device 5 [0x62] NVIDIA GeForce RTX 4090
# Rank 6 Group 0 Pid 121126 on moons device 6 [0x63] NVIDIA GeForce RTX 4090
# Rank 7 Group 0 Pid 121126 on moons device 7 [0x64] NVIDIA GeForce RTX 4090
#
# out-of-place in-place
# size count type redop root time algbw busbw #wrong time algbw busbw #wrong
# (B) (elements) (us) (GB/s) (GB/s) (us) (GB/s) (GB/s)
0 0 float none -1 63.21 0.00 0.00 0 60.02 0.00 0.00 N/A
0 0 float none -1 62.41 0.00 0.00 0 62.10 0.00 0.00 N/A
32 1 float none -1 63.09 0.00 0.00 0 62.15 0.00 0.00 N/A
64 2 float none -1 63.61 0.00 0.00 0 63.87 0.00 0.00 N/A
128 4 float none -1 143.9 0.00 0.00 0 63.67 0.00 0.00 N/A
256 8 float none -1 63.84 0.00 0.00 0 62.88 0.00 0.00 N/A
512 16 float none -1 63.01 0.01 0.01 0 62.95 0.01 0.01 N/A
1024 32 float none -1 63.13 0.02 0.01 0 63.05 0.02 0.01 N/A
2048 64 float none -1 65.01 0.03 0.03 0 63.81 0.03 0.03 N/A
4096 128 float none -1 64.14 0.06 0.06 0 63.38 0.06 0.06 N/A
8192 256 float none -1 62.95 0.13 0.11 0 62.81 0.13 0.11 N/A
16384 512 float none -1 64.39 0.25 0.22 0 63.06 0.26 0.23 N/A
32768 1024 float none -1 63.51 0.52 0.45 0 62.81 0.52 0.46 N/A
65536 2048 float none -1 64.23 1.02 0.89 0 63.03 1.04 0.91 N/A
131072 4096 float none -1 65.07 2.01 1.76 0 64.43 2.03 1.78 N/A
262144 8192 float none -1 75.25 3.48 3.05 0 74.81 3.50 3.07 N/A
524288 16384 float none -1 66.32 7.91 6.92 0 65.05 8.06 7.05 N/A
1048576 32768 float none -1 94.40 11.11 9.72 0 93.75 11.18 9.79 N/A
2097152 65536 float none -1 169.6 12.36 10.82 0 169.4 12.38 10.83 N/A
4194304 131072 float none -1 300.8 13.94 12.20 0 295.4 14.20 12.42 N/A
8388608 262144 float none -1 562.6 14.91 13.05 0 561.0 14.95 13.08 N/A
16777216 524288 float none -1 1058.8 15.85 13.86 0 1060.8 15.82 13.84 N/A
33554432 1048576 float none -1 1982.6 16.92 14.81 0 1988.3 16.88 14.77 N/A
67108864 2097152 float none -1 3902.6 17.20 15.05 0 3901.4 17.20 15.05 N/A
134217728 4194304 float none -1 7447.2 18.02 15.77 0 7472.3 17.96 15.72 N/A
# Out of bounds values : 0 OK
# Avg bus bandwidth : 4.76019
#
The alltoall_perf test data for eight GPUs with P2P enabled:
root@moons:~/nccl-tests-master# NCCL_P2P_LEVEL=SYS ./build/alltoall_perf -b 8 -e 128M -f 2 -g 8
# nThread 1 nGpus 8 minBytes 8 maxBytes 134217728 step: 2(factor) warmup iters: 5 iters: 20 agg iters: 1 validation: 1 graph: 0
#
# Using devices
# Rank 0 Group 0 Pid 121524 on moons device 0 [0x41] NVIDIA GeForce RTX 4090
# Rank 1 Group 0 Pid 121524 on moons device 1 [0x42] NVIDIA GeForce RTX 4090
# Rank 2 Group 0 Pid 121524 on moons device 2 [0x43] NVIDIA GeForce RTX 4090
# Rank 3 Group 0 Pid 121524 on moons device 3 [0x44] NVIDIA GeForce RTX 4090
# Rank 4 Group 0 Pid 121524 on moons device 4 [0x61] NVIDIA GeForce RTX 4090
# Rank 5 Group 0 Pid 121524 on moons device 5 [0x62] NVIDIA GeForce RTX 4090
# Rank 6 Group 0 Pid 121524 on moons device 6 [0x63] NVIDIA GeForce RTX 4090
# Rank 7 Group 0 Pid 121524 on moons device 7 [0x64] NVIDIA GeForce RTX 4090
#
# out-of-place in-place
# size count type redop root time algbw busbw #wrong time algbw busbw #wrong
# (B) (elements) (us) (GB/s) (GB/s) (us) (GB/s) (GB/s)
0 0 float none -1 62.06 0.00 0.00 0 60.23 0.00 0.00 N/A
0 0 float none -1 61.34 0.00 0.00 0 59.87 0.00 0.00 N/A
32 1 float none -1 64.00 0.00 0.00 0 62.55 0.00 0.00 N/A
64 2 float none -1 62.70 0.00 0.00 0 62.47 0.00 0.00 N/A
128 4 float none -1 63.38 0.00 0.00 0 61.81 0.00 0.00 N/A
256 8 float none -1 62.82 0.00 0.00 0 62.12 0.00 0.00 N/A
512 16 float none -1 63.87 0.01 0.01 0 62.01 0.01 0.01 N/A
1024 32 float none -1 62.26 0.02 0.01 0 62.37 0.02 0.01 N/A
2048 64 float none -1 63.28 0.03 0.03 0 63.23 0.03 0.03 N/A
4096 128 float none -1 63.83 0.06 0.06 0 62.95 0.07 0.06 N/A
8192 256 float none -1 63.94 0.13 0.11 0 62.09 0.13 0.12 N/A
16384 512 float none -1 63.99 0.26 0.22 0 63.94 0.26 0.22 N/A
32768 1024 float none -1 66.91 0.49 0.43 0 65.51 0.50 0.44 N/A
65536 2048 float none -1 122.3 0.54 0.47 0 120.3 0.54 0.48 N/A
131072 4096 float none -1 237.7 0.55 0.48 0 235.8 0.56 0.49 N/A
262144 8192 float none -1 464.4 0.56 0.49 0 459.7 0.57 0.50 N/A
524288 16384 float none -1 466.4 1.12 0.98 0 460.3 1.14 1.00 N/A
1048576 32768 float none -1 914.1 1.15 1.00 0 913.9 1.15 1.00 N/A
2097152 65536 float none -1 1776.9 1.18 1.03 0 1786.8 1.17 1.03 N/A
4194304 131072 float none -1 3445.6 1.22 1.07 0 3427.3 1.22 1.07 N/A
8388608 262144 float none -1 6377.0 1.32 1.15 0 6258.2 1.34 1.17 N/A
16777216 524288 float none -1 11991 1.40 1.22 0 11809 1.42 1.24 N/A
33554432 1048576 float none -1 23087 1.45 1.27 0 22581 1.49 1.30 N/A
67108864 2097152 float none -1 53267 1.26 1.10 0 53155 1.26 1.10 N/A
134217728 4194304 float none -1 106721 1.26 1.10 0 106371 1.26 1.10 N/A
# Out of bounds values : 0 OK
# Avg bus bandwidth : 0.492673
#
From the eight-GPU test, it's clear that enabling P2P causes a severe performance drop.
Does anyone have experience in addressing this performance degradation when enabling P2P for eight GPUs?
Bug Incidence
Always
nvidia-bug-report.log.gz
~
More Info
If more information is needed, I can provide it at any time.
NVIDIA Open GPU Kernel Modules Version
NVIDIA-SMI 550.54.15 Driver Version: 550.54.15 CUDA Version: 12.4
Please confirm this issue does not happen with the proprietary driver (of the same version). This issue tracker is only for bugs specific to the open kernel driver.
Operating System and Version
Description: Ubuntu 22.04.1 LTS
Kernel Release
5.15.0-60-generic #66-Ubuntu SMP Fri Jan 20 14:29:49 UTC 2023 x86_64 x86_64 x86_64 GNU/LinuxPlease confirm you are running a stable release kernel (e.g. not a -rc). We do not accept bug reports for unreleased kernels.
Hardware: GPU
NVIDIA GeForce RTX 4090
Describe the bug
Enabling P2P capability on 8 RTX 4090 GPUs results in significantly lower performance in NCCL alltoall_perf tests compared to when P2P capability is disabled.
To Reproduce
Enabling P2P capability on two RTX 4090 GPUs significantly improves performance in the NCCL
alltoall_perftests compared to when P2P is disabled. However, when testing with eight GPUs, the performance gap between enabling and disabling P2P is much larger, with a severe performance drop when P2P is enabled. The relevant test data is as follows:simpleP2Ptest passes.The
alltoall_perftest data for two GPUs with P2P disabled:The
alltoall_perftest data for two GPUs with P2P enabled:From the two-GPU test, it's evident that enabling P2P results in a significant performance boost.
The
alltoall_perftest data for eight GPUs with P2P disabled:The
alltoall_perftest data for eight GPUs with P2P enabled:From the eight-GPU test, it's clear that enabling P2P causes a severe performance drop.
Does anyone have experience in addressing this performance degradation when enabling P2P for eight GPUs?
Bug Incidence
Always
nvidia-bug-report.log.gz
~
More Info
If more information is needed, I can provide it at any time.