Adding new experimental copy-only GFX kernel, gfxsweep update - #277
Merged
Conversation
nileshnegi
reviewed
Apr 29, 2026
Contributor
There was a problem hiding this comment.
Pull request overview
Adds an experimental, copy-only GFX kernel and exposes kernel selection/auto-selection via GFX_KERNEL, plus expands the gfxsweep preset to sweep kernels and timing modes.
Changes:
- Introduce
GpuCopyKernelalongside the existing reduction kernel, with auto/forced selection support. - Add
GFX_KERNELenv var plumbing and config validation for selecting GFX kernels. - Update
gfxsweeppreset to sweep over kernels (KERNELS) and support multiple timing modes (TIMING_MODE).
Reviewed changes
Copilot reviewed 5 out of 5 changed files in this pull request and generated 11 comments.
Show a summary per file
| File | Description |
|---|---|
src/header/TransferBench.hpp |
Adds GFX kernel type enum, kernel eligibility/selection logic, new copy-only kernel implementation, and dispatch support. |
src/client/EnvVars.hpp |
Adds GFX_KERNEL env var parsing/printing and maps it into cfg.gfx.gfxKernel. |
src/client/Presets/GfxSweep.hpp |
Adds KERNELS sweep dimension, TIMING_MODE selection/auto-mode, and updates output formatting. |
src/client/Presets/HbmBandwidth.hpp |
Reorders/adjusts NUM_ITERATIONS validation logic. |
CHANGELOG.md |
Documents the new GFX_KERNEL option. |
💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.
gilbertlee-amd
force-pushed
the
GpuCopyKernel
branch
from
April 30, 2026 01:50
edbaf1f to
aff69e2
Compare
Contributor
There was a problem hiding this comment.
Pull request overview
Copilot reviewed 21 out of 21 changed files in this pull request and generated 5 comments.
💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.
Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>
nileshnegi
reviewed
Apr 30, 2026
nileshnegi
approved these changes
Apr 30, 2026
nileshnegi
added a commit
that referenced
this pull request
May 2, 2026
- Initial pod communication support (#235) - cuda + MNNVL update & pod presets (#241) - Increase CQ size for high qps (#244) - fix hang when NVML is present but fabricmanager isnt (#246) - Adding nica2a preset (#248) - Adding HBM read bandwidth preset (#250) - Pod Ring preset (#251) - gfxsweep preset (#254) (#256) - Adding Batched DMA support (hipMemcpyBatchAsync), and bmasweep preset (#255) - Adding a wallclock consistency detection preset (#258) - Adding smoketest preset for simple correctness tests (#266) - Help / envvars / presets presets (#267) - Modernize CMake build (#268) - Replace version-based pod/amd-smi detection with compile-time API probes (#269) - Fix collective mismatch hangs in multi-rank error paths (#270) - Fix SHOW_ITERATIONS table truncation with multiple transfers per executor (#271) - Reformat a2asweep output to match gfxsweep style (#272) - Gfx sweep update (#274) - Increasing flush frequency in smoketest (#275) - Adding new experimental copy-only GFX kernel, gfxsweep update (#277) - Fixes for cuMem compilation and invalid device ordinal (#278) - Simplifying socket connect, allow for using host address (#279) - Updating podring to run on single node without need to force single pod (#280) - Adding SHOW_PERCENTILES to show extra per-iteration statistics (#281) --------- Co-authored-by: AtlantaPepsi <timhu102@gmail.com> Co-authored-by: Pak Nin Lui <pak.lui@amd.com> Co-authored-by: pierreantoineH <PierreAntoine.Harraud@amd.com> Co-authored-by: Nilesh M Negi <Nilesh.Negi@amd.com> Co-authored-by: Claude <claude@anthropic.com> Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>
nileshnegi
added a commit
that referenced
this pull request
May 2, 2026
- Initial pod communication support (#235) - cuda + MNNVL update & pod presets (#241) - Increase CQ size for high qps (#244) - fix hang when NVML is present but fabricmanager isnt (#246) - Adding nica2a preset (#248) - Adding HBM read bandwidth preset (#250) - Pod Ring preset (#251) - gfxsweep preset (#254) (#256) - Adding Batched DMA support (hipMemcpyBatchAsync), and bmasweep preset (#255) - Adding a wallclock consistency detection preset (#258) - Adding smoketest preset for simple correctness tests (#266) - Help / envvars / presets presets (#267) - Modernize CMake build (#268) - Replace version-based pod/amd-smi detection with compile-time API probes (#269) - Fix collective mismatch hangs in multi-rank error paths (#270) - Fix SHOW_ITERATIONS table truncation with multiple transfers per executor (#271) - Reformat a2asweep output to match gfxsweep style (#272) - Gfx sweep update (#274) - Increasing flush frequency in smoketest (#275) - Adding new experimental copy-only GFX kernel, gfxsweep update (#277) - Fixes for cuMem compilation and invalid device ordinal (#278) - Simplifying socket connect, allow for using host address (#279) - Updating podring to run on single node without need to force single pod (#280) - Adding SHOW_PERCENTILES to show extra per-iteration statistics (#281) --------- Co-authored-by: Tim <43156029+AtlantaPepsi@users.noreply.github.com> Co-authored-by: Pak Nin Lui <pak.lui@amd.com> Co-authored-by: pierreantoineH <PierreAntoine.Harraud@amd.com> Co-authored-by: Nilesh M Negi <Nilesh.Negi@amd.com> Co-authored-by: Claude <claude@anthropic.com> Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Motivation
Adding a new GFX copy kernel that only supports single src single dst copies to investigate performance, as it may have less register pressure than the GFX reduction kernel.
Technical Details
The new copy kernel can be enabled by setting GFX_KERNEL=1. GFX_KERNEL=0 will force the GpuReduceKernel which remain default behavior. Setting GFX_KERNEL=-1 will switch to auto-mode, where GpuCopyKernel will be used if the Transfers are compatible single-source/single-destination copies.
The gfxsweep preset was also updated to allow for KERNELS option to sweep over GFX kernels, as well as adjusted to allow for different timing methods (TIMING_MODE: -1=auto, 0 = Aggregate CPU wall-clock, 1 = HipEvent executor time, 2 = GPU wallclock time).
Submission Checklist