Adding smoketest preset for simple correctness tests - #266
Merged
Conversation
Contributor
There was a problem hiding this comment.
Pull request overview
Adds a new smoketest preset to TransferBench to run a small suite of correctness-focused DMA/GFX transfer checks across common patterns (H2D/D2H/D2D, broadcast, gather, all-to-all), with environment-variable knobs to control sizes, subexecutor counts, and test selection.
Changes:
- Introduces a new
smoketestpreset implementation and registers it in the preset dispatch map. - Adds EnvVars helpers for parsing string arrays and for printing string vectors.
- Updates changelog; includes a small validation-path tweak and whitespace cleanup.
Reviewed changes
Copilot reviewed 6 out of 6 changed files in this pull request and generated 6 comments.
Show a summary per file
| File | Description |
|---|---|
| src/client/Presets/SmokeTest.hpp | New smoketest preset logic, env-var knobs, and tabular output. |
| src/client/Presets/Presets.hpp | Registers the new smoketest preset. |
| src/client/EnvVars.hpp | Adds string-array env-var parsing + string-vector formatting helper. |
| src/header/TransferBench.hpp | Adjusts validation mismatch handling in ValidateAllTransfers(). |
| src/client/Utilities.hpp | Removes trailing whitespace. |
| CHANGELOG.md | Documents the new preset. |
💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.
nileshnegi
reviewed
Apr 25, 2026
nileshnegi
approved these changes
Apr 26, 2026
nileshnegi
added a commit
that referenced
this pull request
May 2, 2026
- Initial pod communication support (#235) - cuda + MNNVL update & pod presets (#241) - Increase CQ size for high qps (#244) - fix hang when NVML is present but fabricmanager isnt (#246) - Adding nica2a preset (#248) - Adding HBM read bandwidth preset (#250) - Pod Ring preset (#251) - gfxsweep preset (#254) (#256) - Adding Batched DMA support (hipMemcpyBatchAsync), and bmasweep preset (#255) - Adding a wallclock consistency detection preset (#258) - Adding smoketest preset for simple correctness tests (#266) - Help / envvars / presets presets (#267) - Modernize CMake build (#268) - Replace version-based pod/amd-smi detection with compile-time API probes (#269) - Fix collective mismatch hangs in multi-rank error paths (#270) - Fix SHOW_ITERATIONS table truncation with multiple transfers per executor (#271) - Reformat a2asweep output to match gfxsweep style (#272) - Gfx sweep update (#274) - Increasing flush frequency in smoketest (#275) - Adding new experimental copy-only GFX kernel, gfxsweep update (#277) - Fixes for cuMem compilation and invalid device ordinal (#278) - Simplifying socket connect, allow for using host address (#279) - Updating podring to run on single node without need to force single pod (#280) - Adding SHOW_PERCENTILES to show extra per-iteration statistics (#281) --------- Co-authored-by: AtlantaPepsi <timhu102@gmail.com> Co-authored-by: Pak Nin Lui <pak.lui@amd.com> Co-authored-by: pierreantoineH <PierreAntoine.Harraud@amd.com> Co-authored-by: Nilesh M Negi <Nilesh.Negi@amd.com> Co-authored-by: Claude <claude@anthropic.com> Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>
nileshnegi
added a commit
that referenced
this pull request
May 2, 2026
- Initial pod communication support (#235) - cuda + MNNVL update & pod presets (#241) - Increase CQ size for high qps (#244) - fix hang when NVML is present but fabricmanager isnt (#246) - Adding nica2a preset (#248) - Adding HBM read bandwidth preset (#250) - Pod Ring preset (#251) - gfxsweep preset (#254) (#256) - Adding Batched DMA support (hipMemcpyBatchAsync), and bmasweep preset (#255) - Adding a wallclock consistency detection preset (#258) - Adding smoketest preset for simple correctness tests (#266) - Help / envvars / presets presets (#267) - Modernize CMake build (#268) - Replace version-based pod/amd-smi detection with compile-time API probes (#269) - Fix collective mismatch hangs in multi-rank error paths (#270) - Fix SHOW_ITERATIONS table truncation with multiple transfers per executor (#271) - Reformat a2asweep output to match gfxsweep style (#272) - Gfx sweep update (#274) - Increasing flush frequency in smoketest (#275) - Adding new experimental copy-only GFX kernel, gfxsweep update (#277) - Fixes for cuMem compilation and invalid device ordinal (#278) - Simplifying socket connect, allow for using host address (#279) - Updating podring to run on single node without need to force single pod (#280) - Adding SHOW_PERCENTILES to show extra per-iteration statistics (#281) --------- Co-authored-by: Tim <43156029+AtlantaPepsi@users.noreply.github.com> Co-authored-by: Pak Nin Lui <pak.lui@amd.com> Co-authored-by: pierreantoineH <PierreAntoine.Harraud@amd.com> Co-authored-by: Nilesh M Negi <Nilesh.Negi@amd.com> Co-authored-by: Claude <claude@anthropic.com> Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Motivation
Adds a simple smoketest preset into TransferBench to perform simple correctness tests. Knobs exist to allow for more thorough tests, however defaults are chosen to try to keep run-times short.
Technical Details
This preset tests DMA/GFX copies between host and device, between two devices, as well as broadcast / gather and all to all.
It supports multi-rank, and has knobs for controlling the Transfer sizes to sweep (SIZE_LIST), as well as how many GFX subexecutors to use (GFX_SE_LIST). For performance reasons, SE_MAX_BYTES is used to limit the total amount of bytes a single GFX Subexecutor (CU/WGP) will Transfers - for example skipping tests copying 1GB using only 1 CU. TEST_LISTS can be used to specify a subset of tests to run, and accepts ranges such as 1-7. Copies are done in parallel
Test Plan
Testing of the various knobs were tried, as well as fault-injection to confirm that failures would be shown.
Test Result
Sample output: