Skip to content

Use a block search for TensorPrimitives.IndexOfMin and IndexOfMax - #134314

Closed
tahakocal wants to merge 1 commit into
dotnet:mainfrom
tahakocal:tensors-indexofminmax-block-search
Closed

tahakocal wants to merge 1 commit into
dotnet:mainfrom
tahakocal:tensors-indexofminmax-block-search

Conversation

@tahakocal

Copy link
Copy Markdown

Fixes #134045.

Why

IndexOfMinMaxCore keeps a running best vector plus a vector of per-lane indices. Every vector costs a compare and two selects, plus the index increment, and the index vector has to fit the lane width, which is why the Size4Plus/Size2/Size1 variants exist. The issue asked for the block search instead: reduce a block of vectors with the hardware min/max, then find the result inside the one block that won.

How

For integer element types IndexOfMinMaxCore now runs a two-pass search:

  • Pass 1 reduces each block of 32 vectors with the operator's vertical min/max into two independent accumulators, merges them, and keeps the first block whose horizontal aggregate beats the running best under the operator's own strict Compare. No index vectors and no blending.
  • Pass 2 scans only the winning block for the first element the best value does not beat, vectorized through the operator's existing vector Compare plus IndexOfFirstMatch.

Because pass 2 uses !TOperator.Compare(best, x[i]) rather than a bitwise match, the operator's own tie rule decides, so the result stays the first index, and the magnitude operators (where equal magnitudes prefer a sign) work unchanged. The vertical operation is the existing MinOperator/MaxOperator/MinMagnitudeOperator/MaxMagnitudeOperator, reached through three new MinMax members on IIndexOfMinMaxOperator<T>, so each index operator keeps pairing with the same aggregation operator its Aggregate already uses.

Two restrictions, both measured:

  • float/double keep the existing implementation. The block loop would have to check for NaN per vector anyway (MinMaxCore does), and -0/+0 and the first-NaN rule need care, so that is better as a separate change.
  • The block path requires at least one full block at the widest accelerated vector width (32 * Vector512<T>.Count where Vector512 is accelerated, and so on). Below that the second pass costs more than it saves - at 8-64 elements the block search measured 0.6-0.8x of the current code - and gating on the widest width also keeps every input on at least as wide a vector as it uses today.

Block bounds are computed as blockStart + Math.Min(blockLength, x.Length - blockStart) so that they do not overflow for spans close to int.MaxValue elements.

Test Plan

System.Numerics.Tensors tests pass: 5724 total, 0 failed (dotnet build /t:Test -c Release -f net11.0, osx-arm64).

Added IndexOfMin_FirstOfEqualMinimumsReturned, IndexOfMax_FirstOfEqualMaximumsReturned, IndexOfMinMagnitude_FirstOfEqualMagnitudesReturned, IndexOfMaxMagnitude_FirstOfEqualMagnitudesReturned and an *AcrossBlocks variant of each. Helpers.TensorLengths only goes to 256, which is below the threshold for most element types, so the *AcrossBlocks tests use Helpers.TensorLengthsSpanningBlocks (511 to 4103) to cover the block boundaries, multi-block inputs, and a winner in the scalar tail. The existing IndexOf*_AllLengths tests accept any index holding an equal value, so nothing pinned the first-index rule that the block search has to preserve. Relaxing the block comparison to TOperator.Compare(blockBest, best) || blockBest == best (a plausible way to get this wrong) fails the new tests on 12 instantiations and nothing else.

Differential check against a build of unmodified main: 20,000 random cases per element type for int, uint, long, short, sbyte, byte, nint, with narrow value ranges (many duplicates), full ranges, and planted MinValue/MaxValue, comparing all four APIs. Every result matches the current implementation.

Benchmarks

Apple M-series, osx-arm64, Vector128 (no AVX2/AVX-512 here), release runtime, GElem/s, higher is better. Random data, best of 5 rounds.

T API N before after ratio
int IndexOfMin 65536 4.41 22.44 5.1x
int IndexOfMin 1048576 4.41 22.80 5.2x
int IndexOfMax 65536 4.42 18.97 4.3x
int IndexOfMinMagnitude 65536 1.17 3.05 2.6x
long IndexOfMin 65536 2.21 10.96 5.0x
short IndexOfMin 65536 8.81 49.14 5.6x
byte IndexOfMin 65536 8.83 96.39 10.9x
int IndexOfMin 1024 5.39 22.69 4.2x
int IndexOfMin 512 5.86 21.09 3.6x
int IndexOfMin 128 7.79 13.28 1.7x

Below the threshold nothing changes, since those lengths keep taking the existing path (int at N=8/16/32/64: 4.61/7.17/8.61/8.54 before, 4.65/7.28/8.99/9.19 after; byte: 2.02/6.44/9.33/11.63 before, 2.13/6.52/9.61/12.28 after).

A second harness, written independently and alternating the two builds round by round with the median taken, puts the same cases slightly lower: 3.5-4.8x for int/long/short at 1024 and above, 4.9-7.3x for byte, and 0.99-1.01x below the threshold. Right at the threshold the gain is smaller, since the second pass always rescans one block: 1.2-2.0x for int at 128, 1.6-2.2x for short at 256.

Two accumulators are worth keeping: with a single accumulator the same code reaches only 11.97 GElem/s for int at N=65536 versus 21 with two.

IndexOfMinMaxVectorsPerBlock is 32 because that is what the issue's own measurements favour (Blocks32 there, with Blocks256 no better); the worst case is one extra read of a single block, about 1.03x the reads.

The three widths are written out separately, like the IndexOfMinMaxVectorized* helpers they sit next to, because ISimdVector<TSelf, T> is internal to CoreLib and not linked into this assembly.

I do not have AVX-512 hardware, so the mask-register blending the issue describes is not what I measured here; the gain above comes from dropping the index vectors and the per-vector selects.

Reduce each block of 32 vectors with the operator's vertical min/max and
then search only the winning block, instead of carrying per-lane index
vectors and blending them on every vector. Applies to IndexOfMin,
IndexOfMax, IndexOfMinMagnitude and IndexOfMaxMagnitude for integer
element types, for inputs of at least one full block.
@dotnet-policy-service dotnet-policy-service Bot added the community-contribution Indicates that the PR has been added by a community member label Sep 20, 2026
@dotnet-policy-service

Copy link
Copy Markdown
Contributor

Tagging subscribers to this area: @dotnet/area-system-numerics
See info in area-owners.md if you want to be subscribed.

@azure-pipelines

Copy link
Copy Markdown
Azure Pipelines:
Successfully started running 3 pipeline(s).
13 pipeline(s) were filtered out due to trigger conditions.
There may be pipelines that require an authorized user to comment /azp run to run.

@lilinus lilinus left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I believe this is the same change as in #133969

@tahakocal

Copy link
Copy Markdown
Author

Built the runtime locally and ran the suite on this branch — System.Numerics.Tensors.Tests, osx-arm64 Release: 5,724 total, 0 errors, 0 failed, 0 skipped.

@tahakocal

Copy link
Copy Markdown
Author

You are right, thank you for the pointer — #133969 is the same change and it predates this one by five days. Closing in favour of it.

For whatever it is worth to that PR, two things I ran into while doing this that may be worth checking there as well:

  • The block loop needs a subtractive bound (i <= limit - N * Count rather than i + N * Count <= limit), otherwise the index can wrap for spans near int.MaxValue and the last block reads past the end.
  • Comparing a NaN mask against VectorXxx<T>.AllBitsSet rather than Zero is a trap for float/double, since AllBitsSet is itself a NaN bit pattern and the comparison then takes NaN semantics. Differential fuzzing caught that one for me in a different PR.

Sorry for the duplicate work.

@tahakocal tahakocal closed this Sep 22, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

area-System.Numerics community-contribution Indicates that the PR has been added by a community member

Projects

None yet

Development

Successfully merging this pull request may close these issues.

TensorPrimitives IndexOfMin & Max produce suboptimal code and use a suboptimal algorithm with index vector searching

2 participants