Skip to content

Compare a vector at a time when taking the minimum of a float span - #134407

Closed
tahakocal wants to merge 2 commits into
dotnet:mainfrom
tahakocal:perf/linq-min-float
Closed

tahakocal wants to merge 2 commits into
dotnet:mainfrom
tahakocal:perf/linq-min-float

Conversation

@tahakocal

Copy link
Copy Markdown

Min over a span of float or double walks it one element at a time:

value = span[0];
for (int i = 1; (uint)i < (uint)span.Length; i++)
{
    T current = span[i];
    if (current < value) { value = current; }
    else if (T.IsNaN(current)) { return current; }
}

The integer overloads reach the vectorized MemoryExtensions.Min, but the floating-point ones do not, because their NaN behaviour differs from the comparer-based one.

Change

Spans of at least two vectors are now compared a vector at a time, with the lanes reduced once the loop ends. Two details of the sequential behaviour are preserved deliberately:

  • NaN. The walk returns the first NaN it meets. Rather than locating it inside a vector, the loop abandons the vectorized block as soon as a vector contains any NaN and hands the whole span to the sequential walk, which reports the same element it would have reported before. NaN inputs are therefore never slower than a single extra pass, and the common NaN-free case pays one compare per vector.
  • Signed zero. The walk keeps the first of two equal values, and -0.0 == +0.0, so [+0.0, -0.0] returns +0.0 while [-0.0, +0.0] returns -0.0. A vector reduction can keep either, so when the result is zero the first zero in the span is returned.

That second point is not theoretical: before I added it, fuzzing found 5 cases in 80,000 where the vectorized result was -0.0 and the sequential one +0.0, which is observable through double.IsNegative and division.

Max is left alone in this PR. Its span path skips leading NaNs, returns the last element when every element is NaN, and then ignores NaNs — Vector.Max propagates them instead, so it needs a different shape (substituting negative infinity for NaN lanes) and deserves its own change with its own measurements.

Measurements

Standalone harness (I could not build the runtime on this machine), Apple M-series arm64, Vector128. Minimum of 7 rounds, ns per call:

type length before after speedup
float 128 45.6 12.3 3.69x
float 1,024 370.7 95.0 3.90x
float 8,192 2,801.7 909.4 3.08x
float 1,000,000 342,258 117,754 2.91x
double 128 44.9 23.3 1.92x
double 1,024 375.3 210.6 1.78x
double 1,000,000 347,200 236,250 1.47x

float gains more because a 128-bit vector holds four of them against two doubles.

Correctness

Differential against the current implementation in the harness: 80,000 cases over float and double, lengths 1–120, with NaN, +0.0, -0.0, both infinities and random values seeded into the data — 0 value mismatches and 0 signed-zero differences, the latter only after the zero handling described above was added.

I have not run the System.Linq test suite locally for the reason above, so CI is the first full validation.

Min over a span of float or double walks it one element at a time, while the
integer overloads reach the vectorized MemoryExtensions.Min. Compare a vector
at a time instead, reducing the lanes once the loop ends.

The sequential walk returns the first NaN it meets, so the vectorized loop
abandons the block and hands the whole span back to that walk as soon as a
vector contains one, rather than trying to locate it. It also keeps the first
of two equal values, which matters only for zero, since negative and positive
zero compare equal while the reduction may keep either: when the result is
zero, the first zero in the span is returned.

Measured on arm64 with Vector128: 3.9x for float and 1.8x for double at 1K
elements, 2.9x and 1.5x at a million.
@dotnet-policy-service dotnet-policy-service Bot added the community-contribution Indicates that the PR has been added by a community member label Sep 22, 2026
@github-actions github-actions Bot added the area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI label Sep 22, 2026
@azure-pipelines

Copy link
Copy Markdown
Azure Pipelines:
Successfully started running 3 pipeline(s).
13 pipeline(s) were filtered out due to trigger conditions.
There may be pipelines that require an authorized user to comment /azp run to run.

@dotnet-policy-service

Copy link
Copy Markdown
Contributor

Tagging subscribers to this area: @JulieLeeMSFT, @jakobbotsch
See info in area-owners.md if you want to be subscribed.

// appears, since the first NaN is the result and the walk already reports it.
if (Vector128.IsHardwareAccelerated && Vector128<T>.IsSupported && span.Length >= Vector128<T>.Count * 2)
{
ref T first = ref MemoryMarshal.GetReference(span);

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Same here: Rewrite to safe code

PS: Is it worth splitting this into 3 PRs ?

@EgorBo EgorBo added area-System.Runtime and removed area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI labels Sep 22, 2026
@dotnet-policy-service

Copy link
Copy Markdown
Contributor

Tagging subscribers to this area: @dotnet/area-system-runtime
See info in area-owners.md if you want to be subscribed.

@tahakocal

Copy link
Copy Markdown
Author

Switched to safe loads: the vectors now come from Vector128.Create(span.Slice(i)) instead of MemoryMarshal.GetReference plus LoadUnsafe, matching how the integer path in MemoryExtensions.MinMax.cs loads. The System.Runtime.InteropServices using is gone with it.

Correctness re-checked on the safe form: 80,000 cases, 0 value mismatches and 0 signed-zero differences.

One thing you should know before deciding, because it changes the numbers in the description: the bounds check is not free at small and medium lengths. Measuring both load forms in the same process (arm64, Vector128, minimum of 7 rounds, ns per call):

scalar safe load unsafe load
float, 128 48.4 18.1 (2.68x) 12.7 (3.80x)
float, 1,024 415.6 143.1 (2.90x) 102.4 (4.06x)
float, 8,192 3,150.7 1,049.0 (3.00x) 1,010.9 (3.12x)
float, 1,000,000 382,566 128,500 (2.98x) 127,850 (2.99x)
double, 128 48.5 33.6 (1.44x) 24.7 (1.96x)
double, 1,024 411.4 271.2 (1.52x) 229.0 (1.80x)
double, 1,000,000 385,378 257,274 (1.50x) 253,828 (1.52x)

So the two converge once the span is large and the check is amortized, and the safe form gives up roughly a quarter of the gain at a thousand elements. (My 65,536 row came out at 2.00x safe against 3.28x unsafe, which does not fit its neighbours on either side, so I am treating that one as noise rather than signal on this machine.)

The description's numbers were taken on the unsafe form and I have not edited them yet — tell me which form you want and I will make the description match the code.

@EgorBo

EgorBo commented Sep 22, 2026

Copy link
Copy Markdown
Member

One thing you should know before deciding, because it changes the numbers in the description: the bounds check is not free at small and medium lengths.

There should be no bound checks if it's written in a more idiomatic and safe way (not how this PR is written). see #127506

@tahakocal

Copy link
Copy Markdown
Author

Also built the runtime locally since my earlier note: System.Linq.Tests on osx-arm64 Release from this branch, with the safe loads in place: 52,260 total, 0 errors, 0 failed, 8 skipped.

@tahakocal

Copy link
Copy Markdown
Author

Folded into #134409 as you suggested, so the whole float min/max change is in one place rather than spread over three PRs. The bounds-check point is addressed there too — walking the span forward instead of re-slicing by index removes the check, and the numbers are in that PR. Closing this one.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

area-System.Runtime community-contribution Indicates that the PR has been added by a community member

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants