Conversation
Max over a span of float or double walks it one element at a time. Compare a vector at a time instead, once the leading NaNs have been skipped and a first candidate is in hand. NaN lanes are replaced with negative infinity before the comparison, because the sequential walk ignores a NaN that is not the leading one while Vector128.Max would propagate it. That substitution costs two operations per vector, which only pays off when a vector holds at least four elements, so double keeps the sequential walk: measured 0.93x to 1.01x for double against 1.97x to 3.05x for float. The walk also keeps the first of two equal values, which matters only for zero, since negative and positive zero compare equal while the reduction may keep either: when the result is zero, the first zero in the span is returned.
|
Azure Pipelines: Successfully started running 3 pipeline(s). 13 pipeline(s) were filtered out due to trigger conditions. There may be pipelines that require an authorized user to comment /azp run to run. |
|
Tagging subscribers to this area: @agocke |
| if (Vector128.IsHardwareAccelerated && Vector128<T>.IsSupported && | ||
| Vector128<T>.Count >= 4 && span.Length - i >= Vector128<T>.Count * 2) | ||
| { | ||
| ref T first = ref MemoryMarshal.GetReference(span); |
|
Tagging subscribers to this area: @dotnet/area-system-runtime |
|
Switched to safe loads — Correctness re-checked on the safe form: 80,080 cases, 0 value mismatches and 0 signed-zero differences, including every all-NaN and leading-NaN span from length 1 to 40. The float numbers come out slightly lower than the description's, which were taken on the unsafe form: 2.56x at 128, 2.67x at 1,024 and 2.10x at 8,192. I measured the two load forms side by side only for #134407, where the bounds check costs about a quarter of the gain at a thousand elements and nothing at a million — I would expect the same shape here, but I have not measured it for Happy to update the description to the safe-form numbers once you confirm which form you want to keep. |
Note: I already commented about this in the other PR - so I highly recommend either working PR by PR or group similar chnages in one PR. |
|
Also built the runtime locally since my earlier note: |
|
Folded into #134409 as you suggested, so the whole float min/max change is in one place rather than spread over three PRs. The bounds-check point is addressed there too — walking the span forward instead of re-slicing by index removes the check, and the numbers are in that PR. Closing this one. |
Companion to #134407, which did the same for
Min.Maxover a span offloatordoublewalks it one element at a time:Change
Once the leading NaNs are skipped and a first candidate is in hand, the rest of the span is compared a vector at a time.
Two details of the sequential behaviour are preserved:
span[i] > valueis false for them, whileVector128.Maxwould propagate one into the result. NaN lanes are therefore replaced with negative infinity before the comparison. The leading-NaN scan and the all-NaN case (which returns the last element) are untouched.-0.0 == +0.0, so when the result is zero the first zero in the span is returned rather than whichever one the reduction kept.The substitution is why this is guarded on
Vector128<T>.Count >= 4, which in practice meansfloatand notdouble: replacing the NaN lanes costs a compare and a select per vector, and with only two elements per vector that cancels the gain. Measured fordouble: 1.43x at 1K, 1.05x at 8K, 1.01x at 64K and 0.93x at a million — sodoublekeeps the sequential walk. I would rather scope it by the measurement than ship a change that is a wash or a small regression on half the types it touches.Measurements
Standalone harness (I could not build the runtime on this machine), Apple M-series arm64,
Vector128. Minimum of 7 rounds, ns per call:Correctness
Differential against the current implementation in the harness: 80,080 cases over
floatanddouble, lengths 1–120, with NaN,+0.0,-0.0, both infinities and random values seeded in, plus every all-NaN span and every leading-NaN span from length 1 to 40 — 0 value mismatches and 0 signed-zero differences.I have not run the System.Linq test suite locally for the reason above, so CI is the first full validation.