Conversation
CountValueType extracts a bitmask from the comparison result and population counts it once per vector. Subtracting the comparison mask from a vector of counts does the same work with one instruction and no round trip out of the vector registers, and the lanes are summed once at the end. A lane cannot overflow for 32-bit and wider elements, since a span holds fewer than int.MaxValue elements and each lane sees at most one of every VectorXx<T>.Count of them. Byte and 16-bit elements keep the existing path, where a lane would overflow after 255 or 32767 matches. CountValueType is only instantiated with byte, short, int and long, as MemoryExtensions.Count reinterprets the span by element size before calling it.
|
Azure Pipelines: Successfully started running 3 pipeline(s). 13 pipeline(s) were filtered out due to trigger conditions. There may be pipelines that require an authorized user to comment /azp run to run. |
|
Tagging subscribers to this area: @JulieLeeMSFT, @jakobbotsch |
|
Tagging subscribers to this area: @dotnet/area-system-runtime |
|
The description says I could not run the suite locally; that is no longer true. Built the runtime and ran |
|
Closing this one: measured end to end against a locally built runtime, it is a regression for the most common element type, and my earlier numbers were wrong. Method: a console app calling only
So counting Scoping the change to 8-byte elements would leave a real but small gain on the less common type, which does not seem worth the extra branch in |
SpanHelpers.CountValueType, which backsMemoryExtensions.Count(span, value), extracts a bitmask from each comparison and population counts it:A comparison result is all-ones for a match, so subtracting it from a vector of counts accumulates the matches directly, with one instruction and without leaving the vector registers each iteration. The lanes are summed once at the end.
Change
For elements of 32 bits and wider, each of the three widths now accumulates
counts -= Vector.Equals(...)and finishes withVector.Sum. Byte and 16-bit elements keep the existing bitmask path, where a lane would overflow after 255 or 32,767 matches; for 32-bit and wider lanes overflow is impossible, since a span holds fewer thanint.MaxValueelements and each lane sees at most one of everyVectorXx<T>.Countof them.CountValueTypeis only ever instantiated withbyte,short,intandlong—MemoryExtensions.Countreinterprets the span by element size before calling it — so the vector arithmetic is always on a supported primitive.The trailing masked vector is unchanged.
Measurements
Benchmarked as an extraction into a standalone harness (I could not build the runtime on this machine), on Apple M-series arm64,
Vector128. Minimum of 7 rounds, ns per call:Caveat I want to be explicit about: these numbers are arm64, where
ExtractMostSignificantBitshas no single-instruction equivalent and expands to a sequence. On x64 it maps to amovmskinstruction, so the win there should be smaller — my change still replaces two operations (extract + popcount) with one (subtract) inside the loop and moves the reduction out of it, but I have no x64 hardware to verify, and I would not want this taken on the arm64 numbers alone. If the perf lab shows a regression on x64 the type guard could be narrowed further, or the change dropped.Correctness
Differential against a scalar oracle in the harness: 80,000 cases over
intandlong, lengths 8–207, values chosen so matches are dense (about one in three) — 0 mismatches. The lengths cover spans shorter than one vector, exact multiples, and every remainder.I have not run the CoreLib test suite locally for the reason above, so CI is the first full validation.