Skip to content

Accumulate matches in a vector when counting 32-bit and wider elements - #134401

Closed
tahakocal wants to merge 1 commit into
dotnet:mainfrom
tahakocal:perf/count-vector-accumulator
Closed

tahakocal wants to merge 1 commit into
dotnet:mainfrom
tahakocal:perf/count-vector-accumulator

Conversation

@tahakocal

Copy link
Copy Markdown

SpanHelpers.CountValueType, which backs MemoryExtensions.Count(span, value), extracts a bitmask from each comparison and population counts it:

while (Unsafe.IsAddressLessThan(ref current, ref oneVectorAwayFromEnd))
{
    count += BitOperations.PopCount(Vector128.Equals(Vector128.LoadUnsafe(ref current), targetVector).ExtractMostSignificantBits());
    current = ref Unsafe.Add(ref current, Vector128<T>.Count);
}

A comparison result is all-ones for a match, so subtracting it from a vector of counts accumulates the matches directly, with one instruction and without leaving the vector registers each iteration. The lanes are summed once at the end.

Change

For elements of 32 bits and wider, each of the three widths now accumulates counts -= Vector.Equals(...) and finishes with Vector.Sum. Byte and 16-bit elements keep the existing bitmask path, where a lane would overflow after 255 or 32,767 matches; for 32-bit and wider lanes overflow is impossible, since a span holds fewer than int.MaxValue elements and each lane sees at most one of every VectorXx<T>.Count of them.

CountValueType is only ever instantiated with byte, short, int and long — MemoryExtensions.Count reinterprets the span by element size before calling it — so the vector arithmetic is always on a supported primitive.

The trailing masked vector is unchanged.

Measurements

Benchmarked as an extraction into a standalone harness (I could not build the runtime on this machine), on Apple M-series arm64, Vector128. Minimum of 7 rounds, ns per call:

type length bitmask + popcount vector counts speedup
int 128 19.3 8.8 2.20x
int 1,024 154.0 82.5 1.87x
int 8,192 1,260.6 882.2 1.43x
int 1,000,000 152,644 113,240 1.35x
long 128 35.5 15.7 2.26x
long 8,192 2,288.2 1,806.0 1.27x
long 1,000,000 280,166 224,060 1.25x

Caveat I want to be explicit about: these numbers are arm64, where ExtractMostSignificantBits has no single-instruction equivalent and expands to a sequence. On x64 it maps to a movmsk instruction, so the win there should be smaller — my change still replaces two operations (extract + popcount) with one (subtract) inside the loop and moves the reduction out of it, but I have no x64 hardware to verify, and I would not want this taken on the arm64 numbers alone. If the perf lab shows a regression on x64 the type guard could be narrowed further, or the change dropped.

Correctness

Differential against a scalar oracle in the harness: 80,000 cases over int and long, lengths 8–207, values chosen so matches are dense (about one in three) — 0 mismatches. The lengths cover spans shorter than one vector, exact multiples, and every remainder.

I have not run the CoreLib test suite locally for the reason above, so CI is the first full validation.

CountValueType extracts a bitmask from the comparison result and population
counts it once per vector. Subtracting the comparison mask from a vector of
counts does the same work with one instruction and no round trip out of the
vector registers, and the lanes are summed once at the end.

A lane cannot overflow for 32-bit and wider elements, since a span holds fewer
than int.MaxValue elements and each lane sees at most one of every
VectorXx<T>.Count of them. Byte and 16-bit elements keep the existing path,
where a lane would overflow after 255 or 32767 matches.

CountValueType is only instantiated with byte, short, int and long, as
MemoryExtensions.Count reinterprets the span by element size before calling it.
@dotnet-policy-service dotnet-policy-service Bot added the community-contribution Indicates that the PR has been added by a community member label Sep 22, 2026
@github-actions github-actions Bot added the area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI label Sep 22, 2026
@azure-pipelines

Copy link
Copy Markdown
Azure Pipelines:
Successfully started running 3 pipeline(s).
13 pipeline(s) were filtered out due to trigger conditions.
There may be pipelines that require an authorized user to comment /azp run to run.

@dotnet-policy-service

Copy link
Copy Markdown
Contributor

Tagging subscribers to this area: @JulieLeeMSFT, @jakobbotsch
See info in area-owners.md if you want to be subscribed.

@EgorBo EgorBo added area-System.Runtime and removed area-CodeGen-coreclr CLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI labels Sep 22, 2026
@dotnet-policy-service

Copy link
Copy Markdown
Contributor

Tagging subscribers to this area: @dotnet/area-system-runtime
See info in area-owners.md if you want to be subscribed.

@tahakocal

Copy link
Copy Markdown
Author

The description says I could not run the suite locally; that is no longer true. Built the runtime and ran System.Memory.Tests on osx-arm64 Release against a System.Private.CoreLib rebuilt from this branch: 52,929 total, 0 errors, 0 failed, 1 skipped. I verified the md5 of the CoreLib in artifacts/bin/testhost matched the freshly compiled one first, since a plain test-project build does not rebuild CoreLib.

@tahakocal

Copy link
Copy Markdown
Author

Closing this one: measured end to end against a locally built runtime, it is a regression for the most common element type, and my earlier numbers were wrong.

Method: a console app calling only MemoryExtensions.Count(span, value), run with artifacts/bin/testhost/.../dotnet so it binds to the locally built CoreLib, DOTNET_TieredCompilation=0, minimum of 9 rounds, on osx-arm64. Baseline is a clean build of main measured the same way.

element length main this branch
int 1,024 77.9 ns 93.9 ns 0.83x
int 8,192 591.4 ns 974.1 ns 0.61x
int 1,000,000 75,082 ns 123,724 ns 0.61x
long 1,024 288.7 ns 213.6 ns 1.35x
long 8,192 2,316.1 ns 1,987.0 ns 1.17x
long 1,000,000 290,008 ns 247,830 ns 1.17x
byte 1,000,000 18,918 ns 18,698 ns 1.01x

byte is the control — it keeps the existing path by design and does not move, which says the measurement is picking up this change and not something else.

So counting int gets about 40% slower while long gains a little. My earlier figures came from an extracted copy of the loop rather than the real one, and that copy did not reproduce this at all — it reported 1.35x for int. Same lesson as in #134400: the extracted-loop harness was not representative.

Scoping the change to 8-byte elements would leave a real but small gain on the less common type, which does not seem worth the extra branch in CountValueType, so I would rather withdraw it than argue for it. Sorry for the noise.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

area-System.Runtime community-contribution Indicates that the PR has been added by a community member

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants