Skip to content

[Validation Approach] Improve ConcurrentStack.TryPopRange performance under contention. - #134306

Draft
aw0lid wants to merge 1 commit into
dotnet:mainfrom
aw0lid:improve/ConcurrentStack-TryPopRange
Draft

aw0lid wants to merge 1 commit into
dotnet:mainfrom
aw0lid:improve/ConcurrentStack-TryPopRange

Conversation

@aw0lid

@aw0lid aw0lid commented Sep 20, 2026 •

Copy link
Copy Markdown
Contributor

Try to improve ConcurrentStack.TryPopRange performance under contention described in #100083.

Benchmark Setup

public class ConcurrentStackBenchmarks
{
    private readonly ConcurrentStack<int> _stack = new();
    private int[] _popRangeBuffer = null!;
    private bool _running;

    [Params(3, 10, 100, 1000, 10000, 100000, 1_000_000)]
    public int TryPopCount;

    [Params(0, 1, 2, 5, 8)]
    public int ParallelThreads;

    [GlobalSetup]
    public void Setup()
    {
        _popRangeBuffer = new int[TryPopCount];
        for (int i = 0; i < TryPopCount; i++)
        {
            _stack.Push(i);
        }

        _running = true;
        for (int i = 0; i < ParallelThreads; i++)
        {
            Task.Run(() =>
            {
                while (Volatile.Read(ref _running))
                {
                    _stack.Push(42);
                    _stack.TryPop(out _);
                }
            });
        }
    }

    [GlobalCleanup]
    public void Cleanup() => _running = false;

    [Benchmark]
    public void PopRange()
    {
        int popped = _stack.TryPopRange(_popRangeBuffer);
        if (popped > 0)
        {
            _stack.PushRange(_popRangeBuffer, 0, popped);
        }
    }
}

Environment

  • OS: Fedora Linux 44 (Workstation Edition).
  • CPU: Intel Core i5-6300U CPU @ 2.40GHz (Skylake, 1 CPU, 4 logical cores).
  • Runtime: .NET 11.0.0 (11.0.0-rc.1.26420.103), X64 RyuJIT x86-64-v3.

@dotnet-policy-service dotnet-policy-service Bot added the community-contribution Indicates that the PR has been added by a community member label Sep 20, 2026
@azure-pipelines

Copy link
Copy Markdown
Azure Pipelines:
Successfully started running 3 pipeline(s).
13 pipeline(s) were filtered out due to trigger conditions.
There may be pipelines that require an authorized user to comment /azp run to run.

@dotnet-policy-service

Copy link
Copy Markdown
Contributor

Tagging subscribers to this area: @dotnet/area-system-collections
See info in area-owners.md if you want to be subscribed.

@aw0lid
aw0lid force-pushed the improve/ConcurrentStack-TryPopRange branch from ff05973 to 31ffa76 Compare September 26, 2026 13:43
@aw0lid

aw0lid commented Sep 26, 2026 •

Copy link
Copy Markdown
Contributor Author

Approach

The current implementation walks the requested number of nodes before attempting the CAS. Under contention, the head can change while we are traversing the list, making the work we have already done more likely to be discarded by a failed CAS.

This change periodically checks whether _head is still the same as the head we started with. If it has changed, we stop the traversal early and retry with the latest head instead of continuing to walk the remaining nodes.

The frequency of these checks is scaled based on count, so smaller pops perform fewer checks while large pops check often enough to avoid spending a significant amount of time traversing a stale chain.

int shift = Numerics.BitOperations.Log2(int.MaxValue / (uint)count);
int checkHeadMask = (1 << shift) - 1;

Benchmarks Results

TryPopCount = 3

ParallelThreads Before (Mean) After (Mean) Delta / Change
0 64.49 ns 66.98 ns +2.49 ns (Regression)
1 383.14 ns 266.28 ns -116.86 ns (Faster)
2 901.59 ns 1,456.63 ns +555.04 ns (Regression)
5 5,388.87 ns 6,208.12 ns +819.25 ns (Regression)
8 9,768.07 ns 9,195.54 ns -572.53 ns (Faster)

TryPopCount = 10

ParallelThreads Before (Mean) After (Mean) Delta / Change
0 146.40 ns 158.43 ns +12.03 ns (Regression)
1 2,627.07 ns 5,017.36 ns +2,390.29 ns (Regression)
2 2,900.49 ns 3,731.73 ns +831.24 ns (Regression)
5 7,703.26 ns 6,651.47 ns -1,051.79 ns (Faster)
8 11,581.57 ns 10,059.38 ns -1,522.19 ns (Faster)

TryPopCount = 100

ParallelThreads Before (Mean) After (Mean) Delta / Change
0 1,380.87 ns 1,430.25 ns +49.38 ns (Regression)
1 7,682.80 ns 8,293.46 ns +610.66 ns (Regression)
2 8,378.94 ns 8,322.32 ns -56.62 ns (Faster)
5 14,525.19 ns 16,004.43 ns +1,479.24 ns (Regression)
8 27,640.60 ns 29,385.75 ns +1,745.15 ns (Regression)

TryPopCount = 1,000

ParallelThreads Before (Mean) After (Mean) Delta / Change
0 16,241.43 ns 16,081.90 ns -159.53 ns (Faster)
1 32,005.97 ns 32,029.28 ns +23.31 ns (Noise)
2 40,362.78 ns 40,113.60 ns -249.18 ns (Faster)
5 51,566.82 ns 53,107.34 ns +1,540.52 ns (Regression)
8 55,901.09 ns 56,165.32 ns +264.23 ns

TryPopCount = 10,000

ParallelThreads Before (Mean) After (Mean) Delta / Change
0 175.11 μs 175.29 μs +0.18 μs
1 293.19 μs 288.11 μs -5.08 μs (Faster)
2 546.36 μs 497.67 μs -48.69 μs (Faster)
5 1,254.12 μs 1,094.44 μs -159.68 μs (Faster)
8 1,369.17 μs 1,167.11 μs -202.06 μs (Faster)

TryPopCount = 100,000

ParallelThreads Before (Mean) After (Mean) Delta / Change
0 2.67 ms 2.48 ms -0.19 ms (Faster)
1 5.98 ms 7.87 ms +1.89 ms (Regression)
2 11.99 ms 12.04 ms +0.05 ms
5 128.01 ms 14.05 ms -113.96 ms (Faster)
8 232.45 ms 16.84 ms -215.61 ms (Faster)

TryPopCount = 1,000,000

ParallelThreads Before (Mean) After (Mean) Delta / Change
0 97.08 ms 89.29 ms -7.79 ms (Faster)
1 136.92 ms 103.73 ms -33.19 ms (Faster)
2 251.16 ms 104.10 ms -147.06 ms (Faster)
5 24,535.96 ms 108.20 ms -24.43 s (Faster)
8 38,800.00 ms 110.30 ms -38.69 s (Faster)

Note

The results show a clear trade-off: the additional _head checks introduce some overhead for small pops and low contention, but significantly reduce wasted traversal under high contention and large pop counts.

Is this trade-off acceptable?

Open Question

I also noticed that skipping the CAS when _head changes during traversal and immediately entering the backoff path causes a severe performance regression under contention—with execution times spiking from milliseconds up to several seconds.

I'm not sure why this happens. Is there a reason we should still attempt the CAS in this case?

@jkotas Since you previously looked at the earlier validation approach, I'd also appreciate your thoughts on the new approach and the open question above.

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟡 Changes recommended

The early-abort path performs a guaranteed-failing CAS, while ordinary TryPop gains unnecessary arithmetic overhead.

Review effort: Balanced
Findings: 1 Medium severity · 1 Low severity

Open (2)
What changed in this PR

Optimizes ConcurrentStack<T>.TryPopRange by detecting contention during long range traversals.

Changes:

  • Adds adaptive head-check intervals.
  • Aborts traversal when the stack head changes.
File Description
ConcurrentStack.cs Adds contention detection to range popping.

Comment on lines +610 to +614
if ((nodesCount & checkHeadMask) == 0)
{
if (head != _head)
{
break;

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Valid feedback

@aw0lid aw0lid Sep 27, 2026 •

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I did some additional experiments around skipping the CAS after detecting that _head has changed. The results are mixed for smaller ranges:

TryPopCount Threads Baseline CAS kept CAS skipped
3 16 24.08 μs 24.46 μs 25.29 μs
10 16 22.67 μs 20.64 μs 21.35 μs
100 16 43.86 μs 45.23 μs 42.65 μs
1,000 16 69.38 μs 68.68 μs 66.66 μs

For 100,000+, skipping the CAS caused a substantial regression under contention, reaching seconds per iteration. These cases also take significantly longer to benchmark.

@tannergooding, I'd be interested in your thoughts on this, and whether there is a way to avoid the CAS instruction in this path.

@jkotas

jkotas commented Sep 27, 2026

Copy link
Copy Markdown
Member

I also noticed that skipping the CAS when _head changes during traversal and immediately entering the backoff path causes a severe performance regression under contention—with execution times spiking from milliseconds up to several seconds.

This is not unusual behavior for spin locks under contention when they are protecting an operation that is non-trivial.

My guess is that this behavior would vary across machines or with small modifications of the microbenchmark.

@@ -586,6 +586,8 @@ private int TryPopCore(int count, out Node? poppedHead)
Node? head;
Node next;
int backoff = 1;
int shift = Numerics.BitOperations.Log2(int.MaxValue / (uint)count);
int checkHeadMask = (1 << shift) - 1;

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Is there some rationale behind this formula?

If I am reading this correctly, these additional checks are only going to kick for count >30k. For example, when count = 10_000, checkHeadMask is going to be 131071 and so we won't execute any additional checks. So it is surprising that you are able to measure an improvement for 10_000.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

You're right — for count = 10,000, the formula does not actually introduce any additional _head checks.

The rationale was to make the check interval adaptive to count— checking _head more frequently as range size grows to prevent heavy traversal on a stale chain, while keeping checks infrequent for smaller counts to minimize overhead. I intentionally used a heuristic rather than hardcoded thresholds, derived purely from experimentation.

Regarding the improvements observed at smaller counts (like 10,000), the repeated runs don't show a consistent benefit, so I don't consider them meaningful evidence for the heuristic:

TryPopCount = 10,000 (After):

ParallelThreads Run 1 Run 2 Run 3
0 179.8 μs 198.5 μs 176.7 μs
1 287.0 μs 333.0 μs 278.9 μs
2 620.3 μs 590.9 μs 563.7 μs
5 1,069.2 μs 1,189.5 μs 1,203.9 μs
8 666.9 μs 1,367.8 μs 1,431.1 μs
16 771.1 μs 1,446.7 μs 1,496.8 μs

Baseline (Before):

ParallelThreads Run 1 Run 2 Run 3
0 176.2 μs 175.5 μs 175.8 μs
1 283.0 μs 285.4 μs 282.6 μs
2 664.8 μs 690.0 μs 584.8 μs
5 1,181.9 μs 1,157.8 μs 1,185.0 μs
8 1,438.0 μs 1,215.8 μs 1,193.2 μs
16 1,539.0 μs 1,946.6 μs 1,561.5 μs

I don't consider small-count improvements meaningful evidence for the heuristic; the main value of this PR lies in large ranges under high contention, where stopping stale traversals can prevent severe performance degradation.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

area-System.Collections community-contribution Indicates that the PR has been added by a community member

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants