Conversation
|
Azure Pipelines: Successfully started running 3 pipeline(s). 13 pipeline(s) were filtered out due to trigger conditions. There may be pipelines that require an authorized user to comment /azp run to run. |
|
Tagging subscribers to this area: @JulieLeeMSFT, @jakobbotsch |
|
Tagging subscribers to this area: @dotnet/area-system-collections |
|
Azure Pipelines: Successfully started running 3 pipeline(s). 13 pipeline(s) were filtered out due to trigger conditions. There may be pipelines that require an authorized user to comment /azp run to run. |
|
Hi all, just following up on this validation PR when you have a chance. I’d appreciate any feedback on the approach, especially the open question around |
This sounds like a problem. From the API documentation at https://learn.microsoft.com/dotnet/api/system.collections.concurrent.concurrentstack-1.trypoprange : "Attempts to pop and return multiple objects from the top of the ConcurrentStack atomically." |
|
Thanks for confirming, @jkotas. I had already looked at the API documentation before opening the PR, but I wanted to confirm my understanding of the atomicity requirement. I also noticed that the
Would this mean that it is valid for For example: Instead of continuing to perform additional CAS operations until all 1,000,000 items have been popped, the operation would return 500,000 after successfully popping that smaller batch with a single CAS. I also noticed the existing test Assert.Equal(numElementsPerThread, res);So I wanted to make sure I’m not overlooking an intended semantic here. Would returning a smaller batch in this way be considered valid for |
I do not think so. We have many |
|
I have another approach I'd like to experiment with. Instead of traversing the entire range before detecting a changed head, periodically check the head during traversal (e.g. every 1/4 of the range). If it changed, discard the current traversal and restart from the new head; otherwise, keep going. The idea is to reduce wasted traversal under contention while preserving the existing atomicity semantics. I'll try it and benchmark it when I get some time. |
8523a56 to
8b9361f
Compare
|
The last push was made by mistake, which resulted in an incorrect diff and caused this PR to be closed. |
Overview
This PR is opened as a validation PR to investigate potential approaches for improving
ConcurrentStack<T>.TryPopRangeperformance under contention, as described in #100083.The current implementation attempts to pop the requested number of items as a single batch. Under contention, repeated CAS failures can cause the same large portion of the stack to be traversed repeatedly, leading to a significant performance degradation.
Approach 1: Adaptive Batch Sizing
The first approach explored in this PR is adaptive batch sizing.
The batch size starts at the requested count.
If the CAS fails 30 times consecutively, the batch size is reduced by half and the operation is retried. After a successful CAS, the popped batch is copied to the destination array and the remaining items are processed using the current batch size.
If contention increases again and the CAS fails another 30 times, the batch size is reduced again.
For example:
The goal of this approach is to avoid repeatedly traversing a large number of nodes when the stack is highly contended, while still allowing large batches when contention is low.
Initial results
Show benchmark code
The most significant improvement can be seen under high contention, where the benchmark goes from ~58 seconds to ~476 ms with 8 threads, while allocations drop from ~15.8 GB to ~268 MB.
These results are encouraging and show that adaptive batch sizing can significantly reduce the performance degradation under contention.
Further approaches may be explored in this PR to evaluate different strategies and trade-offs.
Open Question
One important aspect that still needs to be evaluated is whether performing multiple successful CAS operations within a single
TryPopRangeinvocation is compatible with the atomicity semantics expected from the API.Further validation and discussion are needed before considering any approach as a final implementation.
Note
This is a validation implementation intended to evaluate the approach and its performance.
The code may be refined or refactored later.