Skip to content

Reduce weak ref blocking with java interop - #131952

Merged
BrzVlad merged 4 commits into
dotnet:mainfrom
BrzVlad:feature-no-bridge-wait
Sep 28, 2026
Merged

BrzVlad merged 4 commits into
dotnet:mainfrom
BrzVlad:feature-no-bridge-wait

Conversation

@BrzVlad

@BrzVlad BrzVlad commented Aug 6, 2026 •

Copy link
Copy Markdown
Member

Bridge objects live in 2 worlds, .net and java so a .net bridge object has a correpsonding java peer. Collection of these objects is triggered by .net. When the .net peer is eligible for collection we build some graph over the set of dead objects and pass it over to java. Java triggers its own collection, collecting the java peers if they are dead as well. .NET android reports which bridge objects died on the java side so we can drop the gchandles for them. This will finally allow .net peers to die in the following collection (since they had to be promoted, given we don't know yet if java peers need to keep them alive or not).

Currently obtaining the target of a weak ref blocks until the bridge processing is fully completed. This is the case also on mono and prevents 2 issues:

  • normal c# code checks a weak ref for some object. This can't immediately return correct information. If it returns true, the object gets resurrected and we can end up with a ref to a bridge objects that no longer has a java peer. If it returns false then that can be false as well if the object remains alive.
  • java code could call into managed, inserting a reference to a C# peer and afterward it could drop its own java peer. If this happens while C# gc ran but the java gc is still yet to start, both GC would see their peer as dead, even though it is alive. The .NET android interop obtains the C# peer ref also via weak reference, so this safely synchronizes with bridge processing.

This PR keeps the weak reference wait only for bridge objects that are currently processed. For a weak ref target we need to determine whether the underlying object is pending bridge processing which is awkward to do efficiently because we would need to iterate over a set of handles or implement a lookup from obj address to associated cross reference handle. It turns out there is a free bit in the object header that we could use for this purpose.

Addresses #131370

Copilot AI lite review requested due to automatic review settings August 6, 2026 16:10
@azure-pipelines

Copy link
Copy Markdown
Azure Pipelines:
Successfully started running 3 pipeline(s).
13 pipeline(s) were filtered out due to trigger conditions.
There may be pipelines that require an authorized user to comment /azp run to run.

@BrzVlad

BrzVlad commented Aug 6, 2026

Copy link
Copy Markdown
Member Author

This still needs some polishing, but I'm looking for feedback whether it would be feasible to borrow a per object bit from somewhere. Seems doable from the object header, but according to the comment there might be some friction with debug builds. cc @jkotas

@dotnet-policy-service

Copy link
Copy Markdown
Contributor

Tagging subscribers to this area: @anicka-net, @dotnet/gc
See info in area-owners.md if you want to be subscribed.

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

This PR reduces blocking when resolving weak-reference targets during Android Java GC bridge processing by marking bridge objects that are pending client processing using a previously-unused object header bit. Weak-reference resolution only waits (returns “need to wait”) for objects currently marked as bridge-pending, instead of for all weak handles while bridge processing is active.

Changes:

  • Repurposes the high syncblock/header bit as BIT_SBLK_BRIDGE_PENDING and uses it to mark bridge objects awaiting Java-side processing.
  • Tracks “pending bridge” handle cells during bridge graph construction and clears the pending bit when bridge processing completes (or is not triggered).
  • Updates weak-handle fast-path (GCHandleInternalTryGetBridgeWait) to consult bridge-pending state rather than unconditionally blocking when the bridge is active.

Reviewed changes

Copilot reviewed 9 out of 9 changed files in this pull request and generated 2 comments.

Show a summary per file
File Description
src/coreclr/vm/syncblk.h Renames/reassigns the previously-unused header bit to BIT_SBLK_BRIDGE_PENDING.
src/coreclr/vm/marshalnative.cpp Routes weak-handle “try get” through the new bridge-pending-aware helper.
src/coreclr/vm/interoplibinterface.h Adds Interop::TryGetObjectFromHandleWithoutBridgeWait declaration (FEATURE_JAVAMARSHAL).
src/coreclr/vm/interoplibinterface_java.cpp Implements pending-bit check/clear for CoreCLR Java bridge flow.
src/coreclr/nativeaot/Runtime/ObjectLayout.h Adds BIT_SBLK_BRIDGE_PENDING for NativeAOT parity.
src/coreclr/nativeaot/Runtime/interoplibinterface_java.cpp Implements pending-bit check/clear for NativeAOT Java bridge flow.
src/coreclr/gc/objecthandle.cpp Records pending handle cells and notes the active-bridge recomputation race (FIXME).
src/coreclr/gc/gcbridge.h Exposes pending-handle tracking APIs.
src/coreclr/gc/gcbridge.cpp Implements pending-handle tracking and sets BIT_SBLK_BRIDGE_PENDING on candidates.
Suppressed comments (2)

src/coreclr/vm/interoplibinterface_java.cpp:127

  • When g_GCBridgeActive is already true, this early-return path releases args but never clears BIT_SBLK_BRIDGE_PENDING on the objects just marked in ProcessBridgeObjects for this GC. That can leave stale pending bits behind, causing future weak-handle checks to spuriously block (especially on subsequent bridge-active cycles). Clear the pending bits for the current pending handle set before returning.
    if (g_GCBridgeActive)
    {
        // FIXME: This should become unreachable once bridge graph recomputation is skipped while active.
        // Release the memory allocated since the GCBridge
        // is already running and we're not passing them to it.
        ReleaseGCBridgeArgumentsWorker(args);
        return;

src/coreclr/nativeaot/Runtime/interoplibinterface_java.cpp:77

  • When g_GCBridgeActive is already true, this early-return path releases args but never clears BIT_SBLK_BRIDGE_PENDING on the objects just marked in ProcessBridgeObjects for this GC. That can leave stale pending bits behind, causing future weak-handle checks to spuriously block. Clear the pending bits for the current pending handle set before returning.
    if (g_GCBridgeActive)
    {
        // FIXME: This should become unreachable once bridge graph recomputation is skipped while active.
        // Release the memory allocated since the GCBridge
        // is already running and we're not passing them to it.
        ReleaseGCBridgeArgumentsWorker(args);
        return;
    }

Comment thread src/coreclr/vm/interoplibinterface_java.cpp
Comment thread src/coreclr/gc/objecthandle.cpp Outdated
@jkotas

jkotas commented Aug 6, 2026

Copy link
Copy Markdown
Member

Addresses #131370

#131370 analysis says "Ruled out: WeakReference.Target bridge blocking is not a contributor"

Does this make a difference for real apps?

@BrzVlad

BrzVlad commented Aug 7, 2026 •

Copy link
Copy Markdown
Member Author

Does this make a difference for real apps?

I can't comment on copilot's conclusion on its own created benchmark (which also doesn't reflect gc behavior for user app, where GC is triggered once every couple of seconds, both in their repro and in the actual game where Rolf tested - https://gist.github.com/rolfbjarne/5fd7ac636132718170ea30dc1be2452c. The copilot comment reports collection rates of 40 per second, which is absurd in my opinion, especially for a game). I tested on the exact sample provided by the customer (https://github.com/hyvanmielenpelit/Net11FPSBenchmark) where, following each gc, there were waits on the UI thread of ~30ms. The sample extracted logic from their actual game. While doing other investigation, starting from the maui sample, I simply triggered GCs directly from the button callback which resulted in waits as well. Aside from users simply using weak reference that they shouldn't expect to have significant overhead, the C#-Java interop layer is filled with uses of weak reference checks for bridge objects (so any native event that needs to bubble up into C# would be blocking unnecessarily on the java gc finish and there are probably dozens of other scenarios).

I'm also putting into perpsective that the change is rather simplistic and is a nice to have optimization regardless.

Comment thread src/coreclr/vm/syncblk.h Outdated
@jkotas

jkotas commented Aug 8, 2026

Copy link
Copy Markdown
Member

it would be feasible to borrow a per object bit

I think it is ok. It should be under FEATURE_JAVAMARSHAL so that it can be used for other purposes on other OSes.

@steveisok

Copy link
Copy Markdown
Member

Correction — that "ruled out" line is mine and it's wrong. I've edited the original comment on #131370.

The experiment behind it compared a weak-reference arm against a strong-reference control and found
them indistinguishable. That wasn't a control: dotnet/android resolves peers through weak references
on the UI thread irrespective of what my probe reads, so both arms already contained the mechanism I
was trying to isolate.

Symbolized off-CPU profiling says the opposite. GCHandle_InternalGetBridgeWait →
Interop::WaitForGCBridgeFinish is 31.34% of UI-thread off-CPU time with peers, and absent
without them. The probe's own read loop was off in those runs (reads=0), so every one of those
waits came from the binding layer's peer bookkeeping — not from app code touching
WeakReference.Target. That's what makes me think it generalizes past my probe: any Android app with
peers hits this path whether or not it uses weak references itself.

Ablating it confirms the size. A prototype of this same fix moves the probe 42.7 → 49.3 fps and
drops that frame from 30.21% to 0.01% of UI-thread off-CPU time, recovering ~41% of the lost frame
time.

Two caveats:

  • It doesn't close the gap on its own. No bridge at all is 59.2 fps, so the remaining ~59% is the
    force-promotion of dead bridge objects out of gen0 — a separate mechanism this PR doesn't touch.
  • These are emulator numbers at a deliberately high peer rate (1200 peers/frame). I have not
    measured a patched runtime against GnollHack itself, so I can't put a number on the real app.

Details and the raw runs: https://github.com/steveisok/android-gcbridge-investigation

Note

This analysis was generated with GitHub Copilot.

@steveisok

Copy link
Copy Markdown
Member

This fix will help, but the majority of the problem will remain in the cost of the promotion of these objects.

@steveisok
steveisok self-requested a review August 10, 2026 18:23
Copilot AI review requested due to automatic review settings August 11, 2026 10:26
@BrzVlad
BrzVlad force-pushed the feature-no-bridge-wait branch from 7b98ec3 to 00a176a Compare August 11, 2026 10:26

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Copilot reviewed 10 out of 10 changed files in this pull request and generated no new comments.

Suppressed comments (3)

src/coreclr/vm/interoplibinterface_java.cpp:124

  • TriggerClientBridgeProcessing relies on _ASSERTE(!g_GCBridgeActive) but has no retail guard. If this invariant is violated in a non-Debug build (e.g., redundant/overlapping bridge triggering), the code will proceed and can corrupt the pending-bridge bookkeeping and/or double-trigger client processing. The previous implementation handled this safely by releasing args and returning.
    size_t pendingBridgeHandleCount;
    uintptr_t* pendingBridgeHandles = GetPendingBridgeHandles(&pendingBridgeHandleCount);

    _ASSERTE(!g_GCBridgeActive);

    bool gcBridgeTriggered = JavaNative::TriggerClientBridgeProcessing(args);

src/coreclr/nativeaot/Runtime/interoplibinterface_java.cpp:74

  • TriggerClientBridgeProcessing uses _ASSERTE(!g_GCBridgeActive) as the only protection against overlapping bridge triggers. In retail builds this becomes a no-op; if the invariant is ever violated, the function will proceed and can corrupt pending-handle state or trigger client bridge processing twice. A defensive runtime check (as existed before) would make this robust.
    size_t pendingBridgeHandleCount;
    uintptr_t* pendingBridgeHandles = GetPendingBridgeHandles(&pendingBridgeHandleCount);

    _ASSERTE(!g_GCBridgeActive);

src/coreclr/gc/objecthandle.cpp:1533

  • The comment above the HndScanHandlesForGC call no longer matches the new lp2 usage (it is now a boolean indicating whether to record pending bridge handles, not a pointer for promotion data). This is likely to mislead future changes to the scanning callback contract.
                        // or have a local var for bridgeObjectsToPromote/size (instead of NULL) that's passed in as lp2

Copilot AI review requested due to automatic review settings August 11, 2026 10:48

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Copilot reviewed 13 out of 13 changed files in this pull request and generated no new comments.

Suppressed comments (2)

src/coreclr/vm/marshalnative.cpp:402

  • The comment above this fast path is now misleading: this method no longer waits for bridge processing to finish, it only returns false for handles whose target is currently pending bridge processing. Please update the comment (and fix the typo) to match the new behavior, otherwise future readers may assume this blocks or fully synchronizes with the bridge.
FCIMPL2(FC_BOOL_RET, MarshalNative::GCHandleInternalTryGetBridgeWait, OBJECTHANDLE handle, Object **pObjResult)
{
    FCALL_CONTRACT;

    if (!Interop::TryGetObjectFromHandleWithoutBridgeWait(handle, pObjResult))

src/coreclr/vm/interoplibinterface_java.cpp:127

  • Minor grammar: “wasn't trigger” should be “wasn't triggered”.
    if (!gcBridgeTriggered)
    {
        // Release the memory allocated since the GCBridge
        // wasn't trigger for some reason.
        ClearPendingBridgeBits(pendingBridgeHandles, pendingBridgeHandleCount);

@BrzVlad
BrzVlad force-pushed the feature-no-bridge-wait branch from 2b5a92a to fa2867f Compare August 11, 2026 11:04
Copilot AI review requested due to automatic review settings August 11, 2026 11:04

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Copilot reviewed 13 out of 13 changed files in this pull request and generated no new comments.

Suppressed comments (1)

src/coreclr/vm/interoplibinterface_java.cpp:82

  • TryGetObjectFromHandleWithoutBridgeWait declares result as an _Out_ parameter but returns false without writing *result. Even though current callers likely ignore the value on false, this violates the contract and can lead to accidental use of an uninitialized out value by future callers.
    Object* object = OBJECTREFToObject(ObjectFromHandle(handle));
    if (g_GCBridgeActive && object != nullptr &&
        (object->GetHeader()->GetBits() & BIT_SBLK_BRIDGE_PENDING) != 0)
    {
        return false;
    }

@BrzVlad
BrzVlad marked this pull request as ready for review August 11, 2026 11:48
@azure-pipelines

Copy link
Copy Markdown
Azure Pipelines:
Successfully started running 3 pipeline(s).
13 pipeline(s) were filtered out due to trigger conditions.
There may be pipelines that require an authorized user to comment /azp run to run.

@BrzVlad

BrzVlad commented Aug 12, 2026

Copy link
Copy Markdown
Member Author

This addresses an issue reported on .net11 preview, as a regression from .net10 mono. Ideally we would get this in by Friday, in time for RC1 snap.

Comment thread src/coreclr/vm/interoplibinterface_java.cpp
Comment thread src/coreclr/gc/objecthandle.cpp
Copilot AI review requested due to automatic review settings August 12, 2026 16:55

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Copilot reviewed 13 out of 13 changed files in this pull request and generated no new comments.

Suppressed comments (1)

src/coreclr/gc/gcbridge.cpp:1319

  • ProcessBridgeObjects sets BIT_SBLK_BRIDGE_PENDING on all registered bridge objects even if BuildSccCallbackData() fails to allocate and returns nullptr. In that case, the caller skips triggering client bridge processing, so the pending bits may remain set until the objects die, and a later bridge cycle could spuriously treat them as “pending” and force weak-ref waits.

Consider only setting the pending bit when args != NULL (or clearing the bits on the failure path).

    for (int i = 0; i < DynPtrArraySize(&g_registeredBridges); i++)
    {
        Object* object = (Object*)DynPtrArrayGet(&g_registeredBridges, i);
        object->GetHeader()->SetBit(BIT_SBLK_BRIDGE_PENDING);
    }

@agocke
agocke requested a review from jkoritzinsky August 13, 2026 21:11
steveisok added a commit to steveisok/android-gcbridge-investigation that referenced this pull request Aug 14, 2026
Adds the A/B harness and results for dotnet/runtime#131952, the upstream
implementation of the precise weak-reference wait prototyped here.

The two arms differ only by that PR, which branched directly off the merge
commit of #131764, and share one System.Private.CoreLib.dll because the PR
changes no managed code.

It reproduces the prototype (42.8 -> 49.9 fps at 1200 peers/frame) but
recovers only ~31% of lost frame time at nodes=300, the operating point that
matches the shape reported in #131370 -- so the wait accounts for less of the
damage as the peer rate falls toward the realistic regime, not more.

Also records the measurement trap this cost: an incremental app build does not
re-copy a changed libcoreclr.so out of the runtime pack, so both arms silently
run identical binaries and the result looks like a change that did nothing.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Copilot-Session: 91dd8b98-73db-41b0-8d8c-9e1c21b53f5f
@BrzVlad

BrzVlad commented Aug 17, 2026

Copy link
Copy Markdown
Member Author

Any chance for a review to get this into rc1 ?

@simonrozsival

Copy link
Copy Markdown
Member

MAUI startup A/B on Samsung A16

I tested this change against the exact runtime source used by the installed
Android runtime pack.

Build

  • Device: Samsung Galaxy A16 (SM-A165F), Android 16
  • App: dotnet new maui --sample-content
  • App configuration: Release, android-arm64, CoreCLR, trimmable type map
  • Installed runtime: 11.0.0-rc.2.26455.110
  • VMR commit: 52ecb082fd3889636b5793bcaa9f4ca7cb9deb71
  • Corresponding dotnet/runtime commit:
    459f6b60db0a6fd1ed05aedd4ab4c669b9b5bada
  • Patched runtime: the four commits from Reduce weak ref blocking with java interop #131952 applied to that exact commit
  • Runtime build command for both variants:
    ./build.sh clr.runtime -os android -arch arm64 -c Release -rebuild
  • Base libcoreclr.so:
    77eb5c08ef0992bd5f85a4d2b6172b1c427b43f5b3d44e8d707a1c00621b6572
  • Patched libcoreclr.so:
    ffaa7f11beef1e006e48e555aef2e4487d03f1cb14f6dfa313080982960364b0

The two APKs were cloned from one base APK and re-signed after replacing
lib/arm64-v8a/libcoreclr.so. Excluding signatures, the only differing APK
entry was libcoreclr.so
.

Measurement

  • ART compilation: cmd package compile -m speed -f
  • Cold launch: am start -S -W
  • Metric: TotalTime
  • App data cleared after each install
  • Three warmup cold launches per block
  • Twelve measured cold launches per block
  • Two counterbalanced passes, eight install blocks each
  • 96 launches per variant
  • Device temperature during measured runs: 28.1-28.5 C

Results

Variant Mean Median StdDev
Base 2,598.23 ms 2,594.0 ms 42.42 ms
Selective weak wait 2,571.54 ms 2,569.5 ms 33.74 ms
Difference -26.69 ms (-1.03%) -24.5 ms
  • Bootstrap 95% CI for the mean difference: -37.71 to -16.09 ms
  • Two-sided permutation p-value: < 0.00001
  • 10% trimmed-mean difference: -24.15 ms
  • Pass 1: -33.12 ms (-1.27%)
  • Reverse-order pass 2: -20.25 ms (-0.78%)

Bridge confirmation

A separate diagnostic startup with GC logging enabled recorded:

  • 386 bridge SCCs
  • 23 cross-references
  • callback at 14:44:29.887
  • cleanup completion at 14:44:29.929

That is an approximately 42 ms accepted bridge round during startup. The
observed ~27 ms first-display improvement is consistent with removing
unnecessary UI-thread weak-reference waits during part of that round; this PR
does not make the bridge itself complete faster.

Conclusion

On this peer-heavy MAUI sample, selective weak-reference waiting produces a
repeatable, statistically detectable improvement of about 20-30 ms, or
roughly 1% of cold startup time.

These local CoreCLR builds do not have the official runtime pack's PGO/BOLT
optimization, so the absolute startup values should not be compared with
shipping builds. Both sides use identical local build settings, making the
relative A/B result the meaningful value.

Bridge objects live in 2 worlds, .net and java so a .net bridge object has a correpsonding java peer. Collection of these objects is triggered by .net. When the .net peer is eligible for collection we build some graph over the set of dead objects and pass it over to java. Java triggers its own collection, collecting the java peers if they are dead as well. .NET android reports which bridge objects died on the java side so we can drop the gchandles for them. This will finally allow .net peers to die in the following collection (since they had to be promoted, given we don't know yet if java peers need to keep them alive or not).

Currently obtaining the target of a weak ref blocks until the bridge processing is fully completed. This is the case also on mono and prevents 2 issues:
- normal c# code checks a weak ref for some object. This can't immediately return correct information. If it returns true, the object gets resurrected and we can end up with a ref to a bridge objects that no longer has a java peer. If it returns false then that can be false as well if the object remains alive.
- java code could call into managed, inserting a reference to a C# peer and afterward it could drop its own java peer. If this happens while C# gc ran but the java gc is still yet to start, both GC would see their peer as dead, even though it is alive. The .NET android interop obtains the C# peer ref also via weak reference, so this safely synchronizes with bridge processing.

This PR keeps the weak reference wait only for bridge objects that are currently processed. For a weak ref target we need to determine whether the underlying object is pending bridge processing which is awkward to do efficiently because we would need to iterate over a set of handles or implement a lookup from obj address to associated cross reference handle. It turns out there is a free bit in the object header that we could use for this purpose.

FIXME this has a race with redudndant bridge processing, because a new collection would dirty our g_registeredBridgeHandles.
… and clearing of pending bit

The bit clearing already had memory ordering since it was done via InterlockedAnd. For the read we add GetBitsAcquire which does an acquire load, preserving the ordering on the reader side.
@BrzVlad
BrzVlad force-pushed the feature-no-bridge-wait branch from 1db650f to 68531b9 Compare September 24, 2026 15:05
Copilot AI review requested due to automatic review settings September 24, 2026 15:05

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟡 Changes recommended

Unresolved moderate issues remain in pending-bit cleanup and GC interface versioning, with requested regression coverage and documentation updates.

Get a fresh assessment by requesting another Copilot review.

Review effort: Lite
Findings: 3 Low severity

Open (3)

Comment thread src/coreclr/vm/interoplibinterface_java.cpp
Comment thread src/coreclr/vm/syncblk.h
@radimitrov

Copy link
Copy Markdown

This is a critical hotfix for Android, and I'd like to ask for it to be backported to release/11.0 rather than waiting for .NET 12.

I ship a MAUI app on Android (CoreCLR, .NET 10), and this is my number 1 source of ANRs in production since upgrading to CoreCLR about a month ago. Sentry ANR reports consistently show the main thread in pthread_cond_wait under GCHandle_InternalGetBridgeWait while the bridge thread is in java.lang.Runtime.gc(); one session logged 895 ART collections. I found and fixed a peer leak of my own (window insets re-dispatched on every scroll step), which made it worse, but even after that fix my Play user-perceived ANR rate is still around 2.5–4.5% against the 0.47% bad-behaviour threshold, and may drop further once all users have updated. My only mitigation is DOTNET_GCgen0size=32 MB (17 → 2 bridge rounds in a 45 s launch), and it makes each remaining stall longer whenever peers churn.

As @steveisok's profiling above showed, the waits come from the interop layer's own peer bookkeeping, so any Android app with Java peers hits this, whether or not it touches WeakReference itself. The more Java peers an app creates while the UI is active, the worse it gets. Scrolling a CollectionView or ListView does exactly that (see also dotnet/android#11567, explicit Java GCs during scrolling), so most apps will feel it. And the original report is from a game, with ~30 ms UI-thread waits after every GC.

With CoreCLR becoming the default Android runtime in .NET 11, every app that upgrades from Mono will silently pick up this regression: jank and ANRs that weren't there on Mono, with nothing in their own code to explain them. Without a backport they'll live with it until .NET 12 next November. The fix was aimed at the RC1 snap and appears to have missed it only because of review timing.

Ideally this would ship in .NET 11 GA. If that's no longer possible, could it be considered for release/11.0 servicing (11.0.1/11.0.2)?

@BrzVlad

BrzVlad commented Oct 2, 2026

Copy link
Copy Markdown
Member Author

@radimitrov I'm somewhat skeptical about your conclusion for 2 reasons:

  • mono had the same blocking so this PR is an optimization that mono was missing
  • ANRs should be triggered after the app is unresponsive for seconds. I don't see a GC taking more than a few hundred milliseconds, so it is unclear how these GC waits would lead to an ANR.

My worry is that the main reason for what you are experiencing is something else and backporting this would not necessarily fix your issue. I would be interested in doing some GC profiling / comparison with Mono if you are open to share a repro. (Showcasing significantly worse perf compared with Mono is still valuable, ANR crash is not mandatory). Note you can reach me privately on vlbrez@microsoft.com.

@radimitrov

Copy link
Copy Markdown

@BrzVlad In my case it was an error on my part causing thousands of redundant Java peers every time a CollectionView was scrolled. The bridge GC of these objects was causing a UI freeze. Fixing the thousands of reduntant Java peers caused from the edge to edge code seems to have mitigated a lot of the issue, though I can't be sure by how much yet, until some more data is gathered over the next week at least. I think the ANR itself was almost exclusevely on lower end devices after too much background thread work piled up. It is a heavy IDE app, no way around most of the memory pressure. So yes, even I was at best 80% certain of my conclusion when I posted this, but I thought it is better to do it now rather than potentially miss a release window SR1 and SR2.

I still think this fix making it into .NET 11 is a good idea, even if it does nothing to help in my case. I have added additional diagnostics to determine any other ANR and memory issues.

I will try to put together an isolated reproduction to confirm.

@radimitrov

Copy link
Copy Markdown

@BrzVlad Thanks, you were right to push back. Here's an isolated repro: https://github.com/radimitrov/android-coreclr-gcbridge-repro, the results are there as an MD file.

It is a plain .NET for Android app (no MAUI) plus one script (run-repro.ps1). The script builds the same code as Mono (.NET 10), CoreCLR (.NET 10, 11 RC1, and the 12.0.0-alpha.1.26480.102 daily, which includes this PR) and CoreCLR with a fixed gen0 size. It runs the same workloads in alternating order and writes one report. On a rooted device it also takes debuggerd dumps during freezes and symbolizes the libcoreclr.so frames.

On your two points

  • Mono blocks too. You were right, and the repro shows Mono's weak reads waiting as well. I withdraw the "critical hotfix" framing.
  • A Java GC doesn't take seconds. Also right: the bridge's Runtime.gc() takes ~50-75 ms here. The longer freezes come from elsewhere (below).

1. Many small rounds (background allocation): this PR removes most of the UI cost

heavy: 32 MB/s background allocation, 20 dead peers per frame, 30 s bridge rounds fps UI time lost
Mono (.NET 10) 244 27.2 17.9 s
CoreCLR .NET 11 RC1 193 33.5 14.6 s
CoreCLR .NET 12 daily (with #131952) 193 59.5 0.6 s

Same GC work (same managed GC counts, same rounds, same Java GC time), about 24× less lost UI time. With light background work, the worst freeze drops from 155 ms to 68 ms. A large win at no GC cost, so I'd still ask for it in .NET 11 servicing.

2. Large rounds: freezes remain, and this PR doesn't remove them

burst: no background work, 60 dead peers per frame, 60 s bridge rounds managed GCs gen0/1/2 worst freeze
Mono (.NET 10) 4 0/4/4 217 ms
CoreCLR .NET 11 RC1 6 6/2/0 734 ms
CoreCLR .NET 12 daily (with #131952) 4 6/6/6 1,650 ms

Rounds here carry roughly 34k (11) to 50k (12) dead peers, against a similar ~50k on Mono. About 90% of each CoreCLR freeze comes after Runtime.gc() returns. The dumps taken mid-freeze (two per runtime) all show the same thing on 11 and 12:

UI thread:     GCHandle_InternalGetBridgeWait → Interop::WaitForGCBridgeFinish → (waiting)
bridge thread: JavaMarshal_FinishCrossReferenceProcessing → Interop::FinishCrossReferenceProcessing
               → Ref_NullBridgeObjectsWeakRefs → HndEnumHandles → NullBridgeObjectWeakRef

NullBridgeObjectWeakRef scans the whole unreachable array for every short/long weak handle in the process (// FIXME Store these objects in a hashtable in order to optimize lookup; the same in release/10.0, release/11.0 and main). Every registered peer holds a WeakReference in dotnet/android's registry, so that's O(weak handles × dead peers). The freeze lengths are consistent with that: with a ~0.12 s per-round baseline, rounds of ~4k / ~34k / ~50k dead peers freeze 0.15 / 0.73 / 1.5-1.7 s. The round sizes are estimates, though, and I haven't measured a runtime with that loop changed.

What I can't explain yet: with #131952 the UI should only wait for objects pending in the round, and in heavy it doesn't wait at all. In burst on 12 it still ends up in GCHandle_InternalGetBridgeWait, so it is reading something that is pending. It isn't the frame callback itself: holding it in a managed static changes nothing. Threads that don't touch Java interop never stall, so it isn't a stop-the-world GC. The one difference I see is that every GC in burst on 12 is a full gen2 collection (gen0 on 11), so long-lived peers are in scope too.

Separately, for .NET 10

CoreCLR uses the 256 KB gen0 floor on Android: release/10.0 defines TARGET_ANDROID instead of TARGET_LINUX, so the cache-size detection in gcenv.unix.cpp is compiled out (0.21 MB allocated per GC in the repro). That gives 20× Mono's GC count and 603 vs 244 bridge rounds under load. #128826 fixes it in .NET 11; for .NET 10 it may deserve servicing, or at least a documented DOTNET_GCgen0size.

None of this reached ANR length (5 s) on this device, but the large-round freezes grow faster than linearly with dead peers, and a low-end phone does the same work several times slower. Happy to run anything else that helps.

In truth .NET 11 RC1 performs better than .NET 10 so the real world performance will probably already be better even without the fix, also unless I am mistaken there is still work actively being done on the GC area for Android. So I will leave it entirely to your judgement if this PR should make it into a .NET 11 release.

Results

scenario variant managed GCs bridge rounds Java GC ms fps UI time lost ms worst freeze ms freezes >100 ms >700 ms weak-read wait ms
idle Mono 0 0 0 60.0 15 32 0 0 0
idle CoreCLR 18 18 1,324 56.8 1,526 150 11 0 0
idle CoreCLR-gen0-1000000 0 0 0 60.3 10 27 0 0 0
idle CoreCLR-gen0-2000000 1 1 62 58.7 786 777 1 1 0
idle CoreCLR-net11 1 1 68 59.1 543 534 1 0 0
idle CoreCLR-net12 0 0 0 60.3 16 32 0 0 1
idle CoreCLR-net12-gen0-2000000 0 0 0 60.3 16 28 0 0 0
light Mono 31 31 2,448 54.5 2,881 131 21 0 6
light CoreCLR 29 29 2,183 55.3 2,473 141 17 0 1
light CoreCLR-gen0-1000000 7 7 499 58.5 891 183 7 0 0
light CoreCLR-gen0-2000000 3 3 192 59.3 527 189 3 0 0
light CoreCLR-net11 21 21 1,600 56.3 1,965 155 14 0 1
light CoreCLR-net12 27 27 2,028 59.7 375 68 0 0 17
light CoreCLR-net12-gen0-2000000 3 3 188 59.7 302 134 1 0 1
heavy Mono 243 244 18,264 27.2 17,922 130 81 0 168
heavy CoreCLR 4,919 603 26,486 22.2 19,984 136 8 0 348
heavy CoreCLR-gen0-1000000 61 60 4,231 50.2 5,479 140 40 0 0
heavy CoreCLR-gen0-2000000 31 30 1,893 54.8 2,927 142 21 0 0
heavy CoreCLR-net11 194 193 14,268 33.5 14,612 130 90 0 102
heavy CoreCLR-net12 194 193 14,197 59.5 645 66 0 0 1
heavy CoreCLR-net12-gen0-2000000 30 30 1,907 58.6 1,032 75 0 0 5
heavy-nopeers Mono 243 2 159 59.7 305 109 1 0 4
heavy-nopeers CoreCLR 4,823 4 229 59.7 179 78 0 0 35
heavy-nopeers CoreCLR-gen0-1000000 61 3 170 58.7 1,154 107 1 0 1
heavy-nopeers CoreCLR-gen0-2000000 31 3 157 58.7 965 92 0 0 0
heavy-nopeers CoreCLR-net11 194 3 167 59.6 736 72 0 0 2
heavy-nopeers CoreCLR-net12 194 3 164 59.5 859 68 0 0 11
heavy-nopeers CoreCLR-net12-gen0-2000000 30 2 108 58.6 985 74 0 0 1
burst Mono 0 4 176 59.2 768 217 4 0 134
burst CoreCLR 48 48 3,710 54.7 5,289 151 43 0 4
burst CoreCLR-gen0-1000000 4 4 201 54.3 5,785 1,700 4 4 14
burst CoreCLR-gen0-2000000 4 4 191 54.7 5,430 1,716 4 4 12
burst CoreCLR-net11 6 6 359 56.2 3,958 734 6 3 25
burst CoreCLR-net12 6 4 206 54.2 5,945 1,650 4 4 29
burst CoreCLR-net12-gen0-2000000 6 4 209 54.7 5,498 1,534 4 4 6
burst-bgpeers Mono 0 4 173 59.3 719 205 4 0 1
burst-bgpeers CoreCLR 48 48 3,731 55.2 5,135 164 40 0 3
burst-bgpeers CoreCLR-gen0-1000000 4 4 211 54.3 5,826 1,720 4 4 1
burst-bgpeers CoreCLR-gen0-2000000 4 4 201 55.0 5,248 1,664 4 4 1
burst-bgpeers CoreCLR-net11 6 6 336 56.4 3,837 747 6 1 0
burst-bgpeers CoreCLR-net12 7 5 254 54.6 5,532 1,438 5 4 1
burst-bgpeers CoreCLR-net12-gen0-2000000 5 4 207 54.4 5,743 1,564 5 4 1

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

8 participants