Skip to content

[NativeAOT] Fast interface dispatch helper for riscv64 - #134057

Merged
MichalStrehovsky merged 4 commits into
dotnet:mainfrom
maximmenshikov:maximmenshikov/cached_dispatch
Sep 24, 2026
Merged

MichalStrehovsky merged 4 commits into
dotnet:mainfrom
maximmenshikov:maximmenshikov/cached_dispatch

Conversation

@maximmenshikov

Copy link
Copy Markdown
Contributor

Interface dispatch moved to a helper with a global dispatch cache in #123252. amd64 and arm64 got optimized helpers that check the monomorphic cell inline and probe the cache in assembly; riscv64 was left with the slow path, so every interface call goes through RhpUniversalTransition and the managed resolver.

Port the arm64 INTERFACE_DISPATCH macro: monomorphic cell check, then a GenericCache<Key, nint> probe using the same hash, the same quadratic reprobe and the same seqlock version check, falling back to RhpCidResolve and RhpCidResolve_Worker. Only temporaries are used, so nothing has to be spilled around the probe.

Three things differ from arm64, all forced by the ISA:

  • arm64 gets its acquire loads from ldar and orders the value read against the version re-read with dmb ishld. RVWMO does not order a plain ld against later loads, so each of those becomes a fence r, r after the load.
  • at the key compare only one temporary is still free, so the probe count travels in the upper bits of the index register instead of having a register of its own.
  • the tail call to RhpUniversalTransitionTailCall is written out by hand: the tail pseudo-instruction expands through t1, which is carrying the thunk parameter.

Interface dispatch moved to a helper with a global dispatch cache in
dotnet#123252. amd64 and arm64 got optimized helpers that check the
monomorphic cell inline and probe the cache in assembly; riscv64 was left
with the slow path, so every interface call goes through
RhpUniversalTransition and the managed resolver.

Port the arm64 INTERFACE_DISPATCH macro: monomorphic cell check, then a
GenericCache<Key, nint> probe using the same hash, the same quadratic
reprobe and the same seqlock version check, falling back to RhpCidResolve
and RhpCidResolve_Worker. Only temporaries are used, so nothing has to be
spilled around the probe.

Three things differ from arm64, all forced by the ISA:

  - arm64 gets its acquire loads from ldar and orders the value read
    against the version re-read with dmb ishld. RVWMO does not order a
    plain ld against later loads, so each of those becomes a fence r, r
    after the load.
  - at the key compare only one temporary is still free, so the probe
    count travels in the upper bits of the index register instead of
    having a register of its own.
  - the tail call to RhpUniversalTransitionTailCall is written out by
    hand: the tail pseudo-instruction expands through t1, which is
    carrying the thunk parameter.

Signed-off-by: Maxim Menshikov <maksim.menshikov@nethermind.io>
@dotnet-policy-service dotnet-policy-service Bot added the community-contribution Indicates that the PR has been added by a community member label Sep 16, 2026
@azure-pipelines

Copy link
Copy Markdown
Azure Pipelines:
Successfully started running 3 pipeline(s).
13 pipeline(s) were filtered out due to trigger conditions.
There may be pipelines that require an authorized user to comment /azp run to run.

@dotnet-policy-service

Copy link
Copy Markdown
Contributor

Tagging subscribers to this area: @agocke, @dotnet/ilc-contrib
See info in area-owners.md if you want to be subscribed.

@jkotas jkotas added the arch-riscv Related to the RISC-V architecture label Sep 16, 2026

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🔵 Needs a closer look

The hand-written concurrent assembly requires RISC-V hardware validation of ABI and memory-ordering behavior.

Pull request overview

Adds optimized NativeAOT interface dispatch for RISC-V 64, matching existing ARM64 cache behavior.

Changes:

  • Adds monomorphic dispatch-cell checks.
  • Probes the global dispatch cache with RISC-V memory ordering.
  • Preserves slow-path resolution for cache misses.
File summaries
File Description
src/coreclr/nativeaot/Runtime/riscv64/DispatchResolve.S Implements cached RISC-V 64 interface dispatch.
Review details
  • Files reviewed: 1/1 changed files
  • Comments generated: 0
  • Review effort level: Balanced

@azure-pipelines

Copy link
Copy Markdown
Azure Pipelines:
Successfully started running 3 pipeline(s).
13 pipeline(s) were filtered out due to trigger conditions.
There may be pipelines that require an authorized user to comment /azp run to run.

@MichalStrehovsky
MichalStrehovsky merged commit ace2c51 into dotnet:main Sep 24, 2026
107 of 110 checks passed
@dotnet-milestone-bot dotnet-milestone-bot Bot added this to the 12.0-preview1 milestone Sep 24, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

arch-riscv Related to the RISC-V architecture area-NativeAOT-coreclr community-contribution Indicates that the PR has been added by a community member

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants