Skip to content

Fix LMUL overflow mismatch on GPU - #205

Closed
Horace-Maxwell wants to merge 1 commit into
Syncleus:masterfrom
Horace-Maxwell:fix/bounty-38-long-mul-cast
Closed

Horace-Maxwell wants to merge 1 commit into
Syncleus:masterfrom
Horace-Maxwell:fix/bounty-38-long-mul-cast

Conversation

@Horace-Maxwell

Copy link
Copy Markdown

Fixes #38

Some OpenCL drivers/devices appear to miscompile 64-bit integer multiplication, producing results where the high 32 bits of a 64-bit product are incorrect/zero. This shows up as CPU/GPU mismatches for expressions like (long) tc * 100 when the mathematical result does not fit in 32 bits.

This change lowers LMUL to a helper (aparapi_lmul) implemented in terms of 32-bit multiplies (and mul_hi) so kernels do not rely on native 64-bit multiply support. The helper is only emitted when the entrypoint (or any called method) contains LMUL.

Tests run:

  • mvn -q -Dtest=com.aparapi.codegen.test.LongMultiplyCastOverflowTest test
  • mvn -q -Dtest='com.aparapi.codegen.test.*Test' -DfailIfNoTests=false test

Build tweaks (for newer JDKs):

  • Remove -XX:MaxPermSize from Surefire argLine
  • Bump Scala tooling/deps to 2.13.18 / scala-maven-plugin 4.9.10
  • Bump JaCoCo to 0.8.13 (Java 25 support)

Work around miscompiled 64-bit multiply by lowering LMUL to 32-bit operations and add regression test for issue #38.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Bounty $50] Inconsistent results between GPU and CPU when integers overflow.

1 participant