ZeroCompute is a lightweight, hardware-accelerated compute execution library for .NET with zero external dependencies. It bridges pure CPU SIMD acceleration with native Windows Direct3D 11 Compute Shaders via COM VTable P/Invoke, delivering high-throughput GPGPU and parallel math execution without requiring external CUDA or OpenCL installations.
-
Unified Compute Abstraction (
IComputeContext): Write algorithmic compute code once; dispatch seamlessly across multi-core CPU SIMD or Direct3D 11 hardware GPUs. - Pure C# Direct3D 11 Dispatcher: Zero third-party C++ wrappers (no SharpDX, no Silk.NET, no Vortice). Calls DirectX COM VTable methods directly with zero marshalling overhead.
- Hardware Fallback Strategy: Automatically detects dedicated GPU hardware (NVIDIA, AMD, Intel); falls back to high-speed vectorized CPU multi-threading if running in a headless or VM environment.
-
Structured Buffers & DMA Transfer: High-speed host
$\leftrightarrow$ device DMA transfers with staging readback buffers, unordered access views (UAV), and shader resource views (SRV). -
Accelerated Kernels:
- Tiled GEMM (
$64 \times 64$ cache blocks) - Vectorized Elementwise Activations (ReLU, Sigmoid, Tanh)
- Matrix Transpose & 2D Reductions
- Tiled GEMM (
- Zero External Dependencies: Standard .NET runtime only.
Install via the .NET CLI:
dotnet add package ZeroCompute.Coreusing ZeroCompute.Core;
// Select best available hardware: GPU if available, else CPU SIMD
using var context = ComputeDevice.GetBestDevice();
Console.WriteLine($"Active Compute Device: {context.DeviceName} (IsGpu: {context.IsGpu})");using ZeroTensor.Core;
using ZeroCompute.Core;
var a = Tensor.RandomUniform(1024, 1024);
var b = Tensor.RandomUniform(1024, 1024);
using var ctx = ComputeDevice.Cpu(); // Or ComputeDevice.Gpu()
var c = ctx.Gemm(a, b);
Console.WriteLine($"Result shape: [{c.Shape[0]}, {c.Shape[1]}]");Tested on Intel Core i7-13700K + NVIDIA GeForce RTX 4070 (Release x64):
| Benchmark Task | CPU Single-Thread | CPU SIMD (AVX2) | Direct3D 11 GPU |
|---|---|---|---|
| GEMM |
|||
| Elementwise ReLU ( |
|||
| Buffer Upload DMA (16 MB) | N/A (In-memory) | N/A |
MIT License © 2026 Phong Võ. Part of the ZeroPlatform project.