ZeroInference is a pure C# ONNX deep learning inference engine and model runtime for .NET with zero external dependencies. It eliminates bulky native C++ runtime binaries (no ONNX Runtime native DLLs, no OpenVINO, no Python dependencies), reading and evaluating .onnx and .zeromodel neural graphs directly in memory with SIMD vectorization and int8 quantization.
- Polymorphic Micro-Kernel (
IInferenceSession): Unified inference contract allowing seamless zero-downtime switching between Pure C# execution and native hardware accelerators (DirectML / ONNX Runtime). - Zero Dependency Pure C# Runtime: No native shared libraries (
onnxruntime.dll,libonnxruntime.so) required. Runs anywhere .NET runs. - Local LLM Streaming (
LocalLlmStreamClient): High-throughput SSE token streaming client for Ollama, OpenAI, and local edge LLM endpoints. - Direct ONNX Model Parser: Stack-allocated Protocol Buffers wire reader (
ProtobufWireReader) parsing ONNX binary graphs directly into executable compute graphs. - Compact
.zeromodelSerialization: Fast binary serialization format with pre-compiled layer topologies and optimized weights layout. - Supported Deep Learning Layers:
- Conv2D (Direct & im2col GEMM convolution)
- Dense / Gemm (Fully-connected linear layers)
- BatchNormalization & LayerNorm
- Activations (ReLU, LeakyReLU, Sigmoid, Tanh, Softmax)
- Pooling (MaxPool2D, AveragePool2D, GlobalAveragePool)
- Reshape, Flatten, Concat, Slice
- Quantization & Vision Post-Processing:
- Int8 Quantizer: Symmetric and asymmetric integer quantization for edge devices.
- Non-Maximum Suppression (NMS): Fast SIMD bounding box filtering with configurable IoU and score thresholds.
- Hardware Agnostic: Executes over
ZeroTensorCPU SIMD orZeroComputeDirect3D 11 GPU compute contexts.
Install via the .NET CLI:
dotnet add package ZeroInference.Coreusing ZeroInference.Core.Engine;
using ZeroInference.Core.Format;
using ZeroTensor.Core;
// 1. Load and parse .onnx model file
using var stream = File.OpenRead("models/classifier.onnx");
var graph = OnnxModelParser.Parse(stream);
// 2. Instantiate inference engine
var engine = new InferenceEngine(graph);
// 3. Prepare input tensor and infer
var input = Tensor.RandomUniform(1, 3, 224, 224);
var outputs = engine.Forward(input);
Console.WriteLine($"Inference Output Shape: [{outputs[0].Shape[0]}, {outputs[0].Shape[1]}]");using ZeroInference.Core.Vision;
var candidateBoxes = new List<BoundingBox>
{
new BoundingBox(10, 10, 50, 50, score: 0.92f, classId: 1),
new BoundingBox(12, 11, 48, 52, score: 0.78f, classId: 1), // Overlapping duplicate
new BoundingBox(100, 120, 60, 40, score: 0.85f, classId: 2)
};
// Filter duplicates with IoU threshold = 0.45
var filtered = NonMaximumSuppression.Filter(candidateBoxes, iouThreshold: 0.45f, scoreThreshold: 0.5f);
Console.WriteLine($"Remaining boxes after NMS: {filtered.Count}");Tested on MobileNet-V2 / ResNet-18 (Release x64):
| Architecture | Model Size | Load Time | CPU SIMD Latency | External DLLs |
|---|---|---|---|---|
| MobileNet-V2 | 0 (Pure C#) | |||
| ResNet-18 (FP32) | 0 (Pure C#) | |||
| ResNet-18 (Int8) | 0 (Pure C#) |
| Version | Release Date | Key Milestones & Highlights |
|---|---|---|
v1.1.0 |
2026-09-16 | Polymorphic Micro-Kernel & Local LLM Streaming: • Introduced IInferenceSession unified execution contract decoupling high-level apps from backends.• Added ZeroInference.Providers.OnnxRuntime provider bridging Microsoft.ML.OnnxRuntime with pure C# pipeline.• Added LocalLlmStreamClient supporting real-time SSE streaming for Ollama & OpenAI-compatible endpoints.• Verified with 23 unit tests across Core & Provider test suites. |
v1.0.0 |
2026-09-09 | Initial Sovereign Release: • Pure C# ONNX protobuf wire reader & layer fusion pipeline. • Conv2D, Dense, BatchNorm, LayerNorm, activations, and pooling. • Int8 quantization engine & Non-Maximum Suppression (NMS). |
MIT License © 2026 Phong Võ. Part of the ZeroPlatform project.