ZeroCompute.Core
1.3.0
dotnet add package ZeroCompute.Core --version 1.3.0
NuGet\Install-Package ZeroCompute.Core -Version 1.3.0
<PackageReference Include="ZeroCompute.Core" Version="1.3.0" />
<PackageVersion Include="ZeroCompute.Core" Version="1.3.0" />
<PackageReference Include="ZeroCompute.Core" />
paket add ZeroCompute.Core --version 1.3.0
#r "nuget: ZeroCompute.Core, 1.3.0"
#:package ZeroCompute.Core@1.3.0
#addin nuget:?package=ZeroCompute.Core&version=1.3.0
#tool nuget:?package=ZeroCompute.Core&version=1.3.0
ZeroCompute
ZeroCompute is a high-performance, dual-engine compute execution framework for .NET with zero external dependencies. It provides:
- A CPU Parallel Compute Runtime that systematically exploits all 5 levels of CPU parallelism (Core, Cache Tiling, SIMD, Instruction-Level Parallelism, and Memory-Level Parallelism).
- A Direct3D 11 GPU Compute Engine utilizing native COM VTable P/Invoke compute shaders for massive GPGPU acceleration without third-party C++ wrappers or external CUDA runtimes.
🌟 Architectural Highlights
Sequential For
↓
Level 1: Multi-Core / Task (Worker Pool, Sovereign Physical Core Pinning, Dynamic Chunk Stealing)
↓
Level 2: Cache-Aware Blocking (L1/L2 64x64 Tiling, 64-Byte Line Padding, Kernel Fusion)
↓
Level 3: SIMD Vectorization (SSE / AVX2 / AVX-512, Padé [5/5] Rational Approximations)
↓
Level 4: Instruction-Level Parallelism (4-Way Loop Unrolling, Multi-Accumulator Chains)
↓
Level 5: Memory-Level Parallelism (Line Fill Buffer Saturation, Prefetching & Streaming Stores)
- Axiom: Parallelism $\neq$ Performance: Never naively assumes more threads equals faster throughput. The
ExecutionPlannerprofiles arithmetic intensity ($I = \text{FLOPs}/\text{Byte}$), working-set cache sizing, and CPU microarchitecture to choose between Sequential, Single-Core SIMD, or Multi-Core Cache-Tiled execution. - Sovereign Worker Pool (
ComputeWorkerPool): Pre-spawned, background threads pinned to discrete physical cores via cross-platform affinity (Win32SetThreadAffinityMaskon Windows andlibc sched_setaffinityon Linux with/proc/cpuinfotopology detection). SMT/Hyperthreading is selectively bypassed for compute-bound kernels to eliminate contention on shared FPU execution pipelines. - Dynamic Lock-Free Chunk Stealing: Atomic chunk consumption balances heterogeneous architectures (e.g. Intel Alder Lake / Raptor Lake P-cores and E-cores) with zero task queue locking.
- Microsecond Dispatch Latency: Hybrid spin-wait barriers (10 spins $\to$
Thread.Yield()$\to$ privateAutoResetEvent) achieve $1.85 \ \mu\text{s}$ dispatch latency ($2.0\times$ faster than standard .NETParallel.For). - SIMD Padé Rational Approximation: Vectorizes transcendental activations (
GELU,Tanh) using a high-precision rational polynomial, converting scalar branching into pure SIMD FMA instructions ($5.99\times$ speedup). - Cache-Aware 2D Tiling: $64 \times 64$ sub-matrix tiling ($16\text{ KB}$) ensures temporal working sets fit entirely inside private L1/L2 data caches ($4.18\times$ speedup on GEMM).
- NUMA-Node Awareness & Sovereign Memory Affinity: Dynamic multi-socket NUMA discovery (Windows
GetNumaHighestNodeNumber/ Linux/sys/devices/system/node) and node-local unmanaged memory allocation (VirtualAllocExNuma/ Linuxmmap) eliminating cross-socket QPI/UPI interconnect latency bottlenecks. - Direct3D 11 GPGPU Acceleration: Seamless GPU offloading for massive tensor matrix operations via native DirectX 11 compute shaders.
- 100% Pure C#: Multi-targeting
net8.0,netstandard2.0, andnet462with zero native DLL dependencies.
📦 Installation
dotnet add package ZeroCompute.Core
🚀 Developer API & Usage Examples
1. High-Performance CPU Parallel Loops (Compute.For)
Automatic cost-based execution planning and dynamic chunk stealing:
using ZeroCompute.Core.Cpu;
int n = 1_000_000;
float[] a = new float[n];
float[] b = new float[n];
float[] c = new float[n];
// Range-based parallel loop with optimal chunking
Compute.For(n, (start, end) =>
{
for (int i = start; i < end; i++)
{
c[i] = a[i] * 2.5f + b[i];
}
});
// Item-based parallel loop with deterministic sequential cutoff for small N
Compute.For(100, i =>
{
c[i] = a[i] + b[i];
});
2. Cache-Aware 2D Grid / Tiling (Compute.Tile2D)
Eliminates cache thrashing by restricting active working sets to L1/L2 cache capacity:
int rows = 1024, cols = 1024;
const int blockR = 64, blockC = 64;
Compute.Tile2D(rows, cols, blockR, blockC, (r0, r1, c0, c1) =>
{
// Inner tile [r0..r1, c0..c1] fits completely inside 16 KB L1 Data Cache
for (int r = r0; r < r1; r++)
{
for (int c = c0; c < c1; c++)
{
matrixC[r, c] += matrixA[r, c] * matrixB[r, c];
}
}
});
3. Vectorized Math & Transcendental GELU (Compute.Vector)
4-way unrolled SIMD vectorization with Padé rational approximation:
// High-throughput SIMD vector addition & multiplication
Compute.Vector.Add(aSpan, bSpan, destSpan);
Compute.Vector.Multiply(aSpan, bSpan, destSpan);
// High-precision vectorized GELU activation (5.99x speedup vs scalar Math.Tanh)
Compute.Vector.Gelu(inputSpan, outputSpan);
4. Single-Pass Kernel Fusion (Compute.FusedMap)
Saves 66.7% DRAM bandwidth by fusing multiple operations in L1/registers instead of round-tripping through memory:
// 3 operations executed in a single memory pass
Compute.FusedMap(src, dst, x =>
{
float v = x * 2.0f + 10.0f;
return v > 0f ? v : 0f; // Fused Scale + Bias + ReLU
});
5. Multi-Accumulator Parallel Reduction (Compute.ReduceSum)
Saturates multiple execution ports (Port 0 & Port 1) by unrolling across 4 independent vector accumulators:
float sum = Compute.ReduceSum(dataSpan);
float max = Compute.ReduceMax(dataSpan);
6. Unified Compute Device Selection (CPU / GPU)
using ZeroCompute.Core;
using ZeroTensor.Core;
// Automatically selects Direct3D 11 GPU if present, else fallback to CPU SIMD
using var ctx = ComputeDevice.GetBestDevice();
var a = Tensor.RandomUniform(1024, 1024);
var b = Tensor.RandomUniform(1024, 1024);
// Executes GEMM on selected hardware
var c = ctx.Gemm(a, b);
7. NUMA-Node Topology & Domain Allocation
using ZeroCompute.Core.Cpu;
// Query system NUMA topology (Windows & Linux dual-socket / multi-socket)
int nodeCount = NumaTopology.GetNodeCount();
Console.WriteLine($"System NUMA Nodes: {nodeCount}, Dual-Socket Detected: {NumaTopology.IsDualSocketOrGreater()}");
// Allocate 64MB pin-bound directly on NUMA node 0 memory controller
using var memory = NumaTopology.Allocate(64 * 1024 * 1024, preferredNode: 0);
Span<byte> span = memory.AsSpan();
span.Fill(0xAA);
📊 Empirical Benchmarks
Evaluated on an Intel 12-Logical-Core processor under Windows 11 (Release x64 build):
| Benchmark Workload | Dataset Size ($N$) | Sequential (Scalar) | Standard Parallel.For |
ZeroCompute CPU Runtime | Speedup vs Seq | Architectural Driver |
|---|---|---|---|---|---|---|
| Small Array Add | $N = 1,000$ | $0.001\text{ ms}$ | $0.172\text{ ms}$ | $0.001\text{ ms}$ | $1.00\times$ ($258\times$ vs Par.For) | Zero dispatch overhead via Planner cutoff |
| Medium Array Add | $N = 100,000$ | $0.180\text{ ms}$ | $0.179\text{ ms}$ | $0.037\text{ ms}$ | $4.89\times$ vs Par.For | L2 Cache-resident SIMD vectorization |
| Huge Array Add | $N = 10,000,000$ | $19.46\text{ ms}$ | $14.19\text{ ms}$ | $12.44\text{ ms}$ | $1.56\times$ | Memory Wall (DRAM bandwidth saturation ~75 GB/s) |
| GELU Activation | $N = 2,000,000$ | $32.92\text{ ms}$ | $6.91\text{ ms}$ | $5.49\text{ ms}$ | $5.99\times$ ($1.26\times$ vs Par.For) | SIMD Padé rational approximation + SMT bypass |
| GEMM Matrix Multiply | $512 \times 512$ | $334.11\text{ ms}$ | $334.11\text{ ms}$ | $79.86\text{ ms}$ | $4.18\times$ vs Par.For | $64 \times 64$ L1/L2 Cache Tiling + FMA unrolling |
| Kernel Fusion | $N = 2,000,000$ | $7.85\text{ ms}$ | $7.85\text{ ms}$ | $2.62\text{ ms}$ | $3.00\times$ vs Par.For | Single-pass L1 cache residency (66.7% DRAM saved) |
| Parallel Reduction | $N = 10,000,000$ | $10.12\text{ ms}$ | $12.30\text{ ms}$ | $2.15\text{ ms}$ | $4.71\times$ vs Par.For | 4-way SIMD vector accumulators + ILP unrolling |
| Dispatch Latency | $50,000 \times N = 100$ | $18.1\text{ ms}$ | $181.0\text{ ms}$ ($3.62 \ \mu\text{s}$) | $92.5\text{ ms}$ ($1.85 \ \mu\text{s}$) | $2.0\times$ faster | Sub-microsecond spin-wait barrier signaling |
🏛️ Placement in ZeroPlatform Ecosystem
graph TD
subgraph Tier 0: Primitives
ZP[ZeroPrimitives.Core]
end
subgraph Tier 1: Hardware Compute & Data
ZT[ZeroTensor.Core: N-D Tensor & Striding]
ZC[ZeroCompute.Core: CPU Multi-Level Runtime & D3D11 GPU]
end
subgraph Tier 2: AI & Signal Engines
ZI[ZeroInference.Core: Edge AI ONNX Engine]
ZN[ZeroNeural.Core: Deep Learning & Autograd]
ZS[ZeroSignal.Core: DSP & FFT Filtering]
end
ZP --> ZT
ZP --> ZC
ZT --> ZC
ZC --> ZI
ZC --> ZN
ZC --> ZS
📄 License
MIT License © 2026 Phong Võ. Part of the ZeroPlatform project.
| Product | Versions Compatible and additional computed target framework versions. |
|---|---|
| .NET | net5.0 was computed. net5.0-windows was computed. net6.0 was computed. net6.0-android was computed. net6.0-ios was computed. net6.0-maccatalyst was computed. net6.0-macos was computed. net6.0-tvos was computed. net6.0-windows was computed. net7.0 was computed. net7.0-android was computed. net7.0-ios was computed. net7.0-maccatalyst was computed. net7.0-macos was computed. net7.0-tvos was computed. net7.0-windows was computed. net8.0 is compatible. net8.0-android was computed. net8.0-browser was computed. net8.0-ios was computed. net8.0-maccatalyst was computed. net8.0-macos was computed. net8.0-tvos was computed. net8.0-windows was computed. net9.0 was computed. net9.0-android was computed. net9.0-browser was computed. net9.0-ios was computed. net9.0-maccatalyst was computed. net9.0-macos was computed. net9.0-tvos was computed. net9.0-windows was computed. net10.0 was computed. net10.0-android was computed. net10.0-browser was computed. net10.0-ios was computed. net10.0-maccatalyst was computed. net10.0-macos was computed. net10.0-tvos was computed. net10.0-windows was computed. |
| .NET Core | netcoreapp2.0 was computed. netcoreapp2.1 was computed. netcoreapp2.2 was computed. netcoreapp3.0 was computed. netcoreapp3.1 was computed. |
| .NET Standard | netstandard2.0 is compatible. netstandard2.1 was computed. |
| .NET Framework | net461 was computed. net462 is compatible. net463 was computed. net47 was computed. net471 was computed. net472 was computed. net48 was computed. net481 was computed. |
| MonoAndroid | monoandroid was computed. |
| MonoMac | monomac was computed. |
| MonoTouch | monotouch was computed. |
| Tizen | tizen40 was computed. tizen60 was computed. |
| Xamarin.iOS | xamarinios was computed. |
| Xamarin.Mac | xamarinmac was computed. |
| Xamarin.TVOS | xamarintvos was computed. |
| Xamarin.WatchOS | xamarinwatchos was computed. |
-
.NETFramework 4.6.2
- ZeroPrimitives.Core (>= 1.7.0)
- ZeroTensor.Core (>= 1.5.0)
-
.NETStandard 2.0
- ZeroPrimitives.Core (>= 1.7.0)
- ZeroTensor.Core (>= 1.5.0)
-
net8.0
- ZeroPrimitives.Core (>= 1.7.0)
- ZeroTensor.Core (>= 1.5.0)
NuGet packages (1)
Showing the top 1 NuGet packages that depend on ZeroCompute.Core:
| Package | Downloads |
|---|---|
|
ZeroInference.Core
Pure C# Edge AI model inference engine: ONNX parser, layer fusion, zero-allocation memory planner, INT8 quantization, and fast NMS vision detection for .NET. |
GitHub repositories
This package is not used by any popular GitHub repositories.