ZeroLlm.Core
1.3.0
dotnet add package ZeroLlm.Core --version 1.3.0
NuGet\Install-Package ZeroLlm.Core -Version 1.3.0
<PackageReference Include="ZeroLlm.Core" Version="1.3.0" />
<PackageVersion Include="ZeroLlm.Core" Version="1.3.0" />
<PackageReference Include="ZeroLlm.Core" />
paket add ZeroLlm.Core --version 1.3.0
#r "nuget: ZeroLlm.Core, 1.3.0"
#:package ZeroLlm.Core@1.3.0
#addin nuget:?package=ZeroLlm.Core&version=1.3.0
#tool nuget:?package=ZeroLlm.Core&version=1.3.0
🧠 ZeroLlm: Sovereign Pure C# Small Language Model (SLM) Runtime
Architectural Standard: 100% Pure C#, Zero Unmanaged Wrappers (no
llama.cppor Python runtime required), Zero External Dependencies, Multi-Targeting across.NET 8.0,.NET Framework 4.6.2, and.NET Standard 2.0.
ZeroLlm is an in-process, high-throughput Small Language Model (SLM) inference runtime and Transformer execution engine engineered for mission-critical edge gateways, industrial controllers, and enterprise agent systems. Operating within Tier 3 (Perception & AI) of the ZeroPlatform ecosystem, it bridges high-speed tokenization (ZeroTokenizer), prompt grammar constraints (ZeroPrompt), and cognitive agent swarms (ZeroAgent).
⚡ Performance Benchmarks (Pure C# .NET 8 CPU)
Through rigorous algorithmic and low-level kernel optimizations (Precomputed RoPE, Zero-Lock KV SequenceView, Pinned SwiGLU unrolling, and 4-way pipelined accumulator streams), ZeroLlm achieves 20.3x speedup on pure CPU:
| Optimization Level | Data Type | Prompt Tokens | Generated Tokens | Decoding Speed | Total Latency | Speedup | Working Memory |
|---|---|---|---|---|---|---|---|
| Baseline (Scalar) | FP32 | 52 | 20 | 8.7 tok/s | 2,298 ms | 1.0x | ~80 MB |
| SIMD Vectorized | FP32 | 52 | 100 | 147.5 tok/s | 677 ms | 17.0x | ~48 MB |
| Deep-Optimized (v1.3.0) | FP32 | 52 | 100 | 176.4 tok/s | 566.8 ms | 20.3x 🚀 | ~48 MB |
| Deep-Optimized Quantized | Q8_0 | 52 | 100 | 99.0 tok/s | 1,010.5 ms | 11.4x | 13.68 MB ⚡ |
🏛️ Key Subsystem Capabilities
| Component | Namespace | Description |
|---|---|---|
MoELayer |
ZeroLlm.Core.Layers |
DeepSeek-style Sparse Mixture-of-Experts with 1 Shared Expert + Top-2 Routed Experts and auxiliary load-balancing loss. |
PagedKvCache & KvSequenceView |
ZeroLlm.Core.Memory |
Virtual memory page-block allocator (16 tokens/block). KvSequenceView resolves physical blocks once per layer, eliminating 2,400 lock acquisitions/token. |
RoPE (Precomputed Cache) |
ZeroLlm.Core.Layers |
Precomputed trigonometric $(\cos, \sin)$ table lookup. Eliminates 1,152 transcendental mathematical calls/token. |
QuantizedKernels |
ZeroLlm.Core.Quantization |
High-performance Q8_0 and Q4_0 dot-product kernels unrolled with 4 independent accumulation streams to eliminate CPU pipeline latency stalls. |
SwiGLU |
ZeroLlm.Core.Layers |
Pinned-pointer unrolled Swish-Gated Linear Unit kernel ($Swish(gate) \odot up$) with zero array bounds checking overhead. |
GqaAttention |
ZeroLlm.Core.Layers |
Grouped-Query Attention (GQA) kernel with multi-query head mapping and vectorized SIMD Value accumulation. |
LlmSampler |
ZeroLlm.Core.Sampling |
Token sampling engine supporting Greedy, Temperature scaling, Top-K, Top-P (Nucleus), Min-P, Repetition penalty, and Grammar Logit Processors. |
LlmEngine |
ZeroLlm.Core.Engine |
Zero-allocation autoregressive execution pipeline supporting synchronous completions and asynchronous token streaming (IAsyncEnumerable<string>). |
GgufReader & GgufWriter |
ZeroLlm.Core.Format |
Pure C# binary reader and serializer for GGUF v2/v3 model formats, supporting FP32, Q8_0, and Q4_0 quantized tensors. |
LlmTrainer |
ZeroLlm.Core.Training |
Masked Causal LM trainer with AdamW optimizer supporting Full parameter SFT, Head/Embedding adaptation, and MoE routing loss. |
🚀 Quickstart
1. Execute Text Completion with Pre-Quantized Model (Q8_0)
using ZeroLlm.Core.Engine;
using ZeroLlm.Core.Format;
using ZeroLlm.Core.Sampling;
using ZeroTokenizer.Core.Vietnamese;
// 1. Load GGUF v3 Model
var model = GgufReader.LoadModel("models/vietnamese_erp_mds_official_q8_0.gguf");
var tokenizer = VietnameseErpTokenizer.CreateDefault();
// 2. Initialize Engine
using var engine = new LlmEngine(model, tokenizer, SamplingConfig.Greedy);
// 3. Generate Completion
string response = await engine.CompleteAsync("Kiểm tra tồn kho mặt hàng thép cuộn 10mm tại Kho Tổng");
Console.WriteLine(response);
2. Stream Tokens Asynchronously
await foreach (string token in engine.GenerateStreamAsync("Tra cứu đơn hàng bán SO-2026-MDS01"))
{
Console.Write(token);
}
3. Grammar-Constrained Tool Calling Decoding
using ZeroPrompt.Core.Grammar;
var grammarProcessor = registry.CreateGrammarProcessor(tokenizer);
var sampling = new SamplingConfig
{
Temperature = 0.0f,
MaxTokens = 128,
ContextLogitProcessor = grammarProcessor.Process
};
string toolCall = await engine.CompleteAsync(prompt, sampling);
// Guaranteed to produce 100% syntactically valid JSON tool call
📄 License
Architected and developed by Phong Võ (kzxl) for the ZeroUniverse / ZeroPlatform ecosystem. Released under the MIT License.
| Product | Versions Compatible and additional computed target framework versions. |
|---|---|
| .NET | net5.0 was computed. net5.0-windows was computed. net6.0 was computed. net6.0-android was computed. net6.0-ios was computed. net6.0-maccatalyst was computed. net6.0-macos was computed. net6.0-tvos was computed. net6.0-windows was computed. net7.0 was computed. net7.0-android was computed. net7.0-ios was computed. net7.0-maccatalyst was computed. net7.0-macos was computed. net7.0-tvos was computed. net7.0-windows was computed. net8.0 is compatible. net8.0-android was computed. net8.0-browser was computed. net8.0-ios was computed. net8.0-maccatalyst was computed. net8.0-macos was computed. net8.0-tvos was computed. net8.0-windows was computed. net9.0 was computed. net9.0-android was computed. net9.0-browser was computed. net9.0-ios was computed. net9.0-maccatalyst was computed. net9.0-macos was computed. net9.0-tvos was computed. net9.0-windows was computed. net10.0 was computed. net10.0-android was computed. net10.0-browser was computed. net10.0-ios was computed. net10.0-maccatalyst was computed. net10.0-macos was computed. net10.0-tvos was computed. net10.0-windows was computed. |
| .NET Core | netcoreapp2.0 was computed. netcoreapp2.1 was computed. netcoreapp2.2 was computed. netcoreapp3.0 was computed. netcoreapp3.1 was computed. |
| .NET Standard | netstandard2.0 is compatible. netstandard2.1 was computed. |
| .NET Framework | net461 was computed. net462 is compatible. net463 was computed. net47 was computed. net471 was computed. net472 was computed. net48 was computed. net481 was computed. |
| MonoAndroid | monoandroid was computed. |
| MonoMac | monomac was computed. |
| MonoTouch | monotouch was computed. |
| Tizen | tizen40 was computed. tizen60 was computed. |
| Xamarin.iOS | xamarinios was computed. |
| Xamarin.Mac | xamarinmac was computed. |
| Xamarin.TVOS | xamarintvos was computed. |
| Xamarin.WatchOS | xamarinwatchos was computed. |
-
.NETFramework 4.6.2
- Microsoft.Bcl.AsyncInterfaces (>= 8.0.0)
- System.Buffers (>= 4.5.1)
- System.Memory (>= 4.5.5)
- System.Runtime.CompilerServices.Unsafe (>= 6.0.0)
- ZeroConcurrency (>= 1.4.0)
- ZeroPrimitives.Core (>= 1.7.0)
- ZeroTensor.Core (>= 1.5.0)
- ZeroTokenizer.Core (>= 1.0.0)
-
.NETStandard 2.0
- Microsoft.Bcl.AsyncInterfaces (>= 8.0.0)
- System.Buffers (>= 4.5.1)
- System.Memory (>= 4.5.5)
- System.Runtime.CompilerServices.Unsafe (>= 6.0.0)
- ZeroConcurrency (>= 1.4.0)
- ZeroPrimitives.Core (>= 1.7.0)
- ZeroTensor.Core (>= 1.5.0)
- ZeroTokenizer.Core (>= 1.0.0)
-
net8.0
- ZeroConcurrency (>= 1.4.0)
- ZeroPrimitives.Core (>= 1.7.0)
- ZeroTensor.Core (>= 1.5.0)
- ZeroTokenizer.Core (>= 1.0.0)
NuGet packages (1)
Showing the top 1 NuGet packages that depend on ZeroLlm.Core:
| Package | Downloads |
|---|---|
|
ZeroAgent.Core
Pure C# enterprise-grade cognitive AI Agent runtime (ReAct loop, tool registry, episodic memory via ZeroVector, and multi-agent swarm) for .NET. |
GitHub repositories
This package is not used by any popular GitHub repositories.