TensorSharp.Backends.MLX 2026.10.3

dotnet add package TensorSharp.Backends.MLX --version 2026.10.3
                    
NuGet\Install-Package TensorSharp.Backends.MLX -Version 2026.10.3
                    
This command is intended to be used within the Package Manager Console in Visual Studio, as it uses the NuGet module's version of Install-Package.
<PackageReference Include="TensorSharp.Backends.MLX" Version="2026.10.3" />
                    
For projects that support PackageReference, copy this XML node into the project file to reference the package.
<PackageVersion Include="TensorSharp.Backends.MLX" Version="2026.10.3" />
                    
Directory.Packages.props
<PackageReference Include="TensorSharp.Backends.MLX" />
                    
Project file
For projects that support Central Package Management (CPM), copy this XML node into the solution Directory.Packages.props file to version the package.
paket add TensorSharp.Backends.MLX --version 2026.10.3
                    
#r "nuget: TensorSharp.Backends.MLX, 2026.10.3"
                    
#r directive can be used in F# Interactive and Polyglot Notebooks. Copy this into the interactive tool or source code of the script to reference the package.
#:package TensorSharp.Backends.MLX@2026.10.3
                    
#:package directive can be used in C# file-based apps starting in .NET 10 preview 4. Copy this into a .cs file before any lines of code to reference the package.
#addin nuget:?package=TensorSharp.Backends.MLX&version=2026.10.3
                    
Install as a Cake Addin
#tool nuget:?package=TensorSharp.Backends.MLX&version=2026.10.3
                    
Install as a Cake Tool

TensorSharp

<p align="center"> <img src="imgs/banner_1.png" alt="TensorSharp logo" width="320"> </p>

English | 中文

TensorSharp is a .NET 10 inference engine for local GGUF models. Run it on Windows, macOS or Linux through the CLI, browser chat, or Ollama/OpenAI-compatible APIs, or embed it in your own .NET application. Use managed C# CPU kernels or native CUDA, Metal and Vulkan backends, with support that varies by model.

Current source covers text and reasoning, multimodal input, embeddings, image generation/editing, video with audio, and Agent Skills with code tools. TensorSharp also powers TensorAgent, the local app for iPhone, iPad, Mac and Windows. Start below, or check the project status for capabilities and validation limits; source changes may be ahead of published packages.

Highlights

  • Local models, several interfaces. One engine for CLI use, browser chat and Ollama/OpenAI-compatible APIs, across managed CPU and native accelerator backends.
  • Text and multimodal models. Dense and MoE GGUF models, reasoning, image/audio input and document questions. See Supported models for each family's capabilities.
  • Embeddings and media generation. Text/code embeddings, Qwen-Image-2.1 generation and masked editing, and video models including MiniMax-H3 with audio.
  • Efficient serving. Continuous batching and a paged, Radix prefix-shared KV cache are on by default. Speculative decoding and multi-GPU placement are available for supported models. See Features.
  • Agent Skills and code tools. TensorSharp.AgentHost adds file, shell and document workflows. Server and TensorAgent chats also support bounded sub-agent delegation with private workspaces and read-only defaults.
  • Recorded comparisons. Benchmarks against llama.cpp use identical GGUF files and hardware; results apply to the measured model, backend and workload. See Benchmarks.
  • TensorAgent apps. Local chat, attachments, skills and saved work on phones and desktops. See the Mac/Windows installation guide and app coverage notes.

Quick Start

TensorSharp CLI and server

The Releases page provides self-contained CLI and Server archives for Windows x64 (CPU/CUDA), Linux x64 (CPU/CUDA), and macOS arm64.

To build from source you need the full .NET 10 SDK (how to install it), git, curl, CMake 3.20+, and the toolchain for your GPU. Then run the verified Gemma 4 E4B model (7.48 GiB). On Windows with an NVIDIA GPU (PowerShell):

git clone https://github.com/zhongkaifu/TensorSharp.git; Set-Location TensorSharp
New-Item -ItemType Directory -Force models | Out-Null
curl.exe -L --fail "https://huggingface.co/ggml-org/gemma-4-E4B-it-GGUF/resolve/main/gemma-4-E4B-it-Q8_0.gguf?download=true" -o models\gemma-4-E4B-it-Q8_0.gguf
'Answer in one short sentence: what is TensorSharp?' | Set-Content prompt.txt
$env:TENSORSHARP_GGML_NATIVE_ENABLE_CUDA = 'ON'
dotnet run --project TensorSharp.Cli -c Release -p:TensorSharpSkipMlxNative=true -- --model models\gemma-4-E4B-it-Q8_0.gguf --input prompt.txt --max-tokens 128 --backend ggml_cuda

On other machines, change the backend (see Pick a backend):

  • macOS (Apple Silicon): drop the CUDA environment variable and use --backend ggml_metal.
  • Linux + NVIDIA: prefix the dotnet run with TENSORSHARP_GGML_NATIVE_ENABLE_CUDA=ON and use --backend ggml_cuda.
  • AMD / Intel / NVIDIA Vulkan: set TENSORSHARP_GGML_NATIVE_ENABLE_VULKAN=ON and use --backend ggml_vulkan.

Host the same model as a server: a browser chat at http://localhost:5000 plus Ollama- and OpenAI-compatible APIs.

dotnet run --project TensorSharp.Server.Host -c Release -p:TensorSharpSkipMlxNative=true -- --model models/gemma-4-E4B-it-Q8_0.gguf --backend ggml_cuda --max-tokens 512

The server listens on 0.0.0.0:5000 with no built-in authentication or TLS; keep it behind a firewall or an authenticated HTTPS reverse proxy.

Pick a backend

Your hardware Backend
Apple Silicon (Mac) --backend ggml_metal
Windows / Linux + NVIDIA GPU --backend ggml_cuda
Windows / Linux + AMD / Intel / NVIDIA GPU --backend ggml_vulkan
No GPU --backend ggml_cpu (native kernels), or --backend cpu (pure C#, no native dependencies)

The Getting started guide has the rest: installing the SDK on each platform, multi-GPU and multi-node runs, NVIDIA DGX Spark, multimodal input, embeddings, and making it fast. Every option is in the CLI and Server references, and both programs print them with --help.

dotnet build TensorSharp.slnx also builds TensorAgent's available desktop heads and the iOS simulator head on Apple Silicon when the selected SDK has the required MAUI workloads and staged native/Python files. Missing prerequisites skip the affected app head with a warning; see TensorAgent build instructions.

TensorAgent Desktop: download, install, chat

Download the latest release and expand Assets. Choose a file beginning with tensoragent-desktop-:

Platform Download and install
macOS 14+ on Apple Silicon tensoragent-desktop-<version>-osx-arm64.dmg: open and drag TensorAgent to Applications. A PKG installer and ZIP are also available.
Windows x64 tensoragent-desktop-<version>-win-x64-cpu.msi, or win-x64-cuda.msi for a compatible NVIDIA GPU/driver: install and open TensorAgent from Start. ZIPs are also available.

The app includes its .NET runtime and native engine. Windows needs WebView2 Evergreen Runtime if missing; Python/Node are optional tools for skills. The current built-in catalog needs at least 12 GB system RAM, and model weights download separately. Open ☰ → Models → Download → Use, then type a message. The Desktop user guide covers package verification, unsigned-app prompts, first-run setup, attachments, skills, updates and troubleshooting. Historical releases may have no Desktop assets until a release runs the updated workflow; iPhone/iPad remain source builds.

Supported model families at a glance

Backend, modality, feature support, and validation coverage vary by model. See the supported models tables, the model cards, and the embedding guide for details.

Recent source additions include Qwen-Image-2.1 masked edits with exact protected pixels and optional processing of the selected region, twelve TensorAgent LoRA plug-ins for speed, style and editing, and Qwen3.8 Flash Next on a 48 GB Mac using SSD-backed weights. Multi-GPU --layer-split and --tp are separate controls; support and performance depend on the architecture and quantization. These source features may be ahead of published packages; Desktop installers appear in releases built with the updated Release Binaries workflow, while older releases may lack them.

See it in action

One engine, four ways to use it, each an unedited capture of a real run.

<table> <tr> <td align="center" width="50%"><img src="website/assets/screenshots/tensorsharp-cli.png" alt="TensorSharp.Cli in a terminal: an interactive chat with Gemma 4 E4B that reads this README and answers questions about it" width="250"><br><b>TensorSharp.Cli</b><br>Models in your terminal</td> <td align="center" width="50%"><img src="website/assets/screenshots/tensorsharp-webui.png" alt="The TensorSharp Web UI: Qwen3.8 27B compared two mortgages by writing and running a Python script" width="400"><br><b>TensorSharp.Server.Host</b><br>Web UI chat and Ollama/OpenAI-compatible APIs</td> </tr> <tr> <td align="center"><img src="website/assets/screenshots/tensoragent-iphone.png" alt="TensorAgent in the iPhone 17 Pro simulator: Gemma 4 E2B uses ggml_cpu and runs an in-app Python script to scale a recipe" width="140"><br><b>TensorAgent · iPhone simulator</b><br>CPU inference and in-app Python; this is a simulator capture</td> <td align="center"><img src="website/assets/screenshots/tensoragent-mac.png" alt="TensorAgent on a Mac: a saved Qwen-Image 2.1 edit changes the TensorSharp banner background to a starry blue night sky" width="400"><br><b>TensorAgent on Mac</b><br>A saved Qwen-Image 2.1 edit in the Mac app</td> </tr> </table>

What each run shows, step by step: Screenshots.

Benchmarks

TensorSharp and llama.cpp run identical GGUF files on the same NVIDIA RTX 3080 Laptop GPU (16 GB), each on its GGML CUDA and Vulkan builds. Each number is TensorSharp's speedup over llama.cpp on the same backend (geomean, single-stream, greedy, MTP off); above 1.0× means TensorSharp is faster.

Model Backend decode prefill TTFT
Gemma 4 E4B it (Q8_0, dense multimodal) CUDA 1.02× 1.28× 1.27×
Gemma 4 E4B it (Q8_0, dense multimodal) Vulkan 1.00× 1.05× 1.03×
Gemma 4 12B it (QAT UD-Q4_K_XL, dense) CUDA 1.04× 1.17× 1.16×
Gemma 4 12B it (QAT UD-Q4_K_XL, dense) Vulkan 1.21× 1.04× 1.03×
Qwen 3.6 35B-A3B (UD-IQ2_XXS, MoE) CUDA 0.98× 1.28× 1.27×
Qwen 3.6 35B-A3B (UD-IQ2_XXS, MoE) Vulkan 0.87× 1.04× 1.03×
Qwen 3.6 27B (UD-IQ2_XXS, dense) CUDA 1.07× 0.96× 0.95×
Qwen 3.6 27B (UD-IQ2_XXS, dense) Vulkan 1.02× 0.85× 0.84×

What these numbers mean, how to rerun them, and the head-to-heads of models too large for this GPU: Benchmarks.

Documentation

New here? The sections above are all you need to get running. Everything else is detailed reference:

Doc What's inside
TensorAgent Desktop user guide Mac/Windows downloads, DMG/PKG/MSI/ZIP installation, model setup, first chat, attachments, skills, updates and troubleshooting
TensorSharp and TensorAgent book guide Building LLM Inference Engines and Agentic Runtimes from Scratch, plus From Tensors to Tokens: introductions, Amazon links, and repository reading paths
Getting started The full first-run guide: the .NET SDK on each platform, every backend, multi-GPU and multi-node runs, NVIDIA DGX Spark, embeddings, choosing a backend, and making it fast
Supported models Implemented model families and their validation scope: example GGUFs, modalities, thinking, tools, and speculative decoding
Benchmarks TensorSharp against llama.cpp on the same GPU and files, and the head-to-heads of larger models
Screenshots The CLI, the Web UI, and TensorAgent on iPhone and Mac at work, with what each run did
Model Downloads Per-model huggingface-cli download + run quick reference (quant tiers, projectors, companions)
Usage Full CLI reference (options, interactive REPL, JSONL batch), server hosting, logging, HTTP API examples, backends, and the env-var matrix
Features Deep dives on continuous batching, speculative decoding, tool calling, thinking mode, multimodal, MoE, KV codecs, and more
Configuration files Put options in a reusable JSON file with ${variables} and auto-downloading models
Development Prerequisites, building the native GGML/MLX libraries, repository layout, package boundaries, internal architecture, and the test harness
Per-model architecture cards End-to-end docs of each architecture (forward graph, components, parameters, prefill/decode optimizations)
Paged attention & continuous batching The vLLM-style paged KV cache, prefix sharing, and iteration-level scheduler
Agent Skills & agentic work The SKILL.md format, progressive disclosure and its budget, the in-process tool loop, sandboxed code execution, workspaces and artifacts, the path/ZIP/exec security model, and the HTTP + C# surfaces
Multiple agents Automatic task delegation, private child workspaces, dependency scheduling, permission limits, server controls, and reproducible evaluation
Browser automation skill (Playwright) Running the bundled playwright skill, which drives a browser through @playwright/cli via skills_run: the flags it needs, the macOS Chromium-sandbox config, account handoff, and TensorAgent desktop hosting (not iOS)
Speculative decoding The three-layer design (model adapter / algorithm / speculator weights), the shipped auto / draft-head / block / ngram algorithms, and what to write to add a new one
Environment variable feature matrix Which high-impact runtime flags affect which models, backends, and prompt types
Engine comparison report Full per-scenario TensorSharp vs llama.cpp tables
ggml_metal vs llama.cpp Head-to-head prefill/decode on Apple Silicon, the four graph-construction gaps it found, and what each was worth
Test/benchmark matrix runner Sweep model × backend × feature × env-var cells and generate regression reports
Server API examples Complete curl and Python examples for the server surface

Current Status

Actively developed, and the source tree runs ahead of the published packages.

Area Where it stands
Models A dozen autoregressive families plus text diffusion, image generation and editing, and video with audio. See Supported models.
Inference hosts CLI, Web UI, Ollama- and OpenAI-compatible APIs, and TensorAgent for iPhone, iPad, Mac and Windows. The release workflow packages Mac/Windows Desktop; historical releases may lack those assets. iPhone/iPad use source builds.
Backends Pure C# CPU, direct CUDA/cuBLAS, MLX Metal, and GGML CPU/Metal/CUDA/Vulkan, with per-architecture exceptions.
Serving features Continuous batching with a shared prefix cache, speculative decoding, tensor parallelism, structured output, and tool calling.
Agentic work Agent Skills, sandboxed file and shell tools, and bounded sub-agents. See Agent Skills and Multiple agents.
TensorAgent Twelve catalog entries, saved chats and artifacts, masked image edits and LoRA choices, eight interface languages, and persisted text-turn statistics. Media generation has been measured on a Mac; iOS media generation and Windows image/audio/video generation remain unverified.

Per-area detail (which architecture runs on which backend, which features each family supports, and the known limits) is in the status matrix.

Author

Zhongkai Fu

License

See LICENSE for details.

Learn with the books

Qwen inference and agentic runtimes Gemma 4 and multimodal inference
<a href="https://www.amazon.com/dp/B0HJQ4VQ31"><img src="website/assets/building-llm-inference-engines-cover.jpg" alt="Building LLM Inference Engines and Agentic Runtimes from Scratch: Qwen Dense and MoE Models with TensorSharp and TensorAgent" width="190"></a> <a href="https://www.amazon.com/dp/B0H9P44QZZ"><img src="website/assets/from-tensors-to-tokens-cover.jpg" alt="From Tensors to Tokens: Building a Multimodal LLM Inference Engine from Scratch with TensorSharp and Gemma 4 E4B" width="190"></a>
Building LLM Inference Engines and Agentic Runtimes from Scratch: Qwen Dense and MoE Models with TensorSharp and TensorAgent From Tensors to Tokens: Building a Multimodal LLM Inference Engine from Scratch with TensorSharp and Gemma 4 E4B
Build Qwen dense/MoE inference and controlled agent workflows in C#. Follow tensors, tokenization, attention, expert routing, quantization, and caching through GPU acceleration, multimodal execution, tools, skills, sandboxed code, and desktop/mobile deployment with TensorSharp and TensorAgent. Build a multimodal inference engine in C#/.NET with Gemma 4 E4B, from tensors, GGUF model loading, quantization, and tokenization to text, image, video, and audio execution. Connect correctness checks and serving optimizations to the running TensorSharp code.
Buy on Amazon Buy on Amazon

Explore both books and their repository reading paths

<p align="center"><a href="https://buymeacoffee.com/zhongkaifu"><img src="https://cdn.buymeacoffee.com/buttons/v2/default-yellow.png" alt="Buy Me A Coffee" height="50"></a><br> <sub>TensorSharp/TensorAgent is free. If you like it, a coffee keeps the work on it going.</sub></p>

Product Compatible and additional computed target framework versions.
.NET net10.0 is compatible.  net10.0-android was computed.  net10.0-browser was computed.  net10.0-ios was computed.  net10.0-maccatalyst was computed.  net10.0-macos was computed.  net10.0-tvos was computed.  net10.0-windows was computed. 
Compatible target framework(s)
Included target framework(s) (in package)
Learn more about Target Frameworks and .NET Standard.

NuGet packages (3)

Showing the top 3 NuGet packages that depend on TensorSharp.Backends.MLX:

Package Downloads
TensorSharp.Models

Model architecture implementations, multimodal encoders, and model execution helpers for TensorSharp.

TensorSharp.Cli

Command-line host for TensorSharp local inference, diagnostics, prompt inspection, and JSONL batch processing.

TensorSharp.Server

ASP.NET Core server package for TensorSharp with embedding and chat inference, web UI, and Ollama/OpenAI-compatible APIs.

GitHub repositories

This package is not used by any popular GitHub repositories.

Version Downloads Last Updated
2026.10.3 0 10/4/2026
2026.9.29 84 9/30/2026
2026.9.1 150 9/17/2026
3.4.0 140 9/12/2026
3.1.2 188 7/21/2026