EasyOcrSharp 3.1.0

dotnet add package EasyOcrSharp --version 3.1.0
                    
NuGet\Install-Package EasyOcrSharp -Version 3.1.0
                    
This command is intended to be used within the Package Manager Console in Visual Studio, as it uses the NuGet module's version of Install-Package.
<PackageReference Include="EasyOcrSharp" Version="3.1.0" />
                    
For projects that support PackageReference, copy this XML node into the project file to reference the package.
<PackageVersion Include="EasyOcrSharp" Version="3.1.0" />
                    
Directory.Packages.props
<PackageReference Include="EasyOcrSharp" />
                    
Project file
For projects that support Central Package Management (CPM), copy this XML node into the solution Directory.Packages.props file to version the package.
paket add EasyOcrSharp --version 3.1.0
                    
#r "nuget: EasyOcrSharp, 3.1.0"
                    
#r directive can be used in F# Interactive and Polyglot Notebooks. Copy this into the interactive tool or source code of the script to reference the package.
#:package EasyOcrSharp@3.1.0
                    
#:package directive can be used in C# file-based apps starting in .NET 10 preview 4. Copy this into a .cs file before any lines of code to reference the package.
#addin nuget:?package=EasyOcrSharp&version=3.1.0
                    
Install as a Cake Addin
#tool nuget:?package=EasyOcrSharp&version=3.1.0
                    
Install as a Cake Tool

<div align="center">

<img src="https://raw.githubusercontent.com/FarhanLodi/EasyOcrSharp/main/src/EasyOcrSharp/Assets/icon.png" alt="EasyOcrSharp" width="120" height="120">

EasyOcrSharp

High-accuracy, fully-offline OCR for .NET — EasyOCR's neural models, running natively on ONNX Runtime. No Python.

NuGet Downloads License: MIT .NET 10 AOT ready

Quick start · Documents & tables · PDF · Languages · Production

</div>


EasyOcrSharp runs EasyOCR's exact CRAFT text detector and CRNN recognizers — exported to ONNX and executed through Microsoft.ML.OnnxRuntime. You get EasyOCR-grade accuracy in a tiny managed package: no Python interpreter, no PyTorch, no native OCR binaries, and nothing ever leaves the machine.

await using var ocr = new EasyOcrService();
var result = await ocr.ExtractTextFromImage("receipt.png", new[] { "en" });
Console.WriteLine(result.FullText);

✨ Highlights

🌍 86 languages 13 script families: Latin, Cyrillic, Arabic, Devanagari, Bengali, Chinese (Simplified & Traditional), Korean, Japanese, Thai, Tamil, Telugu, Kannada
📦 ~3 MB package Models download on demand and cache locally — nothing is bundled
🔒 Verified & private Every model download is SHA256-checked; OCR runs fully offline
Fast Concurrent multi-region recognition; automatic CUDA GPU with CPU fallback; tunable threads
🧩 Flexible input File / Stream / byte[] / Image / PDF, region-of-interest, recognize-from-boxes, word/line/paragraph grouping, auto language detection
🧱 Document structure AnalyzeDocumentAsync: layout regions, tables as HTML, formulas, seals & reading order (PP-StructureV3, built in), with Markdown / JSON export
📄 Document-ready Searchable-PDF output, plus hOCR / ALTO / TSV / JSON exporters
✍️ Handwriting TrOCR encoder/decoder recognition for handwritten text — something Python EasyOCR cannot do at all
🔳 Barcodes & QR Read barcodes and QR codes alongside text in a single pass
🖍️ Redaction Find by regex/keyword and permanently paint over it — Luhn-checked cards, mod-97 IBANs, emails, SSNs
🧾 Fields & tables Anchor-based key/value extraction with invoice presets, plus recovered tables as DataTable / CSV
🎯 Accurate fields Allow/block-lists, beam / word-beam decoders, per-box rotation, custom recognizers, exposed detection/grouping/contrast thresholds, post-OCR correction & CER/WER metrics
📐 Word geometry True per-word and per-character boxes from CTC alignment — not width estimates
🖥️ CLI + web sample dotnet tool install -g EasyOcrSharp.Cli, plus a Dockerized ASP.NET Core service sample
🩺 Scan-ready Deskew, orientation correction, adaptive binarize, denoise & sharpen, plus model-based document orientation & page unwarp
📊 Production-grade OpenTelemetry metrics & tracing, health checks, resilient resumable downloads, batch API
🛠️ Modern .NET AOT- & single-file-friendly, DI-ready, .NET 10
🖼️ MIT all the way down Imaging runs on EasyImageSharp and the structure engine is built in — no build-time licence key, no commercial tier, no split-licensed package anywhere in the graph

🆚 Why EasyOcrSharp?

EasyOcrSharp Python EasyOCR Cloud OCR APIs Tesseract (.NET wrappers)
Runtime Pure .NET + ONNX Python + PyTorch Remote HTTP service Native binary via P/Invoke
Install dotnet add package pip + CUDA toolchain API key + billing Native libs + tessdata files
Privacy 🟢 100% offline 🟢 offline 🔴 data leaves the machine 🟢 offline
Accuracy 🟢 EasyOCR neural models 🟢 EasyOCR neural models 🟢 high 🟡 classical (weaker on hard text)
Tables / layout 🟢 built-in (PP-Structure) 🟡 add-ons 🟢 yes 🔴 no
PDF in & searchable out 🟢 built-in 🔴 DIY 🟢 yes 🔴 DIY
GPU 🟢 CUDA (opt-in package) 🟢 CUDA n/a 🔴 no
Native AOT / trimming 🟢 yes n/a n/a 🟡 limited
Cost 🟢 free (MIT) 🟢 free 🔴 per-call 🟢 free

Same models as Python EasyOCR, none of the Python. Fully local, so no per-page cost and no data egress.


📥 Installation

dotnet add package EasyOcrSharp

PDF input and searchable-PDF output are built in — no extra package. For NVIDIA GPU acceleration (Windows/Linux x64, CUDA 12+):

dotnet add package EasyOcrSharp.Gpu

Upgrading from 1.x? v2 replaced the ~1.5 GB embedded Python + PyTorch runtime with native ONNX. The public API (EasyOcrService, OcrResult, OcrLine, OcrBoundingBox) is unchanged.

Upgrading from 2.x? v3 changes the imaging library behind the pixel types on the public API — see Imaging below and the changelog. Method names, parameters and results are unchanged; what changes is the namespace the Image<Rgb24> in your using directives comes from.

Imaging

Decoding, encoding and every pixel operation run on EasyImageSharp — MIT-licensed, fully managed, AOT- and trimming-friendly, with no build-time licence key and no commercial tier for consumers to inherit. It is maintained by the same author as this library, so a fix OCR needs does not wait on a third party.

It brings PNG, JPEG (baseline, progressive and CMYK), WebP (decode), GIF, BMP, TIFF (including CCITT G3/G4 and JPEG-in-TIFF), TGA, Netpbm, QOI and ICO, so Image.Load accepts anything you are likely to scan or be sent. Pixel types (Image<Rgb24>, Rgba32, ...), geometry (Point, Size, Rectangle) and the Mutate / Clone processing pipeline live in the EasyImageSharp, EasyImageSharp.PixelFormats and EasyImageSharp.Processing namespaces:

using EasyImageSharp;                 // Image, Image.Load, Color, Rectangle
using EasyImageSharp.PixelFormats;    // Rgb24, Rgba32, L8
using EasyImageSharp.Processing;      // Mutate/Clone: Resize, Rotate, Crop, Grayscale, Deskew, ...

using var img = Image.Load<Rgb24>("page.png");
var result = await ocr.ExtractTextFromImage(img, new[] { "en" });

🚀 Quick start

using EasyOcrSharp.Services;

await using var ocr = new EasyOcrService();

var result = await ocr.ExtractTextFromImage("sample.png", new[] { "en" });

Console.WriteLine(result.FullText);

foreach (var line in result.Lines)
    Console.WriteLine($"{line.Text}  (confidence {line.Confidence:P0})");

The first call for a language downloads its model (cached afterwards). Detected text is returned as reading-order lines (top-to-bottom), matching EasyOCR's readtext().

📦 The result model

public sealed record OcrResult
{
    public string FullText { get; }                 // all lines joined by newlines (reading order)
    public IReadOnlyList<OcrLine> Lines { get; }
    public IReadOnlyList<string> Languages { get; }
    public TimeSpan Duration { get; }
    public bool UsedGpu { get; }
    public int SourceWidth { get; }                 // dimensions OCR ran on (0 if unknown) — handy for exporters
    public int SourceHeight { get; }
}

public sealed record OcrLine
{
    public string Text { get; }
    public double Confidence { get; }                       // 0..1
    public IReadOnlyList<OcrPoint> BoundingPolygon { get; } // 4 corners
    public OcrBoundingBox BoundingBox { get; }              // MinX/MinY/MaxX/MaxY + Width/Height/Center
}

<br>

🧭 Core OCR

Input sources

OCR from a file path, a Stream, raw encoded bytes, or an already-decoded EasyImageSharp image:

await ocr.ExtractTextFromImage("photo.jpg",            new[] { "en" });
await ocr.ExtractTextFromImage(stream,                  new[] { "en" });
await ocr.ExtractTextFromImage(File.ReadAllBytes("p"),  new[] { "en" });   // byte[]
await ocr.ExtractTextFromImage(memory,                  new[] { "en" });   // ReadOnlyMemory<byte>
await ocr.ExtractTextFromImage(image,                   new[] { "en" });   // Image<Rgb24> (caller-owned)

All overloads accept an optional RecognitionOptions and a CancellationToken.

Recognition options

var options = new RecognitionOptions
{
    Grouping = TextGrouping.Line,    // Word | Line (default) | Paragraph
    MinConfidence = 0.3,             // drop results below this confidence
    MaxDegreeOfParallelism = 8,      // regions recognized concurrently (default: CPU count)
    AdjustContrast = true,           // low-confidence contrast-retry pass (EasyOCR's 2nd pass)
    Region = null,                   // optional region of interest (see below)
};

var result = await ocr.ExtractTextFromImage("doc.png", new[] { "en" }, options);
Grouping Behaviour
Word One result per detected box (≈ per word)
Line Adjacent boxes merged into lines (default; matches EasyOCR)
Paragraph Nearby lines merged into paragraph blocks

Region of interest

Restrict OCR to a rectangle — ideal for a fixed field (price, license plate, banner) and faster than scanning the whole image. Boxes are always reported in the original image's coordinates.

// Absolute pixels:
var roi = new RecognitionOptions { Region = OcrRegion.Pixels(x: 40, y: 320, width: 500, height: 80) };

// Resolution-independent fractions — e.g. the bottom third:
var bottom = new RecognitionOptions { Region = OcrRegion.Fraction(0, 0.66, 1, 0.34) };

var result = await ocr.ExtractTextFromImage("receipt.png", new[] { "en" }, bottom);

Multiple languages

Pass several codes for mixed-script images — each region is read by every requested script pack and the highest-confidence result wins:

var result = await ocr.ExtractTextFromImage("street_sign.png", new[] { "en", "ch_sim", "ru" });

Each additional script family loads its own model, so request only the scripts you expect.

Right-to-left scripts

Arabic, Persian, Urdu, Uyghur and Hebrew pages are assembled right-to-left automatically — the right-most column is read first, and within a row the right-most fragment leads:

// Auto: every requested language is right-to-left, so the page is ordered right-to-left.
var result = await ocr.ExtractTextFromImage("invoice_ar.png", new[] { "ar" });

A mixed request stays left-to-right, since such a page is usually Latin-majority. Force the direction when a bilingual page really does flow right-to-left:

var result = await ocr.ExtractTextFromImage("form_ar_en.png", new[] { "ar", "en" }, new RecognitionOptions
{
    ReadingDirection = TextReadingDirection.RightToLeft,
});

DocumentAnalysisOptions.ReadingDirection does the same for AnalyzeDocumentAsync.

This orders boxes — which column and which line comes first. It is not the Unicode Bidirectional Algorithm, and does not need to be: each recognized line is already a logical-order string, so bidi stays your renderer's job and this output is correct input for one.

Automatic language detection

Don't know the language? Let the engine detect it — pass no codes and set AutoDetectLanguage:

var result = await ocr.ExtractTextFromImage("unknown.png", Array.Empty<string>(),
    new RecognitionOptions { AutoDetectLanguage = true });

// Or just detect, without recognizing:
IReadOnlyList<string> langs = await ocr.DetectLanguagesAsync("unknown.png");

Detection samples the largest text regions and scores candidate script packs by confidence. Candidates default to a common set (Latin, Cyrillic, Chinese, Japanese, Korean); widen them when you expect heavier scripts:

var opts = new RecognitionOptions
{
    AutoDetectLanguage = true,
    AutoDetectCandidates = new[] { "en", "ar", "hi", "ch_sim" },
};

Word & character geometry

Each line normally carries one polygon. Ask for WordLevelDetail and every line also reports its words and, optionally, its characters — with real boxes derived from the recognizer's CTC alignment (which timestep emitted which glyph), not estimated by splitting the line width.

var result = await ocr.ExtractTextFromImage("receipt.png", new[] { "en" }, new RecognitionOptions
{
    WordLevelDetail = WordLevelDetail.Words,   // or .Characters for per-glyph boxes
});

foreach (var word in result.Lines.SelectMany(l => l.Words))
    Console.WriteLine($"{word.Text,-20} {word.Confidence:P0}  {word.BoundingBox}");

Defaults to WordLevelDetail.None, so nothing changes — and costs nothing — unless you ask. Rotated lines produce genuinely rotated word quads, and the hOCR / ALTO / TSV exporters and the searchable-PDF text layer all sharpen automatically when word detail is present.

Streaming results

For multi-page documents or a responsive UI, take lines as they are recognized instead of waiting for the whole page:

await foreach (var line in ocr.ExtractTextStreamAsync("poster.png", new[] { "en" }))
    Console.WriteLine(line.Text);   // arrives as each region finishes

✍️ Handwriting (TrOCR)

Handwritten text needs a different model than printed text — EasyOCR's CRNN recognizers are trained on printed glyphs and cannot read cursive at all. Switch it on and call RecognizeHandwritingAsync:

var service = new EasyOcrService(new EasyOcrServiceOptions
{
    Handwriting = HandwritingOptions.Default,   // that's it
});

var notes = await service.RecognizeHandwritingAsync("handwritten-note.png");

The TrOCR models download into the same model cache as everything else on the first handwriting call, checksum-verified like the rest. Handwriting is null by default, so a service that doesn't ask for it never downloads a byte and behaves exactly as before.

Handwriting = new HandwritingOptions
{
    Quantize  = false,   // full precision (~1.5 GB) instead of the int8 default (~520 MB)
    BeamWidth = 4,       // 1 = greedy (default)
},

Quantize defaults to true: int8 weights are a third of the size and roughly twice as fast, at a small accuracy cost on unusual words. Full precision is the more accurate of the two — worth the extra download for archival work.

The hosted weights are an ONNX export of Microsoft's MIT-licensed trocr-base-handwritten, produced by tools/export_trocr_onnx.py.

Using your own export instead. Any standard Optimum TrOCR export works — an encoder taking pixel_values, a decoder taking input_ids + encoder_hidden_states (with or without past_key_values caching), and a byte-level BPE vocabulary — so you can swap in trocr-base-printed, a larger checkpoint, or your own fine-tune. Point the three paths at it and nothing is ever downloaded:

// Explicit paths…
Handwriting = new HandwritingOptions
{
    EncoderModelPath = "models/trocr/encoder_model.onnx",
    DecoderModelPath = "models/trocr/decoder_model.onnx",
    TokenizerPath    = "models/trocr/vocab.json",
};

// …or a folder holding those three conventional file names.
Handwriting = HandwritingOptions.FromDirectory("models/trocr"),

Setting only some paths is fine: whatever you leave null is fetched from the hosted set, so you can override just the decoder and keep the rest.

For long lines you can point DecoderModelPath at a decoder_model_merged.onnx — the runtime detects the past_key_values inputs and uses KV caching automatically. Do not use decoder_with_past_model.onnx on its own; it is the second half of a two-file pipeline and consumes caches it never produces.

Note the checkpoint is fine-tuned on handwriting: it reads printed text too, but the dedicated printed recognizers are better at that.

🔳 Barcodes & QR codes

Documents that need OCR usually carry codes too. Read them from the same image, with no OCR model involved:

foreach (var code in await BarcodeScanner.ReadBarcodesAsync("label.png",
             new BarcodeOptions { MultipleCodes = true }))
{
    Console.WriteLine($"{code.Format}: {code.Text}");
}

// …or both in one pass
var page = await ocr.ExtractTextAndBarcodesAsync("label.png", new[] { "en" });
Console.WriteLine($"{page.Ocr.Lines.Count} lines, {page.Barcodes.Count} codes");

BarcodeOptions covers Formats, TryHarder, MultipleCodes, TryInverted, AutoRotate and a Region restriction.

<br>

📄 Documents, scans & PDF

Scanned-document preprocessing

For photos and scans, enable clean-up via RecognitionOptions.Preprocessing:

var opts = new RecognitionOptions
{
    Preprocessing = new PreprocessingOptions
    {
        Deskew = true,            // straighten small tilt (±15°)
        DetectOrientation = true, // fix 90°/180°/270° rotation (≈4× cost)
        Binarize = true,          // adaptive black/white for uneven lighting
        Denoise = true,           // suppress scanner speckle
        Sharpen = true,           // unsharp-mask for soft / low-DPI scans (SharpenAmount tunes strength)
    },
};
var result = await ocr.ExtractTextFromImage("scan.jpg", new[] { "en" }, opts);

Two model-based document steps are also available — small dedicated neural models that download on first use (SHA256-verified like every other model):

var docOpts = new RecognitionOptions
{
    Preprocessing = new PreprocessingOptions
    {
        DocumentOrientation = true, // PP-LCNet doc classifier: fixes 90°/180°/270° in ONE tiny model
                                    // pass — much cheaper than DetectOrientation's 4× OCR
        DocumentUnwarp = true,      // UVDoc: dewarps curved/folded pages (photographed book pages,
                                    // creased receipts) before OCR
    },
};

Coordinate spaces. DetectOrientation reads the page at whichever 90°/180°/270° rotation scores best, then maps the boxes back onto your original image — results stay anchored to the input you passed and agree with the reported SourceWidth/SourceHeight. The model-based DocumentOrientation / DocumentUnwarp steps instead hand the whole pipeline a corrected page, so their boxes (and SourceWidth/SourceHeight) are in that corrected image's coordinate space.

🧱 Document structure & tables

Beyond plain text OCR, AnalyzeDocumentAsync recovers a page's structure — layout regions, tables (as HTML), formulas (as LaTeX), seals/stamps, and reading order — powered by PaddleOCR's PP-StructureV3 models, running on an engine built into this package (no extra dependency; the models download on first use):

using var ocr = new EasyOcrService();

var doc = await ocr.AnalyzeDocumentAsync("report_page.png");

foreach (var block in doc.Blocks)                    // in reading order
{
    Console.WriteLine($"{block.Order}: {block.Type} @ {block.Bounds}");
    if (block.TableHtml is not null)                 // tables come back as structured HTML
        Console.WriteLine(block.TableHtml);
}

string markdown = doc.ToMarkdown();                  // whole page as Markdown (tables included)
string json = doc.ToJson();

Tune what runs with DocumentAnalysisOptions (all models download on demand, SHA256-verified):

var doc = await ocr.AnalyzeDocumentAsync("scan.jpg", new DocumentAnalysisOptions
{
    DocumentOrientation = true,             // upright a rotated page first
    DocumentUnwarp = true,                  // dewarp a curved/folded page first
    RecognizeTables = true,                 // table structure as HTML (default on)
    RecognizeFormulas = false,              // skip LaTeX formula recognition
    RecognizeSeals = false,                 // skip seal/stamp recognition
    TableModel = DocumentTableModel.SlaNeXt,// higher-accuracy table model (default: SlanetPlus)
    Languages = new[] { "en" },             // text language(s); default pack covers ch/en/ja
});

The layout detector's own behaviour is tunable too. It emits a fixed top-k of candidate boxes with no NMS, so the same area of a page is routinely proposed several times under different labels; the duplicate filter is on by default and matters more the further you lower the confidence floor:

var doc = await ocr.AnalyzeDocumentAsync("scan.jpg", new DocumentAnalysisOptions
{
    LayoutScoreThreshold = 0.25f,           // confidence floor, default 0.5; lower keeps faint regions
    FilterOverlappingRegions = true,        // collapse duplicate/near-duplicate regions (default on)
    LayoutNms = true,                       // additionally run NMS over the regions (default off)
    LayoutUnclipRatio = 1.05f,              // grow each region 5% before recognition (default: none)
    LayoutMergeMode = DocumentLayoutMergeMode.Large,  // keep the enclosing block of a nested pair
    ReadingOrder = DocumentReadingOrder.XyCut,        // ignore the model's predicted order
});

ReadingOrder defaults to Auto: PP-DocLayoutV3 predicts a reading-order index alongside each box and that is what orders the blocks; models that emit no such column fall back to the geometric XY-cut orderer, which XyCut selects unconditionally.

The analyzer shares the service's execution provider, thread limits, cache path (when set) and download-resilience settings, loads lazily on first use, and is disposed with the service. Regular OCR calls never touch it.

📑 PDF input & searchable PDF

OCR scanned PDFs and emit searchable PDFs — built into the main package (no extra install needed). Pages are rasterized with PDFium and processed one at a time, so memory stays low even on large documents.

using EasyOcrSharp.Pdf;

await using var ocr = new EasyOcrService();

// 1) Extract text from every page:
PdfOcrResult doc = await ocr.ExtractTextFromPdfAsync("scan.pdf", new[] { "en" });
Console.WriteLine(doc.FullText);
foreach (var page in doc.Pages)
    Console.WriteLine($"Page {page.PageNumber}: {page.Ocr.Lines.Count} lines");

// 2) Produce a searchable PDF (original pages + invisible, selectable text layer):
await ocr.CreateSearchablePdfAsync("scan.pdf", "scan.searchable.pdf", new[] { "en" },
    pdfOptions: new PdfOcrOptions { Dpi = 250, JpegQuality = 80 });

PdfOcrOptions controls render Dpi, searchable-PDF JpegQuality, and a per-page Progress callback.

Unicode text layers

The invisible text layer uses the base-14 Helvetica font for Latin-1 text, and automatically switches to an embedded, subsetted Type0 / Identity-H font (with a ToUnicode CMap) when the recognized text needs it — so Chinese, Japanese, Korean, Arabic, Devanagari, Thai and Greek PDFs are genuinely searchable and copy-pasteable.

var (result, pdf) = await ocr.CreateSearchablePdfAsync(bytes, new[] { "ch_sim" }, pdfOptions: new()
{
    TextLayerFont     = PdfTextLayerFontMode.Auto,   // Auto | Never | Always
    TextLayerFontPath = "/usr/share/fonts/noto/NotoSansCJK-Regular.ttc",  // optional
});

if (result.TextLayerFontStatus == PdfTextLayerFontStatus.Unavailable)
    logger.LogWarning("No font covered this script — the text layer fell back to Latin-1.");

No font is bundled — a CJK font alone is tens of megabytes. Either point TextLayerFontPath at one, or let the built-in probe find an installed system font. If nothing suitable exists the output falls back to the old Helvetica layer rather than failing, and TextLayerFontStatus tells you so.

🖼️ Multi-frame TIFF

Scanners emit multi-page TIFFs. Read every frame, not just the first:

var doc = await ocr.ExtractTextFromFramesAsync("scan.tif", new[] { "en" });
Console.WriteLine($"{doc.Frames.Count} frames in {doc.Duration.TotalSeconds:0.0}s");

// or stream them, so a 200-page TIFF starts producing results immediately
await foreach (var frame in ocr.StreamTextFromFramesAsync("scan.tif", new[] { "en" }))
    Console.WriteLine($"page {frame.FrameIndex}: {frame.Ocr.FullText}");

The pixel-flood guard applies per frame, MaxFrames bounds the document, and a single-frame image flows through the same call and returns exactly one result.

📊 Tables as data

AnalyzeDocumentAsync recovers tables as HTML. Turn them into something .NET can use:

var structure = await ocr.AnalyzeDocumentAsync("invoice.png");

foreach (var table in structure.Tables())
{
    IReadOnlyList<IReadOnlyList<string>> rows = table.ToRows();
    DataTable dt = table.ToDataTable();      // header row detected from <th>
    string csv  = table.ToCsv();             // RFC 4180 quoting
}

Merged cells are expanded into repeated values, HTML entities are decoded, and malformed markup is tolerated rather than thrown on.

<br>

🔌 Output & integration

Output formats (hOCR / ALTO / TSV / JSON)

Any OcrResult converts to the interchange formats DMS and archival pipelines expect:

using EasyOcrSharp.Export;

using var img = Image.Load<Rgb24>("page.png");
var result = await ocr.ExtractTextFromImage(img, new[] { "en" });

string hocr = result.ToHocr(pageWidth: img.Width, pageHeight: img.Height); // hOCR (HTML)
string alto = result.ToAlto(pageWidth: img.Width, pageHeight: img.Height); // ALTO XML v4
string tsv  = result.ToTsv();                                              // Tesseract-style TSV
string json = result.ToJson(indented: true);                              // AOT-safe JSON

ToJson uses a source-generated EasyOcrJsonContext, so it works in trimmed / Native-AOT apps with no reflection warnings. Recognized text is written verbatim in every script — Cyrillic, Greek, Arabic, CJK and the rest stay readable in a plain text editor instead of turning into \uXXXX escapes — while HTML-sensitive characters (< > & ' +) are still escaped, since a block's TableHtml may end up embedded in a page. Pass your own JsonSerializerOptions for a different encoder, indentation or naming policy (StructureResult.ToJson takes the same overload), and the exporter stays reflection-free:

string strict = result.ToJson(new JsonSerializerOptions
{
    Encoder = JavaScriptEncoder.Default,   // back to \uXXXX escaping for every non-ASCII character
    WriteIndented = true,
});

Recognize from known boxes

If you already have regions — from DetectRegionsAsync, a previous run, or your own layout analysis — recognize them directly and skip detection (EasyOCR's recognize()):

using var image = Image.Load<Rgb24>("form.png");

// e.g. reuse a detection pass, or pass your own polygons (pixel coordinates):
IReadOnlyList<DetectedRegion> regions = await ocr.DetectRegionsAsync(image);

OcrResult result = await ocr.RecognizeRegionsAsync(image, regions, new[] { "en" });

There's also an overload taking raw polygons (IEnumerable<IReadOnlyList<OcrPoint>>).

Detection-only & visualization

Locate text regions without recognizing them — fast and language-independent, ideal for layout analysis, redaction, or cropping fields before a targeted recognition pass:

IReadOnlyList<DetectedRegion> regions = await ocr.DetectRegionsAsync("form.png");

Draw the boxes onto a copy of the image for debugging (no extra dependency; original is untouched):

using EasyOcrSharp.Export;

using var img = Image.Load<Rgb24>("page.png");
var result = await ocr.ExtractTextFromImage(img, new[] { "en" });
using var annotated = img.DrawAnnotations(result, new Rgb24(255, 0, 0), thickness: 2);
await annotated.SaveAsync("page.annotated.png");

Batch processing

Process a folder or queue with bounded concurrency. Results stream as they complete; a failed image is captured (not thrown), so one bad file never aborts the batch:

var files = Directory.EnumerateFiles("inbox", "*.png");

await foreach (var item in ocr.ExtractTextFromImagesAsync(files, new[] { "en" }, maxConcurrency: 4))
{
    if (item.Succeeded) Console.WriteLine($"{item.Source}: {item.Result!.Lines.Count} lines");
    else                Console.Error.WriteLine($"{item.Source} failed: {item.Error!.Message}");
}

🖍️ Redaction

Find sensitive text and permanently destroy those pixels — the region is painted over, not covered by an annotation someone can remove:

var redacted = await ocr.RedactAsync("statement.png", new[] { "en" }, new RedactionOptions
{
    Rules    = RedactionPatterns.Common,     // email, phone, card, IBAN, SSN, long digit runs
    Keywords = new[] { "Account Holder" },
    Style    = RedactionStyle.FilledBox,     // or Blur / Pixelate
    Scope    = RedactionScope.MatchedWords,  // only the matched words, not the whole line
});

await redacted.Image.SaveAsync("statement.redacted.png");
Console.WriteLine($"{redacted.RedactedRegionCount} regions removed");
Console.WriteLine(redacted.SanitizedText);   // the text with matches masked out

// PDFs too
var safe = await ocr.RedactPdfAsync(pdfBytes, new[] { "en" }, options);

The card and IBAN presets are validated, not just matched: CreditCard applies a Luhn check and Iban a mod-97 check, so a random 16-digit order number is not mistaken for a card number.

🧾 Field extraction

Pull structured values out of invoices, receipts and forms using the label positions OCR already produced — no LLM involved:

var result = await ocr.ExtractTextFromImage("invoice.png", new[] { "en" });

var fields = result.ExtractFields(new[]
{
    FieldPresets.InvoiceNumber,
    FieldPresets.InvoiceDate,
    FieldPresets.Total,
    new FieldDefinition
    {
        Name      = "Customer PO",
        Anchors   = new[] { "Customer PO", "PO Number" },
        Direction = FieldDirection.Right | FieldDirection.Below,
    },
});

foreach (var f in fields)
    Console.WriteLine($"{f.Name}: {f.Value}  ({f.Confidence:P0})");

Anchor matching is fuzzy, so OCR damage like Totai still resolves; distances are expressed as multiples of the anchor's line height, so the same definition works at any resolution; and a plain "Total: 42.00" on one line is extracted without any geometry at all.

<br>

🎯 Accuracy & tuning

Constrained fields (allow/block-lists & detection thresholds)

For fixed-format fields, restrict the character set — this sharply cuts errors:

// Digits only (invoice totals, IDs, meter readings):
var digits = new RecognitionOptions { Allowlist = "0123456789.," };

// License plate (upper-case + digits):
var plate = new RecognitionOptions { Allowlist = "ABCDEFGHIJKLMNOPQRSTUVWXYZ0123456789-" };

var total = await ocr.ExtractTextFromImage("receipt.png", new[] { "en" }, digits);

Use Blocklist to forbid specific characters instead. For hard inputs, the CRAFT detector thresholds are exposed via RecognitionOptions.Detection (defaults match EasyOCR):

var opts = new RecognitionOptions
{
    Detection = new DetectionOptions
    {
        TextThreshold = 0.6,  // lower → catch fainter text
        LowText = 0.3,        // lower → keep more of each glyph
        MagRatio = 1.5,       // enlarge before detection (small text)
    },
};

Decoders, rotation & batching

Switch the CTC decoder, recognize rotated text, or batch boxes through the model — all via RecognitionOptions (every option defaults to the previous behaviour):

var opts = new RecognitionOptions
{
    Decoder = DecoderType.BeamSearch,  // Greedy (default) | BeamSearch | WordBeamSearch
    BeamWidth = 10,                    // explored hypotheses (beam decoders)
    RotationInfo = new[] { 90, 270 },  // also try each box rotated; keep the best reading
    BatchSize = 16,                    // batch boxes through one ONNX run (see note below)
};

var result = await ocr.ExtractTextFromImage("rotated_labels.png", new[] { "en" }, opts);

BatchSize note. Batching needs a batch-capable recognizer export; the currently hosted models are exported with batch fixed at 1, so BatchSize > 1 transparently falls back to per-box inference today (results are unaffected). Per-box recognition already runs concurrently — tune throughput with MaxDegreeOfParallelism.

WordBeamSearch constrains output to a lexicon you supply, which is powerful for closed vocabularies (part numbers, place names, a product catalogue):

var opts = new RecognitionOptions
{
    Decoder = DecoderType.WordBeamSearch,
    Dictionary = new[] { "INVOICE", "TOTAL", "SUBTOTAL", "TAX" },
};

Custom recognizers

Register your own exported CRNN ONNX model (e.g. a fine-tuned EasyOCR recog_network) for chosen language codes. A custom recognizer takes precedence over the built-in pack and is loaded straight from disk — never downloaded:

var options = new EasyOcrServiceOptions();
options.CustomRecognizers.Add(new CustomRecognizer
{
    Name = "my_meter_reader",
    ModelPath = @"D:\models\meter_g2.onnx",
    VocabPath = @"D:\models\meter_g2.vocab.json", // or set Characters = "0123456789." inline
    Languages = new[] { "en" },                    // claim the codes it should handle
});

await using var ocr = new EasyOcrService(options);

Fine-tuning grouping & contrast

The thresholds that merge boxes into lines/paragraphs and trigger the contrast-retry pass are exposed for difficult layouts (defaults reproduce EasyOCR's behaviour):

var opts = new RecognitionOptions
{
    GroupingOptions = new GroupingOptions
    {
        SlopeThreshold = 0.1,        // tolerate gently tilted lines (slope_ths)
        YCenterThreshold = 0.5,      // vertical tolerance for same-line boxes (ycenter_ths)
        WidthThreshold = 1.0,        // max horizontal gap to merge on a line (width_ths)
        ParagraphYThreshold = 1.0,   // vertical reach when forming paragraphs (y_ths)
    },
    ContrastThreshold = 0.1,         // re-recognize below this confidence (contrast_ths)
    AdjustContrastTarget = 0.5,      // grey-stretch target for the retry pass (adjust_contrast)
};

Quantized models

Set Quantize to fetch the int8-quantized recognizers instead of the float ones — EasyOCR's quantize=True, for smaller downloads:

await using var ocr = new EasyOcrService(new EasyOcrServiceOptions { Quantize = true });

The int8 variants are hosted alongside the float models and SHA256-verified on download, and text output is effectively unchanged. The win is vocabulary-dependent: ONNX Runtime (CPU) int8-quantizes the matmul/linear layers but not the BiLSTM/convolutions, so large-vocabulary packs shrink most (e.g. zh_sim ~22 → ~16 MB) while small-vocabulary packs change little. The detector stays float (as in EasyOCR). Opt-in; the float models are the default.

Post-OCR correction

Fix the recognizer's mistakes with a domain lexicon — while leaving text it was confident about completely alone:

var corrected = result.Correct(new CorrectionOptions
{
    Dictionary             = File.ReadLines("part-numbers.txt").ToArray(),
    MaxEditDistance        = 2,
    MinConfidenceToCorrect = 0.85,   // only touch tokens the model itself flagged as shaky
    Normalizers            = new[] { FieldNormalizers.Iban(), FieldNormalizers.Date() },
});

Candidate ranking is weighted by the confusions OCR actually makes (0/O, 1/l/I, 5/S, 8/B, rn/m). The normalizers go further than validation: where a checksum identifies the wrong character — IBAN mod-97, ICAO 9303 MRZ check digits — they repair it. Correct never mutates the input; it returns a new OcrResult.

Measuring accuracy (CER / WER)

double cer = result.CharacterErrorRate(expectedText);
double wer = result.WordErrorRate(expectedText);

var report = result.Compare(expectedText, TextComparisonOptions.Relaxed);
Console.WriteLine($"CER {report.CharacterErrorRate:P2} — " +
                  $"{report.Characters.Substitutions} sub, " +
                  $"{report.Characters.Insertions} ins, " +
                  $"{report.Characters.Deletions} del");

Useful for benchmarking a preprocessing change or a decoder setting against your own documents rather than trusting a generic accuracy claim.

<br>

📊 Production & operations

GPU & execution providers

GPU is automatic. ExecutionProvider defaults to Auto: on the first run EasyOcrSharp asks ONNX Runtime what accelerators are actually present and uses the best one, falling back to CPU when there's none. You don't pick a provider — you just install the package for the hardware you have:

Install this package What Auto enables
EasyOcrSharp (base) CPU only
EasyOcrSharp.Gpu NVIDIA CUDA (needs CUDA 12+ on PATH)

Why a separate package? The ONNX Runtime variants ship the same onnxruntime.dll compiled with different providers, so only one can be referenced at a time — and the CUDA build is several hundred MB. Shipping it as an opt-in package keeps the base library small and cross-platform; Auto then lights up whatever you installed.

// Nothing to configure — add the EasyOcrSharp.Gpu package and a GPU is used if present.
await using var ocr = new EasyOcrService();

You can still pin a provider explicitly (e.g. to force CPU, or to require CUDA) and set ONNX Runtime thread limits:

await using var ocr = new EasyOcrService(new EasyOcrServiceOptions
{
    ExecutionProvider = OcrExecutionProvider.Cuda, // Auto (default) | Cpu | Cuda | CoreMl
    IntraOpNumThreads = 4,   // cap CPU use in multi-tenant servers (null = runtime default)
});
Provider Package Notes
Auto (any) Default. Probes the runtime; uses the best installed accelerator, else CPU
Cpu (built-in) Always available
Cuda EasyOcrSharp.Gpu NVIDIA, CUDA 12+ on PATH
CoreMl CoreML-enabled ORT build macOS / Apple Silicon

Any non-CPU provider falls back to CPU automatically (with a logged warning) if its runtime is missing or the device fails to initialize — your app keeps working. The legacy useGpu: true flag still works and forces CUDA. Check ocr.UseGpu to see whether an accelerator was selected.

GPU upgrade hint. When Auto runs on CPU but a real NVIDIA GPU is physically present, EasyOcrSharp detects it and can tell you the exact package to add — EasyOcrSharp.Gpu. It's silent by default: the hint is exposed as a property you can surface yourself, and nothing is logged unless you opt in with LogGpuHint = true.

// Silent by default — read it only if you want to nudge the user yourself:
await using var ocr = new EasyOcrService();
if (ocr.GpuAccelerationHint is { } hint) Console.WriteLine(hint);
// e.g. "EasyOcrSharp: an NVIDIA GPU was detected but OCR is running on CPU. Install the
//       'EasyOcrSharp.Gpu' NuGet package for CUDA acceleration. ..."

// Opt in to a one-time startup warning in the logs instead:
await using var verbose = new EasyOcrService(new EasyOcrServiceOptions { LogGpuHint = true });

Observability & health checks

EasyOcrSharp emits OpenTelemetry-ready metrics and traces, always-on with near-zero cost when nobody is listening:

builder.Services.AddOpenTelemetry()
    .WithMetrics(m => m.AddMeter(EasyOcrDiagnostics.MeterName))
    .WithTracing(t => t.AddSource(EasyOcrDiagnostics.ActivitySourceName));
Instrument What it tells you
easyocr.operations operation count — including failures, so an erroring deployment shows a rising error rate rather than going quiet
easyocr.duration wall-clock latency per operation (ms)
easyocr.lines / easyocr.pages throughput: lines recognized, pages/frames processed
easyocr.operations.active / .queued live saturation — how many are executing vs waiting for a slot
easyocr.queue.wait how long operations wait for a slot before running or being shed
easyocr.model.loads / .download_bytes ONNX sessions created, model bytes fetched

Every operation instrument is tagged, so a dashboard can slice latency and error rate by easyocr.operation (extract, extract_stream, recognize, detect, handwriting, analyze_document, pdf, multi_frame, …), easyocr.outcome (success, error, canceled, timeout, shed), easyocr.languages, easyocr.provider (Cpu / Cuda / …) and — on failures — error.type. The constants are public (EasyOcrDiagnostics.TagNames, .Outcomes, .OperationNames), so alert rules can reference them instead of hard-coded strings:

// error rate, excluding client cancellations
sum(rate(easyocr_operations_total{easyocr_outcome=~"error|timeout"}[5m]))
  / sum(rate(easyocr_operations_total[5m]))

Add a readiness probe that reports whether the models for your languages are cached (so the first real request won't block on a download):

builder.Services.AddHealthChecks()
    .AddEasyOcrHealthCheck(languages: new[] { "en" });

That check is a file-presence check, which cannot catch a model that is present but broken — a truncated download, a half-copied cache, or a GPU whose provider fails at session init. All three report Healthy and then fail every request. Opt into a deep probe that actually runs a tiny synthetic page through the real pipeline once, caches the verdict, and reports the provider that genuinely resolved:

builder.Services.AddEasyOcrWarmUp("en");        // IHostedService: loads models off the startup path

builder.Services.AddHealthChecks()
    .AddEasyOcrHealthCheck(languages: ["en"], name: "easyocr-live")                    // shallow: liveness
    .AddEasyOcrHealthCheck(new EasyOcrHealthCheckOptions { DeepProbe = true },
                           languages: ["en"], name: "easyocr-ready");                  // deep: readiness

The probe never triggers a download — if the models aren't cached it reports that instead, so Offline deployments behave unchanged. When AddEasyOcrWarmUp is registered, readiness stays not ready until warm-up finishes, so an orchestrator won't route traffic to a pod that is still loading. Warm-up failure is never fatal: it is logged and surfaced through the check, and models still load lazily on first use.

Dependency injection

services.AddEasyOcrSharp(o =>
{
    o.ModelCachePath = "/var/cache/easyocr";
    // GPU is automatic (ExecutionProvider = Auto). To force CPU in a multi-tenant host:
    // o.ExecutionProvider = OcrExecutionProvider.Cpu;
});

public class ReceiptParser(IEasyOcrService ocr) { /* inject anywhere */ }

Registered as a singleton — ONNX sessions are expensive to build and thread-safe to reuse.

Hardening & resource limits

When OCR-ing untrusted images or PDFs, EasyOcrSharp guards against decompression-bomb / pixel-flood denial of service. The defaults are generous; raise them if you legitimately process larger inputs, or set them to 0 to disable a guard.

await using var ocr = new EasyOcrService(new EasyOcrServiceOptions
{
    MaxImagePixels = 100_000_000,   // reject images over 100 MP from the header, before decode (default)
});

var pdfOptions = new PdfOcrOptions
{
    MaxPages = 5000,                // reject documents with more pages (default)
    MaxPageMegapixels = 200,        // reject a page that would rasterize larger at the chosen DPI (default)
};

Failures surface as typed exceptions (all derive EasyOcrSharpException, so a catch-all still works):

Exception When
ImageTooLargeException image exceeds MaxImagePixels
PdfProcessingException corrupt / encrypted PDF, or a page/size guard tripped
ModelDownloadException download failed, or a non-HTTPS / malformed model source
ModelChecksumException downloaded model failed (or lacks) SHA256 verification
OfflineModelMissingException model not cached and Offline = true
OcrBusyException at MaxConcurrentOperations and no slot freed within QueueTimeout
OcrTimeoutException a single operation exceeded OperationTimeout

Concurrency, back-pressure & operation timeouts

ONNX sessions are thread-safe to share, but every concurrent run allocates its own tensors — so concurrency, not session count, is what sets peak memory. An ungoverned service handed fifty 12 MP scans at once doesn't get slower in a straight line: every request crawls, all of them eventually time out, and the working set grows until the container is OOM-killed. Both guards are off by default, so nothing changes until you opt in:

await using var ocr = new EasyOcrService(new EasyOcrServiceOptions
{
    MaxConcurrentOperations = Environment.ProcessorCount,  // 0 (default) = unlimited
    QueueTimeout            = TimeSpan.FromSeconds(30),    // then shed with OcrBusyException
    OperationTimeout        = TimeSpan.FromMinutes(2),     // 0 (default) = no cap
});

QueueTimeout defaults to Timeout.InfiniteTimeSpan — a concurrency limit on its own bounds memory but still queues everything, so shedding is a second, deliberate opt-in. Set it to enable back-pressure; TimeSpan.Zero refuses the moment the limit is saturated. Waits longer than ~49.7 days (the platform's timer ceiling) are treated as "wait forever" rather than throwing.

Once it is set, operations past the limit wait up to QueueTimeout for a slot and are then refused promptly rather than queued without bound — a caller told "busy" in 30 seconds can retry or fail over, one silently queued for four minutes cannot. Map it straight onto HTTP:

catch (OcrBusyException ex)    { return Results.Json(..., statusCode: 503); }  // + Retry-After
catch (OcrTimeoutException ex) { return Results.Json(..., statusCode: 504); }

OperationTimeout exists because a CancellationToken only helps when something signals it — a pathological page (thousands of detected boxes, an image that defeats the detector) otherwise occupies a worker forever, and a handful of those take a fixed-size pool down. The cap is cooperative: it cancels the pipeline at its next checkpoint rather than aborting a thread, so set it comfortably above your p99 page time. Note a multi-page PDF is one operation, not one per page.

OcrTimeoutException is deliberately distinct from OperationCanceledException: the first means this input was too slow and should be quarantined, the second means your own caller hung up. Only the primitive per-image operations take a slot; PDF, multi-frame and batch runs are gated per page as they go, so a long document can't hold a slot hostage for its whole duration.

Warm-up (remove cold-start latency). Preload the detector and recognizer packs so the first real request doesn't pay model-download + session-init latency — ideal for serverless / scale-out:

await ocr.WarmUp(new[] { "en" });   // downloads + initializes once, up front

Resilient & offline model downloads

Model downloads are production-hardened: atomic, SHA256-verified, resumable (HTTP range), and retried with exponential backoff. By default the model source must be HTTPS and every model must have a known checksum. Tune everything via ModelDownloadOptions:

await using var ocr = new EasyOcrService(new EasyOcrServiceOptions
{
    Download = new ModelDownloadOptions
    {
        MaxRetries = 5,
        Offline = false,                              // true = never download; fail fast if not cached
        BaseUrlOverride = "https://mirror.corp/ocr",  // private mirror (must be https unless opted out)
        HttpClientFactory = () => httpClientFactory.CreateClient("ocr"), // proxy / corporate certs
        // AllowInsecureModelSource = true,           // permit a plain-http mirror you control
        // AllowUnverifiedModels   = true,            // permit unlisted models that have no registry checksum
        Progress = new Progress<ModelDownloadProgress>(p =>
            Console.WriteLine($"{p.FileName}: {p.Fraction:P0}")),
    },
});

For air-gapped deployments, pre-seed the cache and set Offline = true — a missing model then throws a clear error instead of attempting a download.

🖥️ Command-line tool

dotnet tool install -g EasyOcrSharp.Cli
# recognize an image, a folder, or a glob
easyocrsharp scan receipt.png
easyocrsharp scan scans/ -r -l en,de --format json -o out/

# make a scanned PDF searchable
easyocrsharp pdf scan.pdf -o searchable.pdf

# pre-download models for an air-gapped host, then check what's there
easyocrsharp models pull en,fr
easyocrsharp models list
easyocrsharp models path

# versions, active execution provider, GPU status, cache location
easyocrsharp info

scan writes results to stdout and errors to stderr so it pipes cleanly, exits non-zero on failure, and accepts the same tuning as the library (--allowlist, --min-confidence, --paragraph, --preprocess deskew,binarize,sharpen, --gpu, --jobs). Add --help to any command.

🌐 Web service sample

samples/EasyOcrSharp.WebApi is a runnable ASP.NET Core service — POST /ocr, POST /ocr/pdf, GET /health and a browser upload page — with bounded concurrency, upload limits, problem-details errors and a Dockerfile that already includes the native prerequisites PDFium and ONNX Runtime need.

dotnet run --project samples/EasyOcrSharp.WebApi
curl -X POST "http://localhost:5000/ocr?lang=en&format=text" -F "file=@receipt.png"

<br>

📚 Reference

🌍 Supported languages

Languages are grouped by script; one recognizer covers an entire group, so ["en","es","fr"] loads a single model. Pack sizes vary widely with the network each script was trained on — some are a few MB, some ~210 MB — which affects first-run download size only, not runtime behaviour.

Pack Size Languages
latin_g2 ~15 MB af, az, bs, cs, cy, da, de, en, es, et, fr, ga, hr, hu, id, is, it, ku, la, lt, lv, mi, ms, mt, nl, no, oc, pi, pl, pt, ro, rs_latin, sk, sl, sq, sv, sw, tl, tr, uz, vi
cyrillic_g2 ~15 MB ru, rs_cyrillic, be, bg, uk, mn, abq, ady, kbd, ava, dar, inh, che, lbe, lez, tab, tjk
zh_sim_g2 ~22 MB ch_sim
korean_g2 ~16 MB ko
japanese_g2 ~17 MB ja
telugu_g2 ~15 MB te
kannada_g2 ~15 MB kn
arabic_g2 ~210 MB ar, fa, ug, ur
devanagari_g2 ~210 MB hi, mr, ne, bh, mai, ang, bho, mah, sck, new, gom, sa, bgc
bengali_g2 ~210 MB bn, as, mni
thai_g1 ~210 MB th
tamil_g1 ~210 MB ta
zh_tra_g1 ~215 MB ch_tra

That's all 86 languages EasyOCR supports, mapped exactly to the model each was trained on.

Not supported: Greek (el) and Hebrew (he) — upstream EasyOCR ships no model for either script, so they cannot be exported.

How model downloads work

EasyOcrSharp ships no models in the NuGet package. On the first call for a language it downloads, into a local cache:

  1. CRAFT detector (craft_mlt_25k.onnx, ~80 MB) — shared by all languages, downloaded once.
  2. CRNN recognizer for the language's script pack (e.g. latin_g2.onnx).
  3. A small vocabulary sidecar (<pack>.vocab.json).

Every file is SHA256-verified against a checksum baked into the library, so corrupted or tampered downloads are rejected. Models are hosted on Hugging Face.

Default cache: %LOCALAPPDATA%\EasyOcrSharp\models (Windows) or the platform equivalent. Override it:

await using var ocr = new EasyOcrService(modelCachePath: @"D:\MyApp\Models");
EASYOCRSHARP_CACHE=/var/cache/easyocr                          # cache directory
EASYOCRSHARP_MODEL_BASE_URL=https://files.mycorp.example/ocr   # private/offline mirror

# AnalyzeDocumentAsync's structure models are cached in the same directory but resolved
# through their own pair, so a mirror can serve them separately:
EASYOCRSHARP_STRUCTURE_CACHE=/var/cache/easyocr
EASYOCRSHARP_STRUCTURE_MODEL_BASE_URL=https://files.mycorp.example/ocr

Offline / air-gapped: pre-seed your cache directory with the .onnx + .vocab.json files from the model repo — no network is needed at runtime.

Accuracy notes

EasyOcrSharp reproduces EasyOCR's pipeline faithfully — aspect-preserving resize, normalization, a low-confidence contrast-retry pass, CRAFT box dilation, perspective de-warping of rotated text, and CTC decoding (greedy by default, with optional beam / word-beam search) — so output matches upstream EasyOCR. On top of that:

  • Reading order is column-aware and bands rows by a line-height-relative tolerance, so headings, high-DPI scans, and multi-column pages come out in natural reading order.
  • Overlapping detections are de-duplicated with IoU NMS (DetectionOptions.NmsIouThreshold, default 0.6; set 0 to disable).
  • On multi-language requests, scoring is biased toward the page's dominant script so an over-confident wrong-script pack can't hijack individual boxes.

As with any OCR:

  • Visually identical glyphs (capital I vs lowercase l, $ vs 8) can be confused.
  • Handwriting and low-resolution / low-contrast text are harder than clean printed text.
  • Right-to-left scripts (Arabic) are returned in the model's character order.

Building & testing

git clone https://github.com/FarhanLodi/EasyOcrSharp.git
cd EasyOcrSharp
dotnet build -c Release

# Everything — unit + real end-to-end integration tests. Downloads the models on first run
# and reports a pass/fail summary:
dotnet test

# Fast unit tests only (no models, no network):
dotnet test --filter "Category!=Integration"

# Only the model-backed integration tests:
dotnet test --filter "Category=Integration"

# Interactive console demo:
dotnet run --project test/EasyOcrSharp.Demo

The integration tests exercise every feature against the real engine (allow/block-lists, detection-only, exporters, batch, metrics/tracing, health check, execution-provider fallback, and the full PDF pipeline) — no mocks. The PDF fixtures live in test/assets/pdf/ and are committed, so those tests run out of the box. A couple of tests still skip (never fail) until you supply an optional fixture — a password-protected PDF named e.g. encrypted_secret.pdf for the encrypted-document path, and EASYOCRSHARP_TROCR_DIR pointing at a TrOCR export for the handwriting integration test.

Path Purpose
src/EasyOcrSharp the core library (includes PDF input + searchable-PDF output)
src/EasyOcrSharp.Gpu CUDA execution-provider package
src/EasyOcrSharp.Cli the easyocrsharp command-line tool (dotnet tool)
samples/EasyOcrSharp.WebApi ASP.NET Core service sample + Dockerfile
test/EasyOcrSharp.Tests xUnit unit + integration tests
test/EasyOcrSharp.Demo interactive console demo
test/assets sample images
tools/ maintainer-only ONNX export + quantization scripts

CI (GitHub Actions) builds and runs the unit tests on Linux and Windows for every push and PR. See CHANGELOG.md for release history.

<br>

🤝 Contributing

Contributions are welcome! New features, accuracy improvements, performance tuning, bug fixes, additional language/model coverage, documentation, and tests are all appreciated.

  • 🐛 Found a bug? Open an issue with a minimal repro (image/PDF + the code and options you used).
  • 💡 Have an idea or feature request? Open an issue to discuss it first, then send a PR.
  • 🔧 Sending a PR? Branch from main, keep changes focused, and make sure dotnet build -c Release and the unit tests (dotnet test --filter "Category!=Integration") pass.

If you're working on something larger, or want to collaborate on a feature, feel free to reach out before starting so we can align on the approach.

💖 Support

If EasyOcrSharp saves you time, consider supporting development:

  • 💳 PayPalpaypal.me/FarhanLodi
  • 📱 UPI (India)farhanlodi5@oksbi
  • 🏦 Bank transfer (USD) — details below

<details> <summary><b>USD bank transfer details (Wise)</b></summary>

<br>

USD account details for Farhan Lodi on Wise. Sending from a bank in the US? Use these details for a domestic transfer. Sending from anywhere else? Make an international SWIFT transfer.

Field Value
Name Farhan Lodi
Account type Deposit
Routing number (wire and ACH) 084009519
Account number 420927686563885
SWIFT/BIC TRWIUS35XXX
Bank address Wise US Inc, 108 W 13th St, Wilmington, DE, 19801, United States

Use the routing and account numbers when sending from the US, and the SWIFT/BIC when sending from outside the US.

</details>

📧 Need more details, a different payment method, or have a question? Email farhanlodi31@gmail.com.

📬 Contact

For work inquiries, collaboration, feature requests, or any questions, reach out to:

Farhan Lodifarhanlodi31@gmail.com

📄 License

MIT — see LICENSE. The code has no copyleft or commercially-tiered dependency at any depth. The neural weights downloaded at runtime keep their own upstream licences — EasyOCR and PaddleOCR models are Apache-2.0, TrOCR is MIT — and are attributed in NOTICE.

🙏 Acknowledgments

  • EasyOCR — the underlying CRAFT + CRNN models
  • ONNX Runtime — neural network execution
  • EasyImageSharp — image decoding, encoding and the processing pipeline
  • PaddleOCR — the PP-StructureV3 models behind AnalyzeDocumentAsync

<div align="center"> <br>

⬆ Back to top

<sub>Built with ❤️ for the .NET community · EasyOCR accuracy, zero Python</sub>

</div>

Product Compatible and additional computed target framework versions.
.NET net10.0 is compatible.  net10.0-android was computed.  net10.0-browser was computed.  net10.0-ios was computed.  net10.0-maccatalyst was computed.  net10.0-macos was computed.  net10.0-tvos was computed.  net10.0-windows was computed. 
Compatible target framework(s)
Included target framework(s) (in package)
Learn more about Target Frameworks and .NET Standard.

NuGet packages (1)

Showing the top 1 NuGet packages that depend on EasyOcrSharp:

Package Downloads
EasyOcrSharp.Gpu

CUDA GPU execution provider for EasyOcrSharp. Adds Microsoft.ML.OnnxRuntime.Gpu so detection and recognition run on NVIDIA GPUs.

GitHub repositories

This package is not used by any popular GitHub repositories.

Version Downloads Last Updated
3.1.0 92 8/30/2026
3.0.0 162 8/27/2026
2.3.2 141 8/16/2026
2.3.1 102 8/16/2026
2.3.0 126 8/16/2026
2.2.4 278 7/3/2026
2.2.3 162 6/19/2026
2.2.2 150 6/19/2026
2.2.1 171 6/6/2026
2.2.0 132 6/6/2026
2.1.1 163 5/30/2026
2.1.0 154 5/30/2026
2.0.1 143 5/30/2026
2.0.0 153 5/30/2026

3.1.0: production-operations pass - all additive, every new guard off by default, nothing renamed or removed. OBSERVABILITY: easyocr.operations, .duration, .lines and .pages are now tagged with easyocr.operation, easyocr.outcome, easyocr.provider, easyocr.languages and error.type, so latency and error rate can be sliced by entry point, language and CPU-vs-GPU. Crucially they are recorded on FAILURE too: metrics were previously emitted only on the success path, so a deployment failing every request reported ZERO operations rather than a 100 percent error rate - the one shape of failure a dashboard cannot see. Timeouts and shed load get their own outcomes rather than folding into error, so a burst the concurrency limit handled as designed does not fire the error-budget alert. New saturation instruments easyocr.queue.wait, easyocr.operations.active and easyocr.operations.queued, plus metrics and spans on the previously uninstrumented PDF, searchable-PDF, multi-frame and batch paths. Tag, outcome and operation constants are public as EasyOcrDiagnostics.TagNames, .Outcomes and .OperationNames so alert rules need no magic strings. BACK-PRESSURE: new EasyOcrServiceOptions.MaxConcurrentOperations, QueueTimeout and OperationTimeout, with typed OcrBusyException (503-shaped: at the limit and the queue wait elapsed) and OcrTimeoutException (504-shaped: one operation blew its budget), both deriving EasyOcrSharpException so existing catch-all handlers keep working. Concurrency, not session count, is what sets peak memory, because every concurrent run allocates its own tensors. The gate is taken once per outermost operation; PDF, multi-frame and batch runs gate per page instead, so a long document cannot hold a slot for its whole duration. STARTUP AND READINESS: AddEasyOcrWarmUp(languages) registers an IHostedService that loads models off the startup path (a host that cannot reach the model mirror still starts and reports state instead of crash-looping) and publishes progress through EasyOcrWarmUpState; EasyOcrHealthCheckOptions with DeepProbe, ProbeInterval and ProbeTimeout adds an opt-in health check that actually runs a tiny synthetic page through the real pipeline, caches the verdict, and reports the execution provider that resolved - a truncated model file, a half-copied cache or a GPU whose provider fails at session init previously reported Healthy and then failed every request. The probe never triggers a download, so offline deployments are unchanged, and readiness stays not-ready while warm-up is still running. FIXED: EasyOcrDiagnostics reported version 2.2.1 while the package was 3.x, mislabelling every metric and span it emitted. See CHANGELOG.md. 3.0.1: three bugs reported against PaddleOcrNet, whose PP-StructureV3 engine 3.0.0 brought in-tree, apply to the copy of that engine here and are fixed. (1) Every document-analysis language pack decoded one character class off: the per-script PP-OCRv5 dictionaries (cyrillic, latin, arabic, devanagari, korean, japan, th, el, te, ta, eslav) open with an empty line that IS the CTC blank, so the class they omit is the trailing space; CharacterDictionary prepended a second blank and shifted every class by one, turning Russian into mojibake. BuildVocab now detects the empty first line. The dictionary files were always correct, so nothing re-downloads; the default ch/en/ja recognizers were never affected. (2) OcrResult.ToJson, StructureResult.ToJson and the CLI's --format json escaped everything outside Basic Latin into \uXXXX; they now serialize through EasyOcrJson.Encoder and write Cyrillic, Greek, Arabic, CJK and the rest verbatim, while still escaping HTML-sensitive characters. New ToJson(JsonSerializerOptions) overloads on both types let callers pick their own encoder, indentation or naming policy without giving up trim/AOT safety. (3) The layout confidence floor was a private const; it is now DocumentAnalysisOptions.LayoutScoreThreshold, default unchanged at 0.5 and still exclusive. That report also turned up two gaps: the layout detectors emit a fixed top-k with no NMS, so duplicate regions reached the caller - a post-processing chain now collapses overlaps, drops sub-6px slivers, reference markers and whole-page image false positives (FilterOverlappingRegions, on by default), with optional LayoutNms, LayoutUnclipRatio and LayoutMergeMode; and PP-DocLayoutV3's predicted reading-order index, previously discarded, now orders the blocks (DocumentAnalysisOptions.ReadingOrder, XY-cut still the fallback). Plain-text OCR is unaffected; every change is additive. 3.0.0: BREAKING - two dependencies leave. (1) Imaging moves to EasyImageSharp (MIT, same author, no build-time licence key and no commercial tier for consumers to inherit); Image<Rgb24> and friends appear on the public API (IEasyOcrService, RedactionResult.Image, RedactionOptions.FillColor, DrawAnnotations, BarcodeScanner, the PDF page handlers), so those types now come from a different assembly. (2) The PP-StructureV3 document-structure engine behind AnalyzeDocumentAsync is no longer the third-party PaddleOcrNet package - it is built into this one, under EasyOcrSharp.Structure. Between them, no split-licensed imaging library remains anywhere in the dependency graph. Structure migration is one using directive: PaddleOcrNet.Structure becomes EasyOcrSharp.Structure; StructureResult, StructureBlock and StructureBlockType keep their names, members and behaviour, and everything else that package exposed was engine internals and is now internal. StructureBlock.Lines is now EasyOcrSharp.Models.OcrLine - the same line type the rest of the API returns. Structure models and their cache directory are unchanged, so nothing re-downloads; EASYOCRSHARP_STRUCTURE_MODEL_BASE_URL / _CACHE override them, with the old PADDLEOCRNET_* names still honoured. New dependency: Clipper2 (previously transitive). No OCR behaviour, method name, parameter, default or result shape changed. Migration is one find-and-replace in your using directives: the old imaging namespaces, root plus .PixelFormats / .Processing / .Formats.*, become the matching EasyImageSharp ones. Input format coverage is unchanged or wider (PNG, JPEG incl. progressive and CMYK, WebP, GIF, BMP, TIFF incl. CCITT G3/G4 and JPEG-in-TIFF, TGA, Netpbm, QOI, ICO); WebP is decode-only, so there is no WebP encoder. ImageInfo exposes FrameCount rather than a frame-metadata collection. Deskew preprocessing now uses EasyImageSharp's projection-profile deskew - same estimator, scored on ink coordinates instead of rotating the page once per candidate angle, so it is markedly faster and leaves an already-straight page untouched. See CHANGELOG.md. 2.3.0: thirteen additive capabilities. Recognition detail: RecognitionOptions.WordLevelDetail (default None) populates OcrLine.Words/Characters with true per-word and per-character geometry from the recognizer's CTC alignment, and hOCR/ALTO/TSV emit real word boxes when present. Documents: Unicode searchable PDF via an embedded Type0/CIDFontType2 Identity-H subset font with ToUnicode CMap (PdfOcrOptions.TextLayerFont/TextLayerFontPath, PdfOcrResult.TextLayerFontStatus; no font is bundled — supply one or rely on the system font probe, with graceful Latin-1 fallback); multi-frame TIFF input (ExtractTextFromFramesAsync/StreamTextFromFramesAsync); recovered tables as data (TableHtmlParser, ToRows/ToDataTable/ToCsv). New modes: handwriting recognition via TrOCR (EasyOcrServiceOptions.Handwriting + RecognizeHandwritingAsync; hosted models download on first use, int8 by default, or point at your own export) and barcode/QR reading via ZXing.Net (BarcodeScanner.ReadBarcodesAsync, combined text+barcode pass). Post-processing: redaction that permanently paints over matched text with Luhn-checked card and mod-97 IBAN presets; SymSpell-style post-OCR correction gated on confidence with IBAN/MRZ checksum repair; anchor-based field extraction with invoice presets; CER/WER accuracy metrics. Integration: streaming ExtractTextStreamAsync (IAsyncEnumerable), an easyocrsharp dotnet-tool CLI, and a Dockerized ASP.NET Core sample. All additive and opt-in; no public method renamed and no existing default changed. New dependency: ZXing.Net. 2.2.4: document-structure analysis and document preprocessing. New AnalyzeDocumentAsync (all input overloads) recovers layout regions, tables as HTML, formulas as LaTeX, seals and reading order via PP-StructureV3 (PaddleOcrNet engine), with ToMarkdown()/ToJson() export and DocumentAnalysisOptions (table model choice, per-feature toggles, languages, page orientation/unwarp). New PreprocessingOptions.Sharpen/SharpenAmount (unsharp mask), DocumentOrientation (PP-LCNet single-pass 90/180/270° page fix) and DocumentUnwarp (UVDoc page dewarp) — all default-off. All additive and opt-in; no public method or default changed. 2.2.3: dependency updates (ONNX Runtime 1.27.0, Microsoft.Extensions 10.0.9) plus the hardening + performance + accuracy pass. Security: image decompression-bomb guard (MaxImagePixels), PDF page/size guards (MaxPages, MaxPageMegapixels), HTTPS-only model source + fail-closed checksum verification, model file-name traversal check. Thread-safety: recognizer cache no longer poisoned by a cancelled/failed load; DisposeAsync drains in-flight OCR before releasing sessions. Performance: contiguous Buffer.Span tensor reads, PerspectiveWarp sub-rect copy, CPU intra-op=1 under box-level parallelism, pooled scratch buffers, new WarmUp() to remove cold-start latency. Accuracy: column/font-aware reading order, IoU NMS box de-dup, dominant-script bias on multi-language requests. Added typed exceptions (ModelDownload/Checksum/OfflineModelMissing/PdfProcessing/ImageTooLarge) and OcrResult.SourceWidth/Height. No public method renamed; all additive or safer defaults. See CHANGELOG.md. 2.2.1: clearer typed errors for bad PDFs. 2.2.0: PDF I/O, exporters, telemetry, resilient downloads, auto-GPU, EasyOCR parity.