Proxeno.Kiln
0.3.0
See the version list below for details.
dotnet add package Proxeno.Kiln --version 0.3.0
NuGet\Install-Package Proxeno.Kiln -Version 0.3.0
<PackageReference Include="Proxeno.Kiln" Version="0.3.0" />
<PackageVersion Include="Proxeno.Kiln" Version="0.3.0" />
<PackageReference Include="Proxeno.Kiln" />
paket add Proxeno.Kiln --version 0.3.0
#r "nuget: Proxeno.Kiln, 0.3.0"
#:package Proxeno.Kiln@0.3.0
#addin nuget:?package=Proxeno.Kiln&version=0.3.0
#tool nuget:?package=Proxeno.Kiln&version=0.3.0

Kiln
A from-scratch H.264 encoder for .NET. Pure managed, SIMD-accelerated C#, zero native dependencies, Apache-2.0. Feed it raw frames and it hands back a standards-compliant H.264 baseline bitstream — the whole codec (bitstream, transforms, intra/inter prediction, motion search, entropy coding, deblocking) is implemented here, in this repository, against the ITU-T H.264 specification. No native codec to cross-compile, ship, or keep patched.
Status: pre-release (0.x). APIs will change. Every capability listed here is backed by a test in this repository — including smoke tests that decode every produced stream with an independent reference decoder as an oracle — and the What Kiln is not section says plainly what isn't here.
What you can build
Kiln is the encode step: one .NET process turns rendered or captured frames into an H.264 stream you can send anywhere, with no native runtime on the box. It's built for low-latency, real-time video, such as:
- Game & cloud-gaming streaming — render and encode on a server, play in a browser or thin client with well-under-a-second glass-to-glass latency.
- Screen capture & remote desktop — a headless host encoding its own output frame by frame.
- Camera & robotics video — live feeds from cameras, drones, or robots to an operator's screen.
- A source for a WebRTC / RTP pipeline — Kiln emits Annex B access units that drop straight into a stack like Keryx or any RTP packetizer.
- Anywhere you need low-latency managed H.264 inside a .NET process without bundling a native codec or taking on GPL/LGPL linkage.
Goals
- Pure managed. 100% C# on .NET 10 with hardware intrinsics — NEON/AdvSimd on arm64, AVX2 and SSSE3 on x64, scalar fallback everywhere else. No native library to cross-compile, ship, or patch.
- Genuinely open. Apache-2.0, original work that links no copyleft code — embed it in commercial or proprietary products where GPL/LGPL codec linkage is a problem. That's exactly why it exists.
- Faithful to the spec. Written against ITU-T H.264 (ISO/IEC 14496-10); every numeric table that
originates in the spec carries its clause/table citation in the source (
Table 9-4,§8.5.9,§9.2.1, …). - Real-time first. Predictable per-frame latency over maximum compression: a steady stream of low-latency frames beats a slow exhaustive search.
- Honest about scope. A 0.x that tells you exactly what is proven and what isn't here yet.
Quick start
using Kiln;
using Kiln.RateControl;
var encoder = new H264BaselineEncoder(1280, 720, new H264BaselineEncoderOptions
{
QuantizationParameter = 28,
KeyframeIntervalFrames = 120,
SliceCount = 4, // parallel slice encoding
SpeedMode = EncoderSpeedMode.Balanced, // measured speed/quality preset
});
var annexB = new byte[1280 * 720 * 2];
// Planar I420 input; u/v are half-resolution planes.
var written = encoder.EncodeFrame(y, u, v, strideY: 1280, strideUv: 640, annexB);
var wasIdr = encoder.LastFrameWasIdr;
// annexB[0..written] is a complete Annex B access unit (SPS/PPS included on IDR).
What you get
A real encoder, not a toy — the parts a low-latency streaming server actually needs:
- Baseline bitstream that just plays. Constrained baseline profile, IDR + P-frames, CAVLC entropy coding, Annex B output (SPS/PPS carried on every IDR) that streams straight into a WebRTC or RTP packetizer and decodes on browsers and hardware decoders.
- Genuine coding tools. Intra 4×4 / 16×16 with RD mode selection, P-slice inter prediction with sub-pel SATD motion search, P_Skip, and multiple reference frames.
- Quality knobs. In-loop deblocking, optional greedy trellis quantization, variance-based spatial adaptive QP, and per-frame rate control.
- A measured speed ladder. Four
SpeedModepresets over the motion-search knobs, each a measured point on the speed/quality curve, plus a deterministic per-frame motion-search work budget that bounds the worst-case frame time without making bitstreams timing-dependent. - Parallelism built in. Multi-slice frames encode their slices in parallel and bound the region a lost packet can damage.
- SIMD with a safety net. NEON/AdvSimd, AVX2 and SSSE3 kernels selected at runtime, each covered by parity tests against a scalar reference; CI runs the full suite on Linux, Windows and macOS so both architectures stay green.
- Streaming companions.
Kiln.RateControl(a low-latency rate controller with network feedback) andKiln.Recovery(IDR budgeting / keyframe recovery policy) are public companion namespaces for server use. - Verified. 2,226 tests — spec-roundtrip decoding, SIMD/scalar parity, golden-frame regression, PSNR fidelity floors, adversarial neighbour-availability sweeps, and independent-decoder smoke tests over every produced stream.
Options reference
| Option | Default | What it does |
|---|---|---|
QuantizationParameter |
28 | Base QP, 0–51. |
SpeedMode |
HighQuality |
Measured speed/quality preset ladder (HighQuality/Balanced/Fast/VeryFast) over MaxReferenceFrames, UseMotionSatd, SubPartitionRangeCap and MotionSearchEffortCapPerMb. Any of those four you assign explicitly wins over the mode; other options are never touched. See Performance for what each rung buys and costs. |
KeyframeIntervalFrames |
60 | IDR every N coded frames (frame 0 is always IDR). EncodeFrame(forceKeyframe: true) overrides. |
SliceCount |
1 | Slices per frame; >1 encodes slices in parallel and bounds loss regions. |
MaxReferenceFrames |
2 | 1 = single-ref (WebRTC / hardware-decoder safe), 2 = multi-reference P. |
TargetBitsPerFrame |
0 (off) | Per-MB QP adaptation toward a per-frame bit budget. |
FastSearch |
true | Hex/diamond integer ME + qpel refinement; false = exhaustive integer search. |
UseMotionSatd |
true | SATD scoring for integer-pel ME candidates (SAD for fractional refinement). |
EnableIntraInPFallback |
true | Allows I16x16/I4x4 macroblocks inside P-frames when inter prediction fails. |
TrellisLevel |
0 | 1 = greedy per-coefficient trellis quantization (better RD, ~5% CPU). |
AdaptiveQuantStrength |
0.0 | Variance-based spatial AQ; 1.0 = standard, typical 0.5–1.5. |
PreferRealtimeLatencyTuning |
false | Skips chroma-DC RD refinement for inter-coded chroma. Does not bound motion-search cost — that's MotionSearchEffortCapPerMb. |
LightweightDeblocking |
false | Disables in-loop deblocking (bitstream-signalled) to cut CPU. |
PreferHardwareIntrinsics |
true | Runtime SIMD kernel selection; false forces scalar. |
SubPartitionRangeCap |
16 | Sub-partition ME radius cap (per-frame complexity budget applies). A speed knob: 8 is ~20% faster per frame and 4 ~30% faster. Quality-neutral (±0.01 dB) on coherent motion; on divergent motion 8 costs about −0.17 dB / +13.5% bits at QP 24. |
MotionSearchEffortCapPerMb |
0 (off) | Deterministic worst-case frame-time bound: caps motion-search work per frame (units/MB; each slice gets an equal share) and degrades the search in steps as the budget depletes, paying bitrate/PSNR only when the cap binds. Bitstreams stay reproducible — the count is algorithmic work, not wall clock. Set by the non-default SpeedMode presets; see Performance for measured bounds and prices. |
ProfileIdc / LevelIdc |
66 / 0 (auto) | Signalled profile (baseline) and level. LevelIdc = 0 auto-selects the lowest level whose MaxFS admits the frame, floored at 3.1; set explicitly to pin a level. |
ChromaDcRdLambda, Intra4x4SadLambda |
derived | Expert RD-lambda overrides; leave null. |
Performance
Measured on Apple M5 Max (arm64, NEON/AdvSimd), .NET 10, quiet machine, QP 28, steady P-frames
over deterministic synthetic content. Every number is reproducible from the harnesses in
bench/Kiln.Benchmarks (--slice-quick, --speed-modes, --speed-modes-timing,
--speed-modes-tiers), with competing configurations interleaved in one process so scheduling
drift hits every arm equally. Three questions, in the order a deployment asks them: what does the
default cost, what can a speed mode buy, and what is the worst case.
Default quality on realistic content
Steady P-frame medians on textured scroll-plus-noise content ("coherent motion" — a camera pan or game scroll; neither a best case nor a stress case), default options:
| Resolution | 1 slice | 2 slices | 4 slices | 8 slices |
|---|---|---|---|---|
| 640x480 | 17.3 ms | 8.6 ms | 5.6 ms | 4.0 ms |
| 1280x720 | 19.9 ms | 16.0 ms | 9.4 ms | 7.5 ms |
| 1920x1080 | 24.4 ms | 16.9 ms | 10.9 ms | 10.8 ms |
Slices do not divide the work cleanly. Four slices buy about 2.2x at 1080p — not 4x — and eight
buy essentially nothing beyond four: per-slice motion cost doesn't balance perfectly, and the
frame still pays serial per-frame work. Slices also cost bits, because slice boundaries reset
MV/skip prediction: measured +5% to +26% bitrate at QP 28 going from 1 to 4 slices depending on
content (up to ~+40% on cheap coherent frames at QP 23). Choose SliceCount for latency and
packet-loss confinement and budget for both costs; do not expect linear returns.
The speed-mode ladder: best achievable and its price
H264BaselineEncoderOptions.SpeedMode selects a measured preset over four motion-search knobs
(MaxReferenceFrames, UseMotionSatd, SubPartitionRangeCap, MotionSearchEffortCapPerMb); any
of those you assign explicitly wins over the mode. All rungs keep bitstreams deterministic — the
effort budgets count algorithmic work, never wall clock. 1080p steady-P medians:
| Mode | Sets | Coherent, s=1 / s=4 | Divergent worst case, s=1 / s=4 |
|---|---|---|---|
HighQuality (default) |
2 refs, SATD ME, full sub-partition range, no cap | 26.6 / 10.0 ms | 205 / 70 ms |
Balanced |
1 ref, effort cap 512/MB | 18.8 / 8.2 ms | 80 / 30 ms |
Fast |
+ sub-partition radius 8, cap 256 | 15.6 / 7.7 ms | 64 / 24 ms |
VeryFast |
+ SAD-scored ME, cap 128 | 8.9 / 4.0 ms | 20 / 7.8 ms |
And the quality price, 1080p QP 28, PSNR / bitrate versus HighQuality by content class:
| Mode | Static | Coherent motion | High motion | Scene cut | Divergent motion |
|---|---|---|---|---|---|
Balanced |
0 | −0.1 dB, +1% | −1.6 dB, +12% | +0.3 dB, −3% | −0.4 dB, +40% |
Fast |
0 | −0.1 dB, +1% | −1.7 dB, +11% | +0.4 dB, −4% | −0.5 dB, +55% |
VeryFast |
0 | −1.2 dB, +7% | −3.5 dB, +21% | −0.9 dB, +2% | −1.5 dB, +64% |
Read it as: on static and coherent content Balanced and Fast are near-free and VeryFast
costs about a decibel; when content turns violent, the capped modes hold their frame time and pay
in quality and bits exactly there. The price concentrates at low QP, where frames that genuinely
need a wide search are cut off hardest — at QP 23 Balanced measures −6.8 dB on the high-motion
generator and −3.7 dB on scene cuts. If you run QP ≤ 23 on demanding content, prefer
HighQuality, or compose your own point on the curve (e.g. SpeedMode = Balanced with an
explicit, higher MotionSearchEffortCapPerMb — the explicit knob wins).
Worst case
The number that breaks a real-time deployment is not the mean but the hostile frame. On the
divergent-motion stress generator (opposing half-screen scrolls plus counter-moving blocks) the
default configuration measures 205 ms/frame at 1080p single-slice (70 ms at 4 slices) against
27/10 ms on coherent content in the same run — a ~7x content-dependent swing with no bound. The bound is
MotionSearchEffortCapPerMb (set by every non-default speed mode): a deterministic per-frame
work budget that degrades the search in steps as it depletes, cutting that worst case to 80 ms
(Balanced) / 64 ms (Fast) / 20 ms (VeryFast) single-slice, while charging quality only when
it actually binds. If your content is unusual, size the cap yourself — the
--speed-modes-tiers harness shows how hard a given cap binds per content class.
Microbenchmarks and the perf gate
The committed perf-gate baselines (perf/, BenchmarkDotNet):
| Benchmark | Mean | Min |
|---|---|---|
SATD 4x4 kernel (Satd4x4_Once) |
11.2 ns | 11.1 ns |
SATD 4x4 x 9 intra modes (SatdMany4x4_9Modes) |
99.3 ns | 99.2 ns |
| SAD 8x8 dispatch | 3.5 ns | 3.4 ns |
| SAD 16x16 dispatch (stride 720) | 6.4 ns | 6.4 ns |
| Full-MB 16x16 ME search, range 8 | 76.4 us | 76.1 us |
| Steady P-frame encode, 1280x720, 1 slice | 2.24 ms | 2.20 ms |
That 2.24 ms row is measured on near-static synthetic content — size deployments from the
realistic-content tables above, not from it. Perf discipline is part of the repo:
bench/Kiln.Benchmarks plus scripts/h264-simd-capture-baseline.sh and
scripts/h264-simd-perf-gate.sh gate changes against the committed baseline in perf/. See
docs/perf-gate.md.
What Kiln is not
Not an x264 competitor. No B-frames, no CABAC, no 8×8 transform, no interlace; 4:2:0 8-bit only. If you need maximum compression at any CPU cost, use a full-profile encoder. If you need clean-licensed, dependency-free, low-latency H.264 inside a .NET process, you are in the right place.
Frame dimensions need not be multiples of 16: sizes like 1920×1080 or 1366×768 are supported via
SPS frame cropping — the encoder pads to the 16×16 macroblock grid internally and signals the true
display size, which is what decoders output. Dimensions must be even (4:2:0 chroma is
subsampled 2×2, so odd extents are unrepresentable). By default the encoder signals the lowest
H.264 level whose frame-size limit (MaxFS, Annex A Table A-1) admits the padded picture, floored
at Level 3.1 — 1920×1080 codes as 1920×1088 (8160 macroblocks) and signals Level 4.0 automatically.
Set H264BaselineEncoderOptions.LevelIdc explicitly to pin a level; an explicit level that is too
small for the frame throws, naming the lowest sufficient level. The chosen level is readable from
H264BaselineEncoder.LevelIdc.
0.x behavioural change:
LevelIdcpreviously defaulted to 31 (Level 3.1) and the constructor threw for frames above 1280×720. The default is now 0 = auto-select. Streams ≤720p with default options are byte-identical to before (the auto floor is still 3.1); larger frames now construct and encode instead of throwing.EncodeFrame's output span now has a documented recommended size,H264BaselineEncoder.RecommendedOutputBufferSize.
On the real-time story, be precise about what is wired and what is not. The encoder side is real:
SpeedMode presets and the deterministic motion-search effort cap give measured, bounded per-frame
cost, and Kiln.RateControl's LowLatencyRateController turns network feedback into per-frame
EncoderAdaptationDecisions (bitrate, QP, fps, resolution, and a recommended SpeedMode, which
now has a defined encoder mapping). What does not exist yet is the automatic loop: nothing in the
library applies a controller decision to a running encoder. Your server reads the decision and acts
on it — construct encoder options with the recommended SpeedMode, resize, drop frames — itself.
Likewise the Adaptation (resolution/fps ladders) and Queue (latest-frame dropping) namespaces
ship inside the library, fully tested but not yet wired into the encoder — experimental, and
their APIs may change or move without notice. See docs/architecture.md.
Installing
Published on nuget.org as Proxeno.Kiln. The package
id is Proxeno.Kiln; the assembly and namespace stay Kiln, so code uses using Kiln;.
dotnet add package Proxeno.Kiln
Try it
samples/Kiln.Capture records your camera to a playable .m4v — capture,
colour conversion, H.264 encode and MP4 muxing, all managed, no native binaries anywhere in the
pipeline:
dotnet run --project samples/Kiln.Capture -- list
dotnet run --project samples/Kiln.Capture -- record --seconds 10 --output capture.m4v
Documentation
- samples/Kiln.Capture — camera →
.m4vsample, and how the MP4 muxing works - docs/architecture.md — pipeline stages, SIMD kernel structure, subsystems
- docs/perf-gate.md — benchmark baseline + regression gate workflow
- CONTRIBUTING.md — build, test, and contribution rules
- SECURITY.md — reporting vulnerabilities
License
Apache-2.0 — see LICENSE. © Kiln contributors.
| Product | Versions Compatible and additional computed target framework versions. |
|---|---|
| .NET | net10.0 is compatible. net10.0-android was computed. net10.0-browser was computed. net10.0-ios was computed. net10.0-maccatalyst was computed. net10.0-macos was computed. net10.0-tvos was computed. net10.0-windows was computed. |
-
net10.0
- Microsoft.Extensions.Logging.Abstractions (>= 10.0.0)
NuGet packages
This package is not used by any NuGet packages.
GitHub repositories
This package is not used by any popular GitHub repositories.
0.3.0 — New: EncoderSpeedMode (HighQuality/Balanced/Fast/VeryFast) now actually drives the encoder, trading measured quality for speed; 1080p 4-slice runs 10.0 ms at HighQuality down to 4.0 ms at VeryFast, and worst-case divergent-motion content is bounded from 70 ms to 7.8 ms. Explicitly-set options always override the mode. MotionSearchEffortCapPerMb bounds worst-case motion-estimation cost deterministically. Performance: 1080p 4-slice steady P-frames are ~21% faster than 0.2.0 at identical output bits, via a quadrant-level SAD lower-bound gate, removal of a no-op fractional refinement round, SIMD kernels for the P_Skip acceptance gate, and effort-balanced slice partitioning; the 169 MB reference transform atlas is gone (~610 MB at 4K). Default output is byte-identical to 0.2.0. Docs: the README now separates default-quality, best-achievable and worst-case performance, and states that slices buy ~2.2x rather than 4x. See the GitHub release notes for the full list.