Tedd.FastNoise
1.0.9
dotnet add package Tedd.FastNoise --version 1.0.9
NuGet\Install-Package Tedd.FastNoise -Version 1.0.9
<PackageReference Include="Tedd.FastNoise" Version="1.0.9" />
<PackageVersion Include="Tedd.FastNoise" Version="1.0.9" />
<PackageReference Include="Tedd.FastNoise" />
paket add Tedd.FastNoise --version 1.0.9
#r "nuget: Tedd.FastNoise, 1.0.9"
#:package Tedd.FastNoise@1.0.9
#addin nuget:?package=Tedd.FastNoise&version=1.0.9
#tool nuget:?package=Tedd.FastNoise&version=1.0.9
Tedd.FastNoise
Deterministic coherent noise for voxel worlds and terrain, built for bulk generation.
Perlin, OpenSimplex2, Value and Cellular noise in 2D and 3D, with fractal layering, domain warping, a fusing layer stack and level-of-detail control. Single-point sampling when you need one value; SIMD and multi-core volume fills when you need a million.
The algorithms started as a port of FastNoiseLite, with the generation loop rebuilt around filling buffers instead of answering one question at a time. Where the two still agree, the port is verified against the original; where this library goes further, it goes further.
var noise = new NoiseGenerator(seed: 1337)
{
NoiseType = NoiseType.OpenSimplex2,
FractalType = FractalType.FBm,
Octaves = 5,
Frequency = 0.005f,
};
// One 16x16x256 world column, vectorised across every core.
var density = new float[16 * 256 * 16];
noise.Fill(density, new GridRegion3D(chunkX * 16, 0, chunkZ * 16, 16, 256, 16));
Install
dotnet add package Tedd.FastNoise
For Vulkan compute generation into GPU buffers:
dotnet add package Tedd.FastNoise.Gpu
The GPU package includes the CPU library as a dependency. The CPU package needs no graphics driver.
Targets .NET 10. Optional .NET 11 preview validation uses -p:EnableNet11=true until it ships.
Why this exists
A noise library that only offers GetNoise(x, y, z) forces you to call it once per voxel. That
throws away the two things that make bulk generation fast: sixteen lanes of a vector register doing
the same arithmetic on adjacent coordinates, and the fact that a chunk's worth of samples is
independent work that can be spread across cores. Neither is available to a function that returns
one float.
So the primary API here is a fill:
Fill(destination, region)for a rectangle or a box of samples- a layer stack that runs several noise sources against coordinates held in registers, instead of a buffer per source
- a level-of-detail policy that drops octaves the sample grid cannot represent
Single-point sampling is still there, and still matches the reference exactly. It is just not where the speed is.
What you get
Every CPU backend produces identical bytes
Scalar, SIMD and parallel fills are bit-for-bit equal, on x86 and on ARM, at any vector width. This is a hard guarantee, and the test suite asserts it directly rather than checking values are close.
It matters because a float is usually compared against a threshold — density > 0 decides whether
a voxel is stone or air — and a one-ULP disagreement between a client with AVX-512 and a server on
the scalar fallback is a whole block of disagreement about the world.
Two consequences fall out of that promise:
- The kernels are written once, generically over an operation set, and instantiated per lane width. Scalar and SIMD cannot drift because they are the same source.
- Fused multiply-add is deliberately not used. It would be faster and it would change results by a fraction of an ULP relative to a machine without it.
A port that was checked, not hoped at
The kernels here are not transcriptions. Branchy corner selection became mask arithmetic, loops were unrolled, the whole thing was made generic over lane width -- and every one of those rewrites is a chance to change a value by an ULP and never notice.
So an unmodified copy of FastNoiseLite is vendored into the test project as an oracle, and
CompatibilityTests compares every kernel against it for exact equality across the full matrix of
noise types, fractal types, cellular variants, rotations and domain warps. That is what makes the
rewrites safe to make. It found two real bugs while this was being built, both float association
differences invisible to a tolerance-based test.
Those tests describe the port as it stands today, not a promise about tomorrow. This library will diverge from upstream as it grows -- new algorithms, better quality, features FastNoiseLite has no reason to carry -- and where it does, the corresponding oracle test goes with it. Pin a version if you need output stability.
Layer stacks that fuse
A world is built from layers: continents, then mountains, then hills, then surface detail, then a mask that keeps the detail out of the ocean.
var stack = new NoiseStack { Lod = LodPolicy.Automatic with { CullLayers = true } };
stack.Add(new NoiseLayer
{
Source = new NoiseGenerator(1) { Frequency = 0.0002f, FractalType = FractalType.FBm, Octaves = 4 },
FeatureSize = 2000f,
Name = "continents",
});
stack.Add(new NoiseLayer
{
Source = new NoiseGenerator(2) { Frequency = 0.002f, FractalType = FractalType.Ridged, Octaves = 5 },
Blend = LayerBlend.Add,
Amplitude = 0.4f,
FeatureSize = 200f,
Name = "mountains",
});
stack.Add(new NoiseLayer
{
Source = new NoiseGenerator(3) { Frequency = 0.05f, NoiseType = NoiseType.Value },
Amplitude = 0.02f,
FeatureSize = 8f,
Name = "surface detail",
});
var world = stack.Compile(); // immutable, thread-safe, hand it to workers
world.Fill(heights, new GridRegion2D(0, 0, 512, 512));
The obvious way to combine layers is to fill a buffer per layer and then walk the buffers adding them up. Eight layers over a 512×512 tile means eight full passes writing a megabyte each, then a ninth reading it all back.
Compile() flattens the stack into a flat array of layer plans, and the fill runs every layer
against the coordinates currently in a vector register, blending into an accumulator that never
leaves the register file. Total memory traffic is one write per output value regardless of layer
count.
Blends: Add, Subtract, Multiply, Min, Max, Replace, Lerp. The first layer to survive
culling initialises the accumulator; its blend is ignored.
Zoom levels that cost what they should
Sampling an eight-octave fractal every 512 world units is not just wasteful, it is wrong. Octaves with a wavelength below the sample spacing contribute aliasing, not detail — and when the camera moves, the aliasing changes, so the distant landscape boils.
LodPolicy drops octaves the sample grid cannot carry, and (with CullLayers) skips whole layers
whose FeatureSize is below the spacing:
var noise = new NoiseGenerator(1337)
{
FractalType = FractalType.FBm,
Octaves = 8,
Lod = LodPolicy.Automatic,
};
noise.Fill(closeUp, new GridRegion2D(0, 0, 256, 256, Step: 1f)); // all 8 octaves
noise.Fill(fromOrbit, new GridRegion2D(0, 0, 256, 256, Step: 4096f)); // 1 octave, and correct
FadeLastOctave ramps the finest surviving octave's amplitude across the cull boundary, so detail
appears smoothly as you approach rather than popping in. The ramp continues past the base octave:
once even the coarsest octave is finer than the sample grid, its amplitude falls to zero and the
field flattens to its mean.
That last part matters more than it sounds. The obvious alternative — keep one octave at full amplitude so the field never vanishes — is what makes a zoomed-out view keep its mountain peaks. Those are not mountains any more; they are the aliased remains of an octave the grid cannot carry, and as the camera moves they slide and change shape instead of flattening. Pulling back should smooth the landscape, exactly as the smallest mip of a texture is its average colour rather than one arbitrarily chosen texel. If a view goes flat, that is the honest answer: nothing in the configuration has features large enough to be visible at that spacing.
The normalisation constant is deliberately not recomputed for the reduced octave count — renormalising would make the coarse rendering of a landscape a different height from the fine one, and the terrain would visibly breathe as you flew toward it.
Off by default, because with it off the output is bit-identical to FastNoiseLite at any step.
CompiledNoiseStack.DescribeActiveLayers(step) tells you what a given zoom level will actually
evaluate, so you can check a policy does what you meant.
Procedural graphs
Tedd.FastNoise.Procedural represents complete mathematical composites, not just noise stacks.
Typed inputs, noise, arithmetic, comparisons, conditionals, integer hashes and constant lookups
form an immutable multi-output program. Compilation folds constants, shares identical expressions
and removes unreachable nodes. Arithmetic order and binary32/binary64 boundaries remain explicit;
algebraic reassociation and multiply-add contraction are not enabled.
using Tedd.FastNoise;
using Tedd.FastNoise.Procedural;
var noise = new NoiseGenerator(1337) { Frequency = 0.003f, FractalType = FractalType.FBm, Octaves = 4 };
var graph = new ProceduralGraph();
var x = graph.Input(0, ProceduralType.Float32);
var y = graph.Input(1, ProceduralType.Float32);
var z = graph.Input(2, ProceduralType.Float32);
var height = graph.Noise3D(noise, x, y, z) * graph.Constant(96f) + graph.Constant(64f);
var terrain = graph.Compile(height - y);
terrain.Evaluate(new double[] { 10, 20, 30 }, new double[1]);
Retain compiled programs across calls. Fill consumes interleaved input records and produces
interleaved output records using a generated SIMD delegate where available. Scalar execution uses
a generated delegate; NativeAOT uses the scalar interpreter fallback. Stack3D imports an existing
compiled stack with its LOD plan resolved at the requested spacing. Expose only necessary outputs
to allow elimination of unused fields. Specialize configuration and LOD constants when building
the graph; ordinary runtime inputs are not treated as constants.
The GPU package compiles the same graph into specialized GLSL/SPIR-V. Graph construction does not make arbitrary C# methods GPU-executable: express their pure operations through graph nodes and retain orchestration and stateful work outside the graph.
Planetary terrain recipes
ProceduralGraph is the generic algorithm builder, not a fixed Earth preset. Applications supply
the seed, layers, relief, climate, biome thresholds and material rules as graph expressions.
FastNoise owns the typed intermediate representation, mathematical folding, shared-expression
elimination, LOD-resolved noise and CPU/GPU compilation. No Forcecraft dependency is required.
The planet recipe sample composes continents, ridged mountains, optional detail, moisture, surface materials and sea level. Its column graph returns elevation and moisture; its voxel graph consumes those fields and signed elevation to select air, stone, grass, sand, snow or water. These are illustrative Earth-like rules, not Forcecraft's complete Earth generator. The same pattern accepts an application's own recipe.
// Host policy becomes a specialization constant, not a per-voxel decision.
var resolvedDetail = graph.Select(graph.Constant(step <= detailCutoff), detail, graph.Constant(0d));
var elevation = continents * graph.Constant(relief) + resolvedDetail;
var columns = graph.Compile(elevation, moisture);
// Reuse columns.Evaluate/Fill on CPU, or pass columns plus a material graph to
// VulkanGraphTerrainProducer for column reuse and device-resident packed bricks.
Folding removes constant arithmetic and unreachable branches, not just intermediate memory
traffic. Repeated expressions share one node. Strict floating-point order is retained: the
compiler does not generally rewrite (x * a) * b to x * (a * b) or assume x * 0 is zero.
CPU delegates and GPU shaders are specialized once per immutable recipe/LOD and then reused.
Runtime GPU specialization generates GLSL and compiles optimized SPIR-V; it is not a graph
interpreter running one instruction at a time on the device. Compilation has an up-front cost,
and fewer operations or transfers do not guarantee a faster complete workload.
GPU, when you have one
INoiseAccelerator is the extension point: register one and NoiseBackend.Gpu routes large fills
to it. With none registered, Gpu silently means Parallel, and an accelerator can decline any
individual fill (too small, unsupported configuration) and get the CPU path instead. The fallback
chain is Gpu → Parallel → Simd → Scalar, and every link is tested.
For GPU-resident output, the optional Tedd.FastNoise.Gpu
project provides a Vulkan compute producer. RecordNoise writes float fields; RecordTerrain
writes compact voxel bricks compatible with Forcecraft10's production renderer. Both use
caller-owned devices, command buffers and storage buffers, with no required readback.
Create a LOD-resolved request with noise.CreateRequest(region) and pass it to the producer.
OpenSimplex2, Perlin and Value support None, FBm, Ridged and PingPong fractals. This explicit GPU
API is separate from NoiseBackend.Gpu; cross-device GPU bit equivalence is not guaranteed.
// Reuse the renderer's Vulkan device and a StorageBuffer | TransferDst output buffer.
// Retain these objects until submitted GPU work has completed.
using var producer = new VulkanNoiseProducer(vk, physicalDevice, device);
using var output = producer.BindOutput(buffer, VoxelTerrainSettings.RequiredBytes(32));
var request = noise.CreateRequest(new GridRegion3D(0, 0, 0, 32, 32, 32, Step: 8));
producer.RecordTerrain(commandBuffer, output, request,
new VoxelTerrainSettings(Height: 64, Amplitude: 96));
Use Tedd.FastNoise.Gpu for the Vulkan types above. A terrain cell is solid where
noise * Amplitude + Height - worldY * VerticalScale > 0. The output uses a directory of 4³
bricks, omits empty bricks and stores uniform bricks as one cell. It provides one opaque material;
the host supplies chunk eligibility, world rules, rendering and resource synchronization.
Runnable samples
The Vulkan sample includes complete headless device setup, float-field generation, terrain generation at three LODs and packed-brick decoding:
dotnet run -c Release --project samples/Tedd.FastNoise.Gpu.Sample -- artifacts/gpu-sample
It also executes the composite planetary recipe at three LODs, checks every GPU material against
the same CPU graph, and saves top-down material maps and inspectable generated GLSL. The graph
sample requires Vulkan shaderFloat64. Readback is used only to validate and save sample output;
the library's generation path does not require it. For CPU examples, generate the documentation gallery:
dotnet run -c Release --project tools/Tedd.FastNoise.Gallery -- artifacts/gallery
API
Sampling one point
float v2 = noise.GetNoise(x, y);
float v3 = noise.GetNoise(x, y, z);
Filling a region
noise.Fill(destination, new GridRegion2D(originX, originY, width, height, step));
noise.Fill(destination, new GridRegion3D(originX, originY, originZ, width, height, depth, step));
float[] created = noise.Create(region); // allocates for you
noise.Fill(destination, region, NoiseBackend.Simd); // force a backend
GridRegion3D chunk = GridRegion3D.Chunk(cx, cy, cz, size: 16); // one chunk of a chunked world
Results are written X-fastest: destination[x + width * (y + height * z)]. Nothing in the library
assumes which world axis is up.
Settings
Same names and defaults as FastNoiseLite: Seed, Frequency, NoiseType, RotationType3D,
FractalType, Octaves, Lacunarity, Gain, WeightedStrength, PingPongStrength,
CellularDistanceFunction, CellularReturnType, CellularJitter, DomainWarpType,
DomainWarpAmplitude. Plus Lod and ParallelThreshold.
Noise types
| Type | Cost | Use it for |
|---|---|---|
OpenSimplex2 |
moderate | The default. No axis alignment, so it holds up in 3D density fields. |
OpenSimplex2S |
high | Smoother variant. Scalar only — no wide kernel (see below). |
Perlin |
low | Heightmaps, where mild axis alignment does not show. |
Value |
lowest | Anything that gets thresholded or quantised: ore scatter, per-block variation. |
ValueCubic |
highest | Smooth low-frequency fields. Reads 64 lattice points per 3D sample. |
Cellular |
high | Caves, ore pockets, biome regions, cracks. |
OpenSimplex2S selects its corners with a rank comparison chain, and lanes in a vector disagree
about which branch to take. Rather than evaluate every arm speculatively, bulk fills run the
reference scalar implementation per sample for that type — still parallelised, just not vectorised.
A layer stack fuses as a unit, so one OpenSimplex2S layer holds the whole stack to the scalar
path; CompiledNoiseStack.IsVectorised tells you when that has happened.
Performance
The CPU benchmark graphs compare point loops and scalar, SIMD, and parallel fills across six kernels, single-octave and four-octave fBm signals, and equally sized 2D/3D fields. The raw observations and reproduction instructions accompany the published medians and dispersion. These are host-specific measurements, not latency guarantees.
What it measures, and against what:
| Benchmark | Question |
|---|---|
PointSampling |
Does writing the kernels generically over a lane-width abstraction cost anything? The scalar path should tie with the hand-written reference running identical arithmetic. |
Heightmap2D |
What does a 2D fill cost across scalar, SIMD and parallel, against FastNoiseLite called per sample and against the frozen 2020 implementation in archive/v1? |
VoxelVolume3D |
What does one 16x16x256 world column cost, per noise type? |
LayerFusion |
Does fusing layers into one pass beat a buffer per layer? Should widen with layer count, since the naive version is bound by memory traffic the fused one never generates. |
LevelOfDetail |
What does band-limiting actually save at each zoom level? |
GatherStrategy |
Hardware vgatherdps against a spill-and-index loop for the gradient table reads. |
Run them:
dotnet run -c Release --project src/Tedd.FastNoise.Benchmark -- --filter "*"
Or one at a time:
dotnet run -c Release --project src/Tedd.FastNoise.Benchmark -- --filter "*Heightmap2D*"
The designer
A Windows app for building a stack visually instead of guessing at frequencies and recompiling.
- 2D map — the field as an image, in greyscale, terrain, diverging or viridis colours, with an optional solid/empty mask at a threshold so you can see exactly where terrain would cut.
- 3D heightmap — the same field displaced into terrain, lit and rotatable.
- 3D volume — a 3D field thresholded into voxels and meshed from its exposed faces, rotatable.
- Layers — add, reorder, blend, and toggle layers with every generator setting live.
- Zoom sweep — drag the sample spacing from sub-block to orbital and watch which layers and octaves survive level-of-detail culling, and what the fill costs. The selected world-space centre stays fixed while terrain elevation contracts with the horizontal scale.
- Generated C# — the code that reproduces whatever is on screen, ready to paste.
Drag to orbit, right-drag to pan, wheel to zoom. Every preview shows its own fill time and throughput, so it doubles as a rough profiler for a configuration.
Download the latest build from Releases, or build it yourself:
dotnet run -c Release --project src/Tedd.FastNoise.Designer
How the repository is laid out
src/Tedd.FastNoise/ the library
src/Tedd.FastNoise.Gpu/ optional Vulkan noise and compact voxel producer
src/Tedd.FastNoise.Gpu.Tests/ shader compilation, GPU agreement and payload tests
src/Tedd.FastNoise.Tests/ xUnit, including the vendored reference used as the oracle
src/Tedd.FastNoise.Benchmark/ BenchmarkDotNet
src/Tedd.FastNoise.Designer/ the WPF designer
src/Tedd.FastNoise.Designer.Tests/ Windows-only designer geometry tests
tools/Tedd.FastNoise.Gallery/ renders the sample images for the documentation site
samples/Tedd.FastNoise.Gpu.Sample/ runnable Vulkan field and terrain examples
docs/ the GitHub Pages site
archive/v1/ the 2020 implementation, frozen
archive/v1 is not dead code kept out of sentiment. It is the fixed reference point every
performance claim is measured against; it is retargeted to a supported framework and otherwise
untouched, because a moving baseline measures nothing.
The same discipline applies to the benchmark project: each class documents the question it answers, and nothing lands in the library on the strength of an argument that it ought to be faster.
Building
dotnet build src/Tedd.FastNoise.slnx -c Release
dotnet test src/Tedd.FastNoise.Tests -c Release
dotnet test src/Tedd.FastNoise.Tests -c Release -f net11.0 -p:EnableNet11=true # needs the .NET 11 SDK
Releasing
Run validation locally before publishing. With PowerShell 7, the .NET 10 SDK and a Vulkan compute device, this command builds, tests, packs both libraries and refreshes the website samples:
pwsh -File tools/Validate-Release.ps1 -Gpu -Gallery
It includes the Windows designer tests on Windows. Omit -Gallery to preserve the existing images;
omit both switches for CPU and shader compilation checks without a Vulkan device. Validate GPU
changes with -Gpu before release. Review and commit the generated docs/gallery/*.png files with
the source changes. Binary sample payloads remain local.
ci.yml is manual-only, for occasional Linux, ARM64, Windows and .NET 11 preview verification.
Pushes and pull requests do not start hosted tests. Routine validation and sample generation run
locally; the deployment workflows only package and publish.
Manually dispatch deploy.yml from the deploy branch to ship:
- matching versions of
Tedd.FastNoiseandTedd.FastNoise.Gpu, using GitHub OIDC trusted publishing - a GitHub release carrying the self-contained Windows designer
pages.yml publishes the committed docs/ files when documentation changes reach deploy.
It installs no .NET or Vulkan runtime and generates no samples on hosted runners. Documentation
updates therefore do not require a package release.
The <Version> value supplies the major and minor release line. Each new deploy workflow run adds
its stable run number to the patch component; rerunning the same workflow retains the same version.
After local validation and merging to main, publish the site and dispatch the package release:
git push origin main:deploy
gh workflow run deploy.yml --ref deploy
NuGet.org must trust repository tedd/Tedd.FastNoise and workflow file deploy.yml for both packages. No persistent
NuGet API key is stored. GitHub Pages must use GitHub Actions as its source.
Not done yet
- GPU coverage. The Vulkan producers support OpenSimplex2, Perlin and Value, including
supported stacks imported through
ProceduralGraph.Stack2D/Stack3D. Graphs requireshaderFloat64; Float64 power and other noise types are unsupported on GPU. A host-readbackINoiseAcceleratoradapter is not implemented. - Domain warp in bulk.
DomainWarpworks per point, via the reference implementation. There is no vectorised warp inside the fill loop yet, so warping a whole region means warping coordinates yourself and sampling per point. - 4D noise. FastNoiseLite does not have it either; it would be useful for looping animation.
Licence
LGPL 2.1 — see LICENSE.
Incorporates FastNoiseLite by Jordan Peck under the MIT licence; see THIRD-PARTY-NOTICES.md for what is used and where.
| Product | Versions Compatible and additional computed target framework versions. |
|---|---|
| .NET | net10.0 is compatible. net10.0-android was computed. net10.0-browser was computed. net10.0-ios was computed. net10.0-maccatalyst was computed. net10.0-macos was computed. net10.0-tvos was computed. net10.0-windows was computed. |
-
net10.0
- No dependencies.
NuGet packages (1)
Showing the top 1 NuGet packages that depend on Tedd.FastNoise:
| Package | Downloads |
|---|---|
|
Tedd.FastNoise.Gpu
Vulkan compute noise and GPU-resident brick voxel terrain generation. |
GitHub repositories
This package is not used by any popular GitHub repositories.
Reduced 3D gradient-table gathers in Perlin and OpenSimplex2 while preserving exact scalar, SIMD and parallel outputs. Includes typed procedural graphs and LOD-resolved noise stacks for CPU and Vulkan terrain generation.