Endatix.Modules.Jobs 0.7.8

dotnet add package Endatix.Modules.Jobs --version 0.7.8
                    
NuGet\Install-Package Endatix.Modules.Jobs -Version 0.7.8
                    
This command is intended to be used within the Package Manager Console in Visual Studio, as it uses the NuGet module's version of Install-Package.
<PackageReference Include="Endatix.Modules.Jobs" Version="0.7.8" />
                    
For projects that support PackageReference, copy this XML node into the project file to reference the package.
<PackageVersion Include="Endatix.Modules.Jobs" Version="0.7.8" />
                    
Directory.Packages.props
<PackageReference Include="Endatix.Modules.Jobs" />
                    
Project file
For projects that support Central Package Management (CPM), copy this XML node into the solution Directory.Packages.props file to version the package.
paket add Endatix.Modules.Jobs --version 0.7.8
                    
#r "nuget: Endatix.Modules.Jobs, 0.7.8"
                    
#r directive can be used in F# Interactive and Polyglot Notebooks. Copy this into the interactive tool or source code of the script to reference the package.
#:package Endatix.Modules.Jobs@0.7.8
                    
#:package directive can be used in C# file-based apps starting in .NET 10 preview 4. Copy this into a .cs file before any lines of code to reference the package.
#addin nuget:?package=Endatix.Modules.Jobs&version=0.7.8
                    
Install as a Cake Addin
#tool nuget:?package=Endatix.Modules.Jobs&version=0.7.8
                    
Install as a Cake Tool

Endatix.Modules.Jobs

Durable background job queue for long-running tenant operations.

Endatix Platform is an open-source data collection and management library for .NET. It is designed for building secure, scalable, and integrated form-centric applications that work with SurveyJS. Endatix empowers business users with advanced workflows, automation, and meaningful insights.

Installation:

dotnet add package Endatix.Modules.Jobs

For running and hosting the Endatix Platform, Endatix.Api.Host is the recommended main package as it simplifies the installation and setup process.

dotnet add package Endatix.Api.Host

The module is registered automatically by EndatixBuilder.UseDefaults() — no host code change is needed.

What it is

Two stores in one jobs schema, with one job each:

  • The BackgroundJobs table is the record of every job: tenant, status, progress, attempts, error, trace and retention. It is what GET jobs/{jobId} reads, and the only place a job's state lives.
  • Quartz.NET 4.2.2, clustered through the same database (tables jobs.qrtz_*), is the queue: it decides when and on which node a job runs, retries it on its trigger's policy, and re-runs the work of a node that stopped.

Every job has exactly one row and one Quartz trigger, keyed by the job's id; the trigger carries nothing but that id.

Module layout

Namespace Contents
Endatix.Core.Abstractions.BackgroundJobs IBackgroundJobQueue, IBackgroundJobHandler, BackgroundJobHandler<TPayload>, IBackgroundJobPayload, BackgroundJobRequest, BackgroundJobPayloadSerializer, JobStatus — referenced by anything that enqueues or handles, and free of any Quartz type
Endatix.Modules.Jobs.Domain BackgroundJob entity and its state machine
Endatix.Modules.Jobs.Persistence jobs-schema contexts, EF configuration, migrations (including Quartz's tables)
Endatix.Modules.Jobs.Features BackgroundJobQueue
Endatix.Modules.Jobs.Runtime BackgroundJobsOptions, the Quartz registration, the handler registry, the job wrapper, the job state repository and IJobMetrics

The abstractions live in Endatix.Core rather than here so that assemblies which cannot reference this module — Endatix.Infrastructure, most notably — can still enqueue and handle.

Schema

Database schema: jobs

Table Purpose
BackgroundJobs One row per unit of work: type, payload, tenant, status, progress, retry state
qrtz_* Quartz.NET's clustered job store: durable jobs, triggers, fired triggers, node check-ins, locks

The schema carries its own __EFMigrationsHistory, so job migrations advance independently of app-schema migrations. Quartz's tables are created by the jobs migrations only, from the PostgreSQL script embedded in the referenced Quartz.NET version, so a fresh database always gets the schema that version expects. Quartz validates them at startup (SchemaProvisioning.Validate) and never creates them, so a node whose tables are missing or outdated fails to start. A Quartz upgrade that changes its schema also ships a jobs migration, for databases created before it. That migration must be idempotent (ADD COLUMN IF NOT EXISTS, CREATE INDEX IF NOT EXISTS, and so on): a database created after the upgrade already has the new schema from InitialBackgroundJobs when it runs.

Status

Pending → Processing → Completed | Failed | DeadLettered | Retrying | Canceled, and Retrying → Processing. A job cancelled before it is claimed never runs.

Status Meaning Terminal
Pending Enqueued, never started no
Processing Claimed by the job wrapper on the node Quartz fired it on no
Retrying An attempt failed retryably; waiting for its next attempt no
Completed Success yes
Failed Deterministic failure — retrying cannot help yes
DeadLettered Retryable failure that exhausted its attempt budget yes
Canceled Cancelled by a user yes

Failed and DeadLettered are separate because one status cannot express both "do not retry this" and "retried and gave up", and operators need to tell them apart.

Enqueueing

Feature code builds a request from a typed payload:

public sealed record SubmissionExportPayload(long FormId, long ExportFormatId) : IBackgroundJobPayload
{
    // Persisted on the row and in Quartz's job keys: never change it.
    public static string JobType => "SubmissionExport";
}

var jobId = await backgroundJobQueue.EnqueueAsync(
    BackgroundJobRequest.Create(new SubmissionExportPayload(formId, formatId), tenantId, userId),
    cancellationToken);

Idempotent enqueue. A request may carry a DedupKey naming its unit of work — for a fan-out, {outboxMessageId}:{subscriber}. It is unique per tenant and job type (a filtered unique index, IX_BackgroundJobs_DedupKey), and enqueueing a key that already exists returns the existing job's id without a second row or trigger, even when two enqueues race. Requests without a key never collide. It names the work, never the attempt: no timestamps or counters in it.

Use EnqueueManyAsync for fan-out — one job per webhook endpoint, say. The batch commits in a single transaction, so a partial fan-out cannot deliver to some destinations and silently drop the rest.

Enqueueing is not transactionally joined to app-schema writes: jobs live on their own DbContext, which cannot enlist in an AppDbContext transaction. To commit a domain change and a job together, raise a domain event and enqueue from the outbox — the outbox already guarantees the event survives the business transaction.

Writing a handler

Derive from BackgroundJobHandler<TPayload> and register it with services.AddBackgroundJobHandler<THandler, TPayload>(), which keys it by job type so a job run builds only its own handler. The job type comes from the payload; handlers may live in any assembly. Two handlers declaring the same job type fail startup.

internal sealed class SubmissionExportJobHandler(..., ILogger<SubmissionExportJobHandler> logger)
    : BackgroundJobHandler<SubmissionExportPayload>(logger)
{
    protected override Task<Result> ExecuteAsync(
        BackgroundJobContext job, SubmissionExportPayload payload, CancellationToken cancellationToken) => ...;
}

Payload rules: the job type is a string literal declared once, on the payload; a payload written by one release must deserialize in the next (adding an optional property is fine; renaming, removing or retyping one is a new job type); keep payloads thin — ids plus the minimum non-personal data. Input that cannot be read ends the job Failed without a retry, and the base class logs why through the logger it is given.

Four obligations, each invisible until it hurts in production:

  1. Return a failure Result for deterministic errors; throw only for transient ones. This is the only retry signal there is. Throwing on a permanent error re-runs expensive work until the attempt budget is gone.
  2. Scope every query to the job's TenantId explicitly. Outside a request the ambient tenant filter is permissive, not restrictive — a handler that queries as if it were in a request reads every tenant's data.
  3. Honour the CancellationToken, or the job cannot be cancelled or time-limited.
  4. Do not hold one DbContext for the length of the job. Open a scope per chunk via IServiceScopeFactory; a change tracker held for minutes accumulates every row streamed through it.

Handlers never reference a Quartz type.

Runtime

Every host that registers the module builds a Quartz scheduler named endatix-jobs against the shared store. Endatix:BackgroundJobs:RunInProcess decides what it does with it:

RunInProcess The host
true (default) Enqueues and executes jobs
false Enqueues only: its scheduler has no threads and is never started

That is what allows API and worker roles to be deployed separately from the same image.

Each job type is its own Quartz execution group, capped per node at JobTypes:{JobType}:MaxConcurrency (default 1). Every group this host has no handler for is capped at 0, so a node never takes a job it cannot run; the job waits for a node that can. The thread pool holds every registered job type's full cap at once, so it is sized as the sum of the caps, and there is no global concurrency setting. Trigger acquisition is narrowed to the groups with a free slot on the node, so a backlog of one job type never holds up another. If a node ever fires a job type it has no handler for, the wrapper declines it without touching the row and offers the trigger again 30 seconds later.

Backlog warning. A job that waited past MisfireThresholdSeconds for a slot increments endatix.jobs.misfired (tag job_type), and again at every threshold while it waits. The warning Background job {JobId} of type {JobType} waited past the misfire threshold; {Misfires} triggers of this type misfired since the last warning. is logged at most once per job type per threshold. A job type no running node can handle shows up here: deploy the module that handles it.

Dashboard. With Dashboard:Enabled, the Quartz dashboard is served at /quartz and its HTTP API at /quartz-api, both behind the PlatformAdmin policy and read-only unless Dashboard:AllowWrites is set. It is an operator tool: a caller that may write can schedule jobs on the host, and it shows every tenant's triggers. Even with writes on, only Endatix's own job class may be named.

Each job type the host has a handler for gets one durable Quartz job, which requests recovery, so a job cut off by a stopped or crashed node runs again on another. Every node sharing the store must run the same Quartz version, and nodes' clocks must agree within about a second.

Retention

JobRetentionJob runs on the Retention:Cron schedule, on one node at a time, in an execution group of its own with one thread beyond the job types' caps. Each run deletes finished rows whose ExpiresAt has passed, BatchSize at a time, for at most MaxBatchesPerRun batches. It never deletes a row that is not finished. Every terminal write sets ExpiresAt to its time plus the job type's RetentionDays when the row has none. Quartz deletes its own finished triggers.

Execution

BackgroundJobExecution is the only Quartz job class. It only orchestrates each firing, through one class per step:

  1. Admits the firing (JobFiringAdmission). A firing that carries no job id, as one an operator fires by hand from the dashboard does, is ignored. A job type this host has no handler for is declined without touching the row, and its trigger is offered again 30 seconds later.
  2. Claims the row (JobAttemptClaimer) with a compare-and-swap from Pending/Retrying to Processing that increments AttemptCount. When Quartz reports a recovered firing, or the firing was scheduled to take over an attempt that left its row unsettled, it re-claims the row from Processing, fenced on the attempt it read, and dead-letters a job that has no attempt left instead. A claim that changes nothing ends the firing. A take-over firing with no retry policy is the job's only trigger, so a claim that throws there re-fires the job rather than ending it.
  3. Runs the handler (JobHandlerRunner) in its own DI scope, under an Endatix.Jobs activity whose parent is the trace captured at enqueue, with one token linked from the runtime ceiling (MaxRuntimeMinutes), the cancellation watcher and the host's shutdown.
  4. Records the outcome (JobOutcomeRecorder) with one write fenced on the claimed attempt, tried again a few times if it throws, and records the lifecycle metrics only when it lands. A firing whose outcome still cannot be written is re-fired shortly by UnrecordedJobRefire. A recovered firing has no retry policy, so a retry it needs gets a trigger of its own, stored before the row says Retrying:
Handler Row Quartz
returns success Completed done
returns a failure Result Failed, with its message done, no retry
throws, attempts left Retrying the trigger's retry policy schedules the next attempt
throws, attempts left, on a recovered firing Retrying a trigger of the job's own runs the next attempt
as above, but that trigger cannot be stored stays Processing, nothing written fires again 5 s later and re-claims the row
throws, attempts spent DeadLettered, with a safe message done
row set to Canceled meanwhile stays Canceled done
host stopped waiting for it stays Processing, nothing written re-run on the next node to check in
its node stopped during the last attempt DeadLettered when recovered, the handler not run again done
row taken over by another attempt meanwhile left to that attempt done; the handler's token is cancelled
its outcome cannot be written (tried 4 times) stays Processing fires again 5 s later and re-claims the row

A re-fire is a trigger too, and storing it can fail. While the scheduler is stopping it refuses every new trigger but still completes the firings it waits for, which would delete the job's last trigger; the firing is held instead until the scheduler lets go of it, and the next node to check in recovers it. A re-fire that fails on a node that keeps running is logged at Error: the row stays Processing, and nothing runs it again unless that node stops before the firing completes.

The row's AttemptCount, not Quartz's retry counter, decides dead-lettering: a run recovered after a crash consumes an attempt Quartz never counts. Exception text never reaches ErrorMessage. After Quartz schedules a retry, NextAttemptAt mirrors the trigger's next fire time.

Shutdown. A stopping host gives running jobs ShutdownWaitSeconds to finish and record their outcome. Jobs still running then are left as they are: their handlers are told to stop, nothing is recorded, and Quartz re-runs them on the next node to check in.

Cancellation. Setting a row to Canceled reaches a running handler through the wrapper's watcher, which re-reads the status every CancellationPollSeconds, on whichever node runs the job.

IJobMetrics is public so a host can replace it. The default records on the Endatix.Jobs meter (JobsModule.MeterName), and EndatixTelemetryBuilder subscribes the host's metrics pipeline to it.

Instrument Type Unit Tags
endatix.jobs.events counter {event} endatix.job.type, endatix.job.event
endatix.jobs.duration histogram s endatix.job.type, endatix.job.outcome

Configuration

Under Endatix:BackgroundJobs, with per-job-type overrides under JobTypes:{JobType}:

Key Default Purpose
RunInProcess true Whether this host executes jobs
IdleWaitTimeSeconds 2 How long an idle node waits before looking for jobs another node scheduled
CancellationPollSeconds 10 How often a running job notices it was cancelled
MisfireThresholdSeconds 60 How long a due job may wait for a slot before the backlog warning
ShutdownWaitSeconds 30 How long a stopping host waits for running jobs before leaving them for recovery; never longer than the host's HostOptions.ShutdownTimeout (30 s by default)
MaxRuntimeMinutes 60 Ceiling on one attempt (per type)
MaxAttempts 3 Attempts before DeadLettered (per type)
BackoffBaseSeconds / BackoffCapSeconds 30 / 900 Retry backoff (per type)
RetentionDays 7 How long a finished job's row is kept (per type); stamped as ExpiresAt when it finishes
Retention:Cron / BatchSize / MaxBatchesPerRun 0 0/15 * * * ? / 1000 / 50 When the retention job runs and how much one run deletes
JobTypes:{JobType}:MaxConcurrency 1 Jobs of this type one node runs at once; 0 declines the type
Clustering:CheckinIntervalSeconds / CheckinMisfireThresholdSeconds 7.5 / 7.5 A node silent for their sum is presumed dead
Dashboard:Enabled / Dashboard:AllowWrites false / false The operator dashboard, for platform admins
Clustering:InstanceId generated This node's identity; set only to a value no other running node uses

Registration

Registered via EndatixBuilder.UseDefaults() → UseModule(JobsModule.Instance). Capabilities: IEndatixModule, IHasFeatureFlag, IHasDbMigrations, IHasFastEndpoints. Do not also Api.ScanAssemblies this assembly (bypasses the flag).

Gated by Endatix:FeatureFlags:JobsModule, off by default. The module owns a DbContext and its own migrations, so registering it where nothing enqueues would create a schema no code writes to.

Migrations

PostgreSQL is currently the only supported provider. With the flag off — the default — nothing is registered and other providers are unaffected. With the flag on and a different provider configured, the host fails at startup naming the constraint.

Persistence is provider-split: JobsPostgreSqlDbContext derives from JobsDbContextBase and owns its migrations and model snapshot under Persistence/Migrations/PostgreSql. EF Core keeps one model snapshot per context type, so adding a provider means adding a derived context, its own design-time factory, its own Config/<Provider>/ configuration and its own migrations folder — never reusing an existing one.

The migrations were reset once, in the change that moved scheduling onto Quartz, to a single InitialBackgroundJobs migration. v0.7.6 and the canaries before this change shipped the earlier ones (AddBackgroundJobs, RequireRealTenantOnBackgroundJobs), and EF cannot upgrade a database from them: it would try to create jobs."BackgroundJobs" again and fail. Any database where Endatix:FeatureFlags:JobsModule was turned on under those releases needs DROP SCHEMA jobs CASCADE once before upgrading. That deletes every job row and the schema's migration history. Those rows never ran, because no release executed jobs, and they cannot be carried over, because each job now needs a Quartz trigger as well as its row. Back up jobs."BackgroundJobs" first if you need a record of them. From here on jobs migrations are append-only.

Run the commands from the repository root, with Endatix.WebHost as the startup project.

dotnet ef migrations add <Name> \
  --startup-project src/Endatix.WebHost \
  --project src/Endatix.Modules.Jobs \
  --context JobsPostgreSqlDbContext \
  --output-dir Persistence/Migrations/PostgreSql

Migrations apply automatically at startup when Endatix:Data:EnableAutoMigrations is enabled.

Third-party licence

Quartz.NET is Apache-2.0 licensed. Its licence text ships in THIRD-PARTY-NOTICES, both in this package and in the API image (/app/THIRD-PARTY-NOTICES).

Product Compatible and additional computed target framework versions.
.NET net10.0 is compatible.  net10.0-android was computed.  net10.0-browser was computed.  net10.0-ios was computed.  net10.0-maccatalyst was computed.  net10.0-macos was computed.  net10.0-tvos was computed.  net10.0-windows was computed. 
Compatible target framework(s)
Included target framework(s) (in package)
Learn more about Target Frameworks and .NET Standard.

NuGet packages (1)

Showing the top 1 NuGet packages that depend on Endatix.Modules.Jobs:

Package Downloads
Endatix.Hosting

Package Description

GitHub repositories

This package is not used by any popular GitHub repositories.

Version Downloads Last Updated
0.7.8 100 10/2/2026
0.7.7 96 10/1/2026
0.7.6 238 9/9/2026