dtpipe 1.4.0

There is a newer version of this package available.
See the version list below for details.
dotnet tool install --global dtpipe --version 1.4.0
                    
This package contains a .NET tool you can call from the shell/command line.
dotnet new tool-manifest
                    
if you are setting up this repo
dotnet tool install --local dtpipe --version 1.4.0
                    
This package contains a .NET tool you can call from the shell/command line.
#tool dotnet:?package=dtpipe&version=1.4.0
                    
nuke :add-package dtpipe --version 1.4.0
                    

DtPipe

A self-contained CLI for streaming, transforming, and anonymizing data across databases and file formats.

DtPipe reads from a source, applies row and columnar transformations in batches, and writes to a destination with no intermediate staging. It is designed for automation and CI/CD workflows where repeatable, observable data pipelines matter.


๐Ÿ“– Recipes & Examples โ†’ COOKBOOK.md ยท Full CLI Reference โ†’ REFERENCE.md


Installation

dotnet tool install -g dtpipe
dtpipe --help

Build from Source

Prerequisite: .NET 10 SDK

# Bash (Mac/Linux/Windows Git Bash)
./build.sh

# PowerShell (Windows/Cross-platform)
./build.ps1

Binary created at: ./dist/release/dtpipe


Quick Start

Export a database table

dtpipe \
  -i "pg:Host=localhost;Database=prod;Username=postgres" \
  --query "SELECT * FROM users" \
  -o users.parquet

Anonymize before export

dtpipe \
  -i "pg:Host=localhost;Database=prod;Username=postgres" \
  --query "SELECT * FROM users" \
  --fake "email:internet.email" \
  --fake "name:name.fullName" \
  --mask "phone:###-****" \
  --null "ssn" \
  -o anonymized_users.parquet

In-memory SQL join

dtpipe \
  -i orders.parquet --alias orders \
  -i customers.csv --alias customers \
  --from orders --ref customers \
  --sql "SELECT o.*, c.name FROM orders o JOIN customers c ON o.customer_id = c.id" \
  -o result.parquet

Run from a YAML job file

# Generate a reusable job file from any CLI command
dtpipe -i "pg:..." --query "SELECT * FROM users" --fake "email:internet.email" \
       -o users.parquet --export-job nightly.yaml

# Run it (with optional overrides)
dtpipe --job nightly.yaml --limit 1000

Incremental loading (cursor-driven)

# First run: Full load, initializes the state file with the max updated_at cursor value
dtpipe \
  -i "pg:Host=localhost;Database=prod" \
  --query "SELECT * FROM users WHERE updated_at >= '${{cursor://state.json|1970-01-01}}'" \
  -o "sqlite:Data Source=dw.db" \
  --table "users" \
  --strategy Recreate \
  --key id \
  --cursor "updated_at" \
  --state "state.json"

# Subsequent runs: Incremental load, only retrieves newer records
dtpipe \
  -i "pg:Host=localhost;Database=prod" \
  --query "SELECT * FROM users WHERE updated_at > '${{cursor://state.json}}'" \
  -o "sqlite:Data Source=dw.db" \
  --table "users" \
  --strategy Upsert \
  --key id \
  --cursor "updated_at" \
  --state "state.json"

Providers

DtPipe detects providers from file extensions (.csv, .parquetโ€ฆ) or explicit prefixes. Explicit prefixes are recommended to avoid ambiguity.

Provider Input Output Prefix
DuckDB โœ… โœ… duck:
SQLite โœ… โœ… sqlite:
PostgreSQL โœ… โœ… pg:
Oracle โœ… โœ… ora:
SQL Server โœ… โœ… mssql:
CSV โœ… โœ… csv: / .csv
JsonL โœ… โœ… jsonl: / .jsonl
XML โœ… โ€” xml: / .xml
Apache Arrow โœ… โœ… arrow: / .arrow
Parquet โœ… โœ… parquet: / .parquet
Data Gen โœ… โ€” generate:N
Null โ€” โœ… null:
Checksum โ€” โœ… checksum:

Use keyring://alias anywhere a connection string is expected. DtPipe resolves it from the OS keychain at runtime. Run dtpipe secret set prod-db "pg:..." to store a secret.

DtPipe's native providers cover common sources and destinations. For everything else โ€” object storage (S3, GCS, Azure Blob), Iceberg, MySQL/MariaDB, HTTP APIs, spatial formats โ€” DuckDB's extension ecosystem serves as a connector multiplier. Load an extension with --duck-init on a DuckDB reader, writer, or --sql branch to reach any source or destination DuckDB supports natively. No additional adapters required.


Key Concepts

Transformers (--fake, --mask, --compute, --filter, โ€ฆ) chain left-to-right. When source and destination are both columnar (Parquet, DuckDB, Arrow), data flows through without row conversion. Multiple --input sources with --from, --sql, or --merge form a DAG executed concurrently. Any CLI command can be saved to a YAML job file with --export-job and replayed with --job.

DuckDB is a remarkable engine โ€” fast, self-contained, with a rich SQL dialect and a thriving extension ecosystem. DtPipe uses it as a first-class component precisely because of that quality. When DuckDB alone covers your use case, use it directly. DtPipe adds value in the scenarios it wasn't designed for: anonymizing or masking data in transit, routing one source to multiple destinations concurrently, writing to target databases with strategies like upsert, auto-migrate, or bulk insert, reading from Oracle, SQL Server, or XML streams, and packaging pipelines as repeatable YAML jobs with integrated secret management. DtPipe contributes the pipeline layer; DuckDB contributes the SQL engine.


Documentation

Document Contents
REFERENCE.md Full CLI option tables, YAML job schema, DAG topology reference, secret management
COOKBOOK.md End-to-end scenarios: anonymization, schema transforms, SQL joins, DAG pipelines, YAML automation
EXTENDING.md Adding adapters (readers/writers) and transformers

Shell Autocompletion (experimental)

dtpipe completion --install

Restart your terminal (or source ~/.zshrc) to activate.


Contributing

See EXTENDING.md for the adapter and transformer patterns.

License

MIT

Product Compatible and additional computed target framework versions.
.NET net10.0 is compatible.  net10.0-android was computed.  net10.0-browser was computed.  net10.0-ios was computed.  net10.0-maccatalyst was computed.  net10.0-macos was computed.  net10.0-tvos was computed.  net10.0-windows was computed. 
Compatible target framework(s)
Included target framework(s) (in package)
Learn more about Target Frameworks and .NET Standard.

This package has no dependencies.

Version Downloads Last Updated
1.9.0 0 10/2/2026
1.8.2 110 9/14/2026
1.8.1 110 9/10/2026
1.8.0 113 9/9/2026
1.7.0 106 9/5/2026
1.6.0 183 8/27/2026
1.5.0 167 8/6/2026
1.4.3 153 7/7/2026
1.4.2 147 6/20/2026
1.4.1 148 6/19/2026
1.4.0 167 6/17/2026
1.3.4 152 6/15/2026
1.3.3 136 6/14/2026
1.3.2 141 6/9/2026
1.3.1 152 6/8/2026
1.3.0 213 5/8/2026
1.2.6 147 4/29/2026
1.2.5 159 4/25/2026
1.2.4 149 4/14/2026
Loading failed