dtpipe 1.3.3

There is a newer version of this package available.
See the version list below for details.
dotnet tool install --global dtpipe --version 1.3.3
                    
This package contains a .NET tool you can call from the shell/command line.
dotnet new tool-manifest
                    
if you are setting up this repo
dotnet tool install --local dtpipe --version 1.3.3
                    
This package contains a .NET tool you can call from the shell/command line.
#tool dotnet:?package=dtpipe&version=1.3.3
                    
nuke :add-package dtpipe --version 1.3.3
                    

DtPipe

A self-contained CLI for streaming, transforming, and anonymizing data across databases and file formats.

DtPipe reads from a source, applies row and columnar transformations in batches, and writes to a destination with no intermediate staging. It is designed for automation and CI/CD workflows where repeatable, observable data pipelines matter.


๐Ÿ“– Recipes & Examples โ†’ COOKBOOK.md ยท Full CLI Reference โ†’ REFERENCE.md


Installation

dotnet tool install -g dtpipe
dtpipe --help

Build from Source

Prerequisite: .NET 10 SDK

# Bash (Mac/Linux/Windows Git Bash)
./build.sh

# PowerShell (Windows/Cross-platform)
./build.ps1

Binary created at: ./dist/release/dtpipe


Quick Start

Export a database table

dtpipe \
  -i "pg:Host=localhost;Database=prod;Username=postgres" \
  --query "SELECT * FROM users" \
  -o users.parquet

Anonymize before export

dtpipe \
  -i "pg:Host=localhost;Database=prod;Username=postgres" \
  --query "SELECT * FROM users" \
  --fake "email:internet.email" \
  --fake "name:name.fullName" \
  --mask "phone:###-****" \
  --null "ssn" \
  -o anonymized_users.parquet

In-memory SQL join

dtpipe \
  -i orders.parquet --alias orders \
  -i customers.csv --alias customers \
  --from orders --ref customers \
  --sql "SELECT o.*, c.name FROM orders o JOIN customers c ON o.customer_id = c.id" \
  -o result.parquet

Run from a YAML job file

# Generate a reusable job file from any CLI command
dtpipe -i "pg:..." --query "SELECT * FROM users" --fake "email:internet.email" \
       -o users.parquet --export-job nightly.yaml

# Run it (with optional overrides)
dtpipe --job nightly.yaml --limit 1000

Providers

DtPipe detects providers from file extensions (.csv, .parquetโ€ฆ) or explicit prefixes. Explicit prefixes are recommended to avoid ambiguity.

Provider Input Output Prefix
DuckDB โœ… โœ… duck:
SQLite โœ… โœ… sqlite:
PostgreSQL โœ… โœ… pg:
Oracle โœ… โœ… ora:
SQL Server โœ… โœ… mssql:
CSV โœ… โœ… csv: / .csv
JsonL โœ… โœ… jsonl: / .jsonl
XML โœ… โ€” xml: / .xml
Apache Arrow โœ… โœ… arrow: / .arrow
Parquet โœ… โœ… parquet: / .parquet
Data Gen โœ… โ€” generate:N
Null โ€” โœ… null:
Checksum โ€” โœ… checksum:

Use keyring://alias anywhere a connection string is expected. DtPipe resolves it from the OS keychain at runtime. Run dtpipe secret set prod-db "pg:..." to store a secret.

DtPipe's native providers cover common sources and destinations. For everything else โ€” object storage (S3, GCS, Azure Blob), Iceberg, MySQL/MariaDB, HTTP APIs, spatial formats โ€” DuckDB's extension ecosystem serves as a connector multiplier. Load an extension with --duck-init on a DuckDB reader, writer, or --sql branch to reach any source or destination DuckDB supports natively. No additional adapters required.


Key Concepts

Transformers (--fake, --mask, --compute, --filter, โ€ฆ) chain left-to-right. When source and destination are both columnar (Parquet, DuckDB, Arrow), data flows through without row conversion. Multiple --input sources with --from, --sql, or --merge form a DAG executed concurrently. Any CLI command can be saved to a YAML job file with --export-job and replayed with --job.

DuckDB is a remarkable engine โ€” fast, self-contained, with a rich SQL dialect and a thriving extension ecosystem. DtPipe uses it as a first-class component precisely because of that quality. When DuckDB alone covers your use case, use it directly. DtPipe adds value in the scenarios it wasn't designed for: anonymizing or masking data in transit, routing one source to multiple destinations concurrently, writing to target databases with strategies like upsert, auto-migrate, or bulk insert, reading from Oracle, SQL Server, or XML streams, and packaging pipelines as repeatable YAML jobs with integrated secret management. DtPipe contributes the pipeline layer; DuckDB contributes the SQL engine.


Documentation

Document Contents
REFERENCE.md Full CLI option tables, YAML job schema, DAG topology reference, secret management
COOKBOOK.md End-to-end scenarios: anonymization, schema transforms, SQL joins, DAG pipelines, YAML automation
EXTENDING.md Adding adapters (readers/writers) and transformers

Shell Autocompletion (experimental)

dtpipe completion --install

Restart your terminal (or source ~/.zshrc) to activate.


Contributing

See EXTENDING.md for the adapter and transformer patterns.

License

MIT

Product Compatible and additional computed target framework versions.
.NET net10.0 is compatible.  net10.0-android was computed.  net10.0-browser was computed.  net10.0-ios was computed.  net10.0-maccatalyst was computed.  net10.0-macos was computed.  net10.0-tvos was computed.  net10.0-windows was computed. 
Compatible target framework(s)
Included target framework(s) (in package)
Learn more about Target Frameworks and .NET Standard.

This package has no dependencies.

Version Downloads Last Updated
1.9.0 43 10/2/2026
1.8.2 111 9/14/2026
1.8.1 110 9/10/2026
1.8.0 113 9/9/2026
1.7.0 106 9/5/2026
1.6.0 183 8/27/2026
1.5.0 169 8/6/2026
1.4.3 153 7/7/2026
1.4.2 147 6/20/2026
1.4.1 148 6/19/2026
1.4.0 167 6/17/2026
1.3.4 154 6/15/2026
1.3.3 136 6/14/2026
1.3.2 141 6/9/2026
1.3.1 152 6/8/2026
1.3.0 213 5/8/2026
1.2.6 147 4/29/2026
1.2.5 159 4/25/2026
1.2.4 149 4/14/2026
Loading failed