dtpipe 1.2.6

There is a newer version of this package available.
See the version list below for details.
dotnet tool install --global dtpipe --version 1.2.6
                    
This package contains a .NET tool you can call from the shell/command line.
dotnet new tool-manifest
                    
if you are setting up this repo
dotnet tool install --local dtpipe --version 1.2.6
                    
This package contains a .NET tool you can call from the shell/command line.
#tool dotnet:?package=dtpipe&version=1.2.6
                    
nuke :add-package dtpipe --version 1.2.6
                    

DtPipe

A self-contained CLI for streaming and anonymizing data across databases and file formats.

DtPipe is an Arrow-native, Zero-Copy data pipeline. It reads from a source, applies transformations in columnar batches, and writes to a destination โ€” with no intermediate staging and minimal memory overhead. It targets automation and CI/CD scenarios where you need repeatable, high-performance data movement.


๐Ÿš€ See the COOKBOOK for Recipes & Examples

Anonymization guides, pipeline examples, and detailed tutorials.


Installation

dotnet tool install -g dtpipe
dtpipe --help

Build from Source

Prerequisite: .NET 10 SDK

# Bash (Mac/Linux/Windows Git Bash)
./build.sh

# PowerShell (Windows/Cross-platform)
./build.ps1

Binary created at: ./dist/release/dtpipe

Note: Pre-compiled binaries in GitHub Releases are self-contained โ€” no .NET runtime required.


Shell Autocompletion

DtPipe supports smart suggestions for bash, zsh, and powershell:

  • strategies (Append, Truncate, Upsertโ€ฆ), providers (pg:, ora:, csv:โ€ฆ), and keyring aliases (keyring://โ€ฆ)
dtpipe completion --install

Restart your terminal (or source ~/.zshrc) to activate.


Quick Reference

CLI Usage

dtpipe --input [SOURCE] --query [SQL] --output [DEST] [OPTIONS]

1. Connection Strings (Input & Output)

DtPipe detects providers from file extensions (.csv, .parquet, .duckdb, .sqlite) or explicit prefixes. Explicit prefixes are recommended to avoid ambiguity.

Provider Input Output Prefix / Format Example
DuckDB โœ… โœ… duck: duck:my.duckdb
SQLite โœ… โœ… sqlite: sqlite:data.sqlite
PostgreSQL โœ… โœ… pg: pg:Host=localhost;Database=mydb
Oracle โœ… โœ… ora: ora:Data Source=PROD;User Id=scott
SQL Server โœ… โœ… mssql: mssql:Server=.;Database=mydb
CSV โœ… โœ… csv: / .csv data.csv
JsonL โœ… โœ… jsonl: / .jsonl data.jsonl
XML โœ… โ€” xml: / .xml data.xml
Apache Arrow โœ… โœ… arrow: / .arrow data.arrow
Parquet โœ… โœ… parquet: / .parquet data.parquet
Data Gen โœ… โ€” generate: generate:1M
Null โ€” โœ… null: null:
STDIN/OUT โœ… โœ… - / {CP}:- csv (for csv:-)

Secure your connection strings: use the keyring://my-alias prefix anywhere a connection string is expected. DtPipe resolves it from the OS keychain at runtime. See Secret Management.

- is required explicitly for standard input/output. csv is shorthand for csv:-. Providers that don't support pipes (like pg:) will raise an error if given a bare name without a connection string.

2. Anonymization & Fakers

Use --fake "Col:Generator" to replace sensitive data. See COOKBOOK.md for the full reference.

Category Key Generators
Identity name.fullName, name.firstName, internet.email
Address address.fullAddress, address.city, address.zipCode
Finance finance.iban, finance.creditCardNumber
Phone phone.phoneNumber
Dates date.past, date.future, date.recent
System random.uuid, random.number, random.boolean

3. Positional CLI Option Scoping (Reader vs Writer)

Options are scoped based on their position relative to the output flag (-o):

  • Before -o: applied to the reader / pipeline globally.
  • After -o: applied to the writer, overriding reader defaults.
# Comma separator for input, semicolon for output
dtpipe -i input.csv --csv-separator "," -o output.csv --csv-separator ";"

4. CLI Options Reference

๐Ÿ”ต Core Essentials

The mandatory flags to define your pipeline.

Flag Example / Syntax Description
-i, --input -i data.csv Required. Source connection string or file path.
-q, --query -q "SELECT *..." Required (for SQL sources). The query to execute.
-o, --output -o target.parquet Required. Target connection string or file path.
--dry-run --dry-run 10 Preview N rows in the terminal without writing.

Example: dtpipe -i orders.csv -o orders.parquet

๐Ÿ“ฅ Source (Reader) Configuration

Adjust how DtPipe pulls data from the source.

Flag Example / Syntax Description
--connection-timeout 30 Connection timeout in seconds.
--query-timeout 0 Timeout in seconds (0 = no timeout).
--unsafe-query Allow non-SELECT queries (e.g. EXEC stored procs).
CSV --csv-separator "," Set separator and header usage.
XML --path "//Item" Set record selector. See hierarchical typing.
JSONL/CSV/XML --encoding ISO-8859-1, --column-types "Id:uuid", --path items Universal text-reader options (encoding, explicit types, navigation path).
๐Ÿงช Data Transformation (Mutations)

Modify the content of your rows.

Flag Example / Syntax Description
--fake "Email:internet.email" Generate fake data. Syntax: COL:Dataset.Method. See Full Dataset List.
--mask "Phone:###-****" Partial masking (# keeps, other replaces).
--null --null "Secret" Force a column value to NULL.
--overwrite "Status:Active" Set a static value for every row in a column.
--format "{First} {Last}" .NET Composite Format using source column names.
--compute "row.Age > 18" Arbitrary JS logic. Returns the last expression or explicit return.
--filter "row.Val > 100" Drop rows where the JS expression returns false.
--expand "row.Items.map..." Multi-row expansion via JS (must return an array).
--window-script Stateful batch processing on rows array. See Windowing Aggregations.
--ignore-nulls Global flag to skip transformations if input is NULL.
๐Ÿ“ Schema & Projection

Modify the structure of the output table.

Flag Example / Syntax Description
--rename "Old:New" Rename a column before writing.
--project "Id,Name" Keep ONLY these columns (Whitelist).
--drop "OldId" Remove specific columns (Blacklist).
๐Ÿ“ค Target (Writer) Configuration

Control how data is persisted to the destination.

Flag Example / Syntax Description
--strategy Upsert Append, Truncate, DeleteThenInsert, Recreate, Upsert, Ignore.
--table --table "users" Override target table name (default is export).
--key "Id,Code" PKs for Upsert/Ignore. Auto-detected from DB if omitted.
--insert-mode Bulk Standard or Bulk (high-speed insert for PG/Ora/MSSQL).
--auto-migrate Automatically ALTER TABLE to add missing columns.
--pre-exec "TRUNCATE..." SQL script to run before the pipe starts (Database writers only).
--post-exec "ANALYZE..." SQL script to run after a successful transfer (Database writers only).
โš™๏ธ Execution & Statistics

Performance tuning and automation.

Flag Example / Syntax Description
--limit --limit 1000 Stop once N rows have been processed.
--throttle --throttle 500 Limit throughput to N rows/sec (e.g. for API safety).
--sampling-rate 0.1 Probability (0.0 to 1.0) to include a row.
--batch-size 10000 Buffer size for columnar batching (default: 50,000).
--no-stats Hide progress bars and transfer statistics.
--job --job my.yaml Run a pipeline from a YAML configuration file.
--metrics-path metrics.json Write structured execution results to a JSON file.

๐Ÿ”’ Secret Management

DtPipe integrates with the OS keyring (Windows Credential Manager, macOS Keychain, Linux Secret Service) to store connection strings without exposing them in scripts, command history, or ps output.

A secret can store a full connection string including its provider prefix (e.g. pg:Host=...).

Store a Secret

dtpipe secret set prod-db "ora:Data Source=PROD;User Id=scott;Password=tiger"

Use it in a Transfer

dtpipe -i keyring://prod-db -q "SELECT * FROM users" -o users.parquet

Manage Secrets

Command Description
dtpipe secret list List all stored aliases.
dtpipe secret get <alias> Print the secret value.
dtpipe secret delete <alias> Delete a specific secret.
dtpipe secret nuke Delete ALL secrets.

๐Ÿ”€ Multi-Stream Pipelines (DAG)

You can run multiple branches in a single command using --input flags and --sql processors. Each branch can be named with --alias so that downstream steps can reference it.

A pipeline with multiple inputs or processors is assembled as a DAG and executed concurrently.

Example: In-Memory Join

dtpipe \
  -i customers.parquet --alias customers \
  -i orders.csv --alias orders \
  --from orders --ref customers \
  --sql "SELECT o.*, c.name FROM orders o JOIN customers c ON o.customer_id = c.id" \
  -o result.parquet
Option Description
--alias [NAME] Name this branch for downstream reference
--from [ALIAS[,ALIAS...]] Start a processor or fan-out branch. Accepts one or more comma-separated streaming aliases. Fan-out consumers use a single alias; multi-stream processors use multiple.
--ref [ALIAS[,ALIAS...]] Materialized reference sources for lookup/join (preloaded into memory, comma-separated). Used with --sql for JOIN queries.
--sql "[QUERY]" SQL query to execute on the upstream sources. Default engine: DuckDB.
--merge UNION ALL processor: concatenates all --from streams in order. Requires at least 2 streaming sources.

Contributing

Adding a new database adapter or transformer? See the Developer Guide.

License

MIT

Product Compatible and additional computed target framework versions.
.NET net10.0 is compatible.  net10.0-android was computed.  net10.0-browser was computed.  net10.0-ios was computed.  net10.0-maccatalyst was computed.  net10.0-macos was computed.  net10.0-tvos was computed.  net10.0-windows was computed. 
Compatible target framework(s)
Included target framework(s) (in package)
Learn more about Target Frameworks and .NET Standard.

This package has no dependencies.

Version Downloads Last Updated
1.9.0 0 10/2/2026
1.8.2 110 9/14/2026
1.8.1 110 9/10/2026
1.8.0 113 9/9/2026
1.7.0 106 9/5/2026
1.6.0 183 8/27/2026
1.5.0 167 8/6/2026
1.4.3 153 7/7/2026
1.4.2 147 6/20/2026
1.4.1 148 6/19/2026
1.4.0 167 6/17/2026
1.3.4 152 6/15/2026
1.3.3 136 6/14/2026
1.3.2 141 6/9/2026
1.3.1 152 6/8/2026
1.3.0 213 5/8/2026
1.2.6 147 4/29/2026
1.2.5 159 4/25/2026
1.2.4 149 4/14/2026
Loading failed