dtpipe 1.7.0

There is a newer version of this package available.
See the version list below for details.
dotnet tool install --global dtpipe --version 1.7.0
                    
This package contains a .NET tool you can call from the shell/command line.
dotnet new tool-manifest
                    
if you are setting up this repo
dotnet tool install --local dtpipe --version 1.7.0
                    
This package contains a .NET tool you can call from the shell/command line.
#tool dotnet:?package=dtpipe&version=1.7.0
                    
nuke :add-package dtpipe --version 1.7.0
                    

DtPipe

A self-contained CLI for streaming, transforming, and anonymizing data across databases and file formats.

DtPipe reads from a source, applies row and columnar transformations in batches, and writes to a destination with no intermediate staging. It is designed for automation and CI/CD workflows where repeatable, observable data pipelines matter.


๐Ÿ“– Recipes & Examples โ†’ COOKBOOK.md ยท Full CLI Reference โ†’ REFERENCE.md


Installation

dotnet tool install -g dtpipe
dtpipe --help

Build from Source

Prerequisite: .NET 10 SDK

# Bash (Mac/Linux/Windows Git Bash)
./build.sh

# PowerShell (Windows/Cross-platform)
./build.ps1

Binary created at: ./dist/release/dtpipe


Quick Start

Export a database table

dtpipe \
  -i "pg:Host=localhost;Database=prod;Username=postgres" \
  --query "SELECT * FROM users" \
  -o users.parquet

Anonymize before export

dtpipe \
  -i "pg:Host=localhost;Database=prod;Username=postgres" \
  --query "SELECT * FROM users" \
  --fake "email:internet.email" \
  --fake "name:name.fullName" \
  --mask "phone:###-****" \
  --null "ssn" \
  -o anonymized_users.parquet

In-memory SQL join

dtpipe \
  -i orders.parquet --alias orders \
  -i customers.csv --alias customers \
  --from orders --ref customers \
  --sql "SELECT o.*, c.name FROM orders o JOIN customers c ON o.customer_id = c.id" \
  -o result.parquet

Run from a YAML job file

# Generate a reusable job file from any CLI command
dtpipe -i "pg:..." --query "SELECT * FROM users" --fake "email:internet.email" \
       -o users.parquet --export-job nightly.yaml

# Run it (with optional overrides)
dtpipe --job nightly.yaml --limit 1000

Incremental loading (cursor-driven)

# Full load on first run (state file does not exist โ†’ uses default 1970-01-01)
dtpipe \
  -i "pg:Host=localhost;Database=prod" \
  --query "SELECT * FROM users WHERE updated_at >= '${{cursor://state.json|1970-01-01}}'" \
  -o "sqlite:Data Source=dw.db" --table "users" --strategy Recreate --key id \
  --cursor "updated_at" --state "state.json"

# Subsequent runs: cursor is resolved from state.json, only newer rows are fetched (switch to Upsert)
# See REFERENCE.md#incremental-loading for full flag table and state file format,
# and COOKBOOK.md#incremental-loading for the complete recipe.

See what a pipeline will do, before it does it

# Runs the pipeline over 10 source rows with the writer neutralised: same reader, same
# transformers, same bridges. A step that expands or aggregates shows its new row count.
dtpipe -i "pg:Host=localhost;Database=prod" --query "SELECT * FROM orders" \
  --mask email -o "parquet:out.parquet" --dry-run 10

# Materialise a point and iterate on it without reading the source again
dtpipe -i "oracle:..." --query "SELECT * FROM big_table" --checkpoint -o null:
dtpipe --from-checkpoint <key> --compute "total=price*qty" -o csv:out.csv

Checkpoints live in .dtpipe/, are encrypted, and expire. See REFERENCE.md#sample-mode-and-materialisation.

Database Resilience (Retry Policy)

# Automatically retry transient database connection and timeout errors with Polly exponential backoff
dtpipe \
  -i "pg:Host=localhost;Database=prod" \
  --query "SELECT * FROM sales" \
  -o "mssql:Server=remote_db;Database=analytics" \
  --retry

Start AI Agent MCP Server (Model Context Protocol)

# Launch native MCP server over STDIO for AI assistants (Cursor, Claude Desktop, Antigravity)
dtpipe mcp

Includes tools: dry-run, suggest-pipeline, list-cursors, execute-yaml-job, and schema discovery.

Interactive AI Agent Mode

# Launch the interactive AI agent (supports Ollama & OpenAI backends)
dtpipe agent --provider openai --api-key "sk-..."

# Or run a one-shot mission
dtpipe agent "Inspect csv:invoices.csv, anonymize email, and output to jsonl:users.jsonl"

Features local Ollama auto-discovery, official OpenAI SDK integration, Spectre.Console TUI, step-by-step trajectory inspector, Spectre DAG topology rendering, and 1-click YAML pipeline export.


Providers

DtPipe detects providers from file extensions (.csv, .parquetโ€ฆ) or explicit prefixes โ€” explicit prefixes are recommended to avoid ambiguity. Full table with capabilities, query requirements and Stdin/Stdout support: REFERENCE.md#providers.

Provider family Examples Prefix
Databases PostgreSQL, MySQL, SQLite, DuckDB, SQL Server, Oracle pg:, mysql:, sqlite:, duck:, mssql:, ora:
Files CSV, JsonL, Parquet, Arrow, XML csv:, jsonl:, parquet:, arrow:, xml:
Object storage S3-compatible, Azure Blob s3://bucket/key.parquet, azure://container/blob.csv
Special Data Gen (source), Null/Checksum (sink) generate:N, null:, checksum:

Use keyring://alias anywhere a connection string is expected. DtPipe resolves it from the OS keychain at runtime. Run dtpipe secret set prod-db "pg:..." to store a secret.

For object storage (S3, GCS, Azure Blob), Iceberg, MySQL/MariaDB, HTTP APIs, spatial formats โ€” use DuckDB's extension ecosystem as a connector multiplier. Load an extension with --duck-init on any DuckDB reader, writer, or --sql branch โ€” no additional adapter required. See REFERENCE.md#provider-specific-options and COOKBOOK.md#duckdb-extensions-and-cloud-storage.


Key Concepts

  • Providers โ€” where data comes from and goes to. DtPipe reads from databases (pg:, mysql:, mssql:, ora:, sqlite:, duck:) and files (csv:, parquet:, jsonl:โ€ฆ), and writes to the same set. The provider is inferred from the file extension or an explicit prefix. See Providers for the full list.
  • Transformers โ€” what happens in between. Flags like --fake, --mask, --compute, --filter, --rename are chained left-to-right on every row. Example: anonymize, then derive a column, then filter: --fake "email:internet.email" --compute "fullName:row.first+' '+row.last" --filter "row.age>=18". See COOKBOOK.md.
  • DAG pipelines โ€” combine or split streams. Use --alias to name a source, then --from / --ref / --sql / --merge to join, union or fan-out without temp files. Typical uses: enrich a stream with a lookup table, or write one source to two sinks at once. See DAG Syntax and DAG recipes.
  • YAML jobs โ€” make it repeatable. Any CLI pipeline can be saved with --export-job pipeline.yaml and replayed with dtpipe --job pipeline.yaml (CI/CD, cron, overrides via CLI). See YAML Job Schema.
  • Secrets โ€” keep credentials out of shell history. dtpipe secret set prod-db "pg:..." stores in the OS keychain; reference as keyring://prod-db or ${{keyring://alias}} anywhere a connection string is expected. See Secret Management.

Why DuckDB? DuckDB is fast, self-contained and speaks rich SQL โ€” DtPipe embeds it as its SQL engine (--sql, duck:). When DuckDB alone covers your use case, use it directly. DtPipe adds value where DuckDB stops: anonymization/masking in transit, concurrent fan-out, database write strategies (upsert, auto-migrate, bulk), Oracle/SQL Server/XML sources, and repeatable YAML jobs with secret management.


Documentation

Document Contents
REFERENCE.md Full CLI option tables, YAML job schema, DAG topology reference, secret management
COOKBOOK.md End-to-end scenarios: anonymization, schema transforms, SQL joins, DAG pipelines, YAML automation
EXTENDING.md Adding adapters (readers/writers) and transformers

Shell Autocompletion (experimental)

dtpipe completion --install

Restart your terminal (or source ~/.zshrc) to activate.


Contributing

See EXTENDING.md for the adapter and transformer patterns.

License

MIT

Product Compatible and additional computed target framework versions.
.NET net10.0 is compatible.  net10.0-android was computed.  net10.0-browser was computed.  net10.0-ios was computed.  net10.0-maccatalyst was computed.  net10.0-macos was computed.  net10.0-tvos was computed.  net10.0-windows was computed. 
Compatible target framework(s)
Included target framework(s) (in package)
Learn more about Target Frameworks and .NET Standard.

This package has no dependencies.

Version Downloads Last Updated
1.9.0 0 10/2/2026
1.8.2 110 9/14/2026
1.8.1 110 9/10/2026
1.8.0 113 9/9/2026
1.7.0 106 9/5/2026
1.6.0 183 8/27/2026
1.5.0 167 8/6/2026
1.4.3 153 7/7/2026
1.4.2 147 6/20/2026
1.4.1 148 6/19/2026
1.4.0 167 6/17/2026
1.3.4 152 6/15/2026
1.3.3 136 6/14/2026
1.3.2 141 6/9/2026
1.3.1 152 6/8/2026
1.3.0 213 5/8/2026
1.2.6 147 4/29/2026
1.2.5 159 4/25/2026
1.2.4 149 4/14/2026
Loading failed