dtpipe 1.7.0
See the version list below for details.
dotnet tool install --global dtpipe --version 1.7.0
dotnet new tool-manifest
dotnet tool install --local dtpipe --version 1.7.0
#tool dotnet:?package=dtpipe&version=1.7.0
nuke :add-package dtpipe --version 1.7.0
DtPipe
A self-contained CLI for streaming, transforming, and anonymizing data across databases and file formats.
DtPipe reads from a source, applies row and columnar transformations in batches, and writes to a destination with no intermediate staging. It is designed for automation and CI/CD workflows where repeatable, observable data pipelines matter.
๐ Recipes & Examples โ COOKBOOK.md ยท Full CLI Reference โ REFERENCE.md
Installation
.NET Global Tool (Recommended)
dotnet tool install -g dtpipe
dtpipe --help
Build from Source
Prerequisite: .NET 10 SDK
# Bash (Mac/Linux/Windows Git Bash)
./build.sh
# PowerShell (Windows/Cross-platform)
./build.ps1
Binary created at: ./dist/release/dtpipe
Quick Start
Export a database table
dtpipe \
-i "pg:Host=localhost;Database=prod;Username=postgres" \
--query "SELECT * FROM users" \
-o users.parquet
Anonymize before export
dtpipe \
-i "pg:Host=localhost;Database=prod;Username=postgres" \
--query "SELECT * FROM users" \
--fake "email:internet.email" \
--fake "name:name.fullName" \
--mask "phone:###-****" \
--null "ssn" \
-o anonymized_users.parquet
In-memory SQL join
dtpipe \
-i orders.parquet --alias orders \
-i customers.csv --alias customers \
--from orders --ref customers \
--sql "SELECT o.*, c.name FROM orders o JOIN customers c ON o.customer_id = c.id" \
-o result.parquet
Run from a YAML job file
# Generate a reusable job file from any CLI command
dtpipe -i "pg:..." --query "SELECT * FROM users" --fake "email:internet.email" \
-o users.parquet --export-job nightly.yaml
# Run it (with optional overrides)
dtpipe --job nightly.yaml --limit 1000
Incremental loading (cursor-driven)
# Full load on first run (state file does not exist โ uses default 1970-01-01)
dtpipe \
-i "pg:Host=localhost;Database=prod" \
--query "SELECT * FROM users WHERE updated_at >= '${{cursor://state.json|1970-01-01}}'" \
-o "sqlite:Data Source=dw.db" --table "users" --strategy Recreate --key id \
--cursor "updated_at" --state "state.json"
# Subsequent runs: cursor is resolved from state.json, only newer rows are fetched (switch to Upsert)
# See REFERENCE.md#incremental-loading for full flag table and state file format,
# and COOKBOOK.md#incremental-loading for the complete recipe.
See what a pipeline will do, before it does it
# Runs the pipeline over 10 source rows with the writer neutralised: same reader, same
# transformers, same bridges. A step that expands or aggregates shows its new row count.
dtpipe -i "pg:Host=localhost;Database=prod" --query "SELECT * FROM orders" \
--mask email -o "parquet:out.parquet" --dry-run 10
# Materialise a point and iterate on it without reading the source again
dtpipe -i "oracle:..." --query "SELECT * FROM big_table" --checkpoint -o null:
dtpipe --from-checkpoint <key> --compute "total=price*qty" -o csv:out.csv
Checkpoints live in .dtpipe/, are encrypted, and expire. See
REFERENCE.md#sample-mode-and-materialisation.
Database Resilience (Retry Policy)
# Automatically retry transient database connection and timeout errors with Polly exponential backoff
dtpipe \
-i "pg:Host=localhost;Database=prod" \
--query "SELECT * FROM sales" \
-o "mssql:Server=remote_db;Database=analytics" \
--retry
Start AI Agent MCP Server (Model Context Protocol)
# Launch native MCP server over STDIO for AI assistants (Cursor, Claude Desktop, Antigravity)
dtpipe mcp
Includes tools: dry-run, suggest-pipeline, list-cursors, execute-yaml-job, and schema discovery.
Interactive AI Agent Mode
# Launch the interactive AI agent (supports Ollama & OpenAI backends)
dtpipe agent --provider openai --api-key "sk-..."
# Or run a one-shot mission
dtpipe agent "Inspect csv:invoices.csv, anonymize email, and output to jsonl:users.jsonl"
Features local Ollama auto-discovery, official OpenAI SDK integration, Spectre.Console TUI, step-by-step trajectory inspector, Spectre DAG topology rendering, and 1-click YAML pipeline export.
Providers
DtPipe detects providers from file extensions (.csv, .parquetโฆ) or explicit prefixes โ explicit prefixes are recommended to avoid ambiguity. Full table with capabilities, query requirements and Stdin/Stdout support: REFERENCE.md#providers.
| Provider family | Examples | Prefix |
|---|---|---|
| Databases | PostgreSQL, MySQL, SQLite, DuckDB, SQL Server, Oracle | pg:, mysql:, sqlite:, duck:, mssql:, ora: |
| Files | CSV, JsonL, Parquet, Arrow, XML | csv:, jsonl:, parquet:, arrow:, xml: |
| Object storage | S3-compatible, Azure Blob | s3://bucket/key.parquet, azure://container/blob.csv |
| Special | Data Gen (source), Null/Checksum (sink) | generate:N, null:, checksum: |
Use
keyring://aliasanywhere a connection string is expected. DtPipe resolves it from the OS keychain at runtime. Rundtpipe secret set prod-db "pg:..."to store a secret.
For object storage (S3, GCS, Azure Blob), Iceberg, MySQL/MariaDB, HTTP APIs, spatial formats โ use DuckDB's extension ecosystem as a connector multiplier. Load an extension with
--duck-initon any DuckDB reader, writer, or--sqlbranch โ no additional adapter required. See REFERENCE.md#provider-specific-options and COOKBOOK.md#duckdb-extensions-and-cloud-storage.
Key Concepts
- Providers โ where data comes from and goes to. DtPipe reads from databases (
pg:,mysql:,mssql:,ora:,sqlite:,duck:) and files (csv:,parquet:,jsonl:โฆ), and writes to the same set. The provider is inferred from the file extension or an explicit prefix. See Providers for the full list. - Transformers โ what happens in between. Flags like
--fake,--mask,--compute,--filter,--renameare chained left-to-right on every row. Example: anonymize, then derive a column, then filter:--fake "email:internet.email" --compute "fullName:row.first+' '+row.last" --filter "row.age>=18". See COOKBOOK.md. - DAG pipelines โ combine or split streams. Use
--aliasto name a source, then--from/--ref/--sql/--mergeto join, union or fan-out without temp files. Typical uses: enrich a stream with a lookup table, or write one source to two sinks at once. See DAG Syntax and DAG recipes. - YAML jobs โ make it repeatable. Any CLI pipeline can be saved with
--export-job pipeline.yamland replayed withdtpipe --job pipeline.yaml(CI/CD, cron, overrides via CLI). See YAML Job Schema. - Secrets โ keep credentials out of shell history.
dtpipe secret set prod-db "pg:..."stores in the OS keychain; reference askeyring://prod-dbor${{keyring://alias}}anywhere a connection string is expected. See Secret Management.
Why DuckDB? DuckDB is fast, self-contained and speaks rich SQL โ DtPipe embeds it as its SQL engine (
--sql,duck:). When DuckDB alone covers your use case, use it directly. DtPipe adds value where DuckDB stops: anonymization/masking in transit, concurrent fan-out, database write strategies (upsert, auto-migrate, bulk), Oracle/SQL Server/XML sources, and repeatable YAML jobs with secret management.
Documentation
| Document | Contents |
|---|---|
| REFERENCE.md | Full CLI option tables, YAML job schema, DAG topology reference, secret management |
| COOKBOOK.md | End-to-end scenarios: anonymization, schema transforms, SQL joins, DAG pipelines, YAML automation |
| EXTENDING.md | Adding adapters (readers/writers) and transformers |
Shell Autocompletion (experimental)
dtpipe completion --install
Restart your terminal (or source ~/.zshrc) to activate.
Contributing
See EXTENDING.md for the adapter and transformer patterns.
License
MIT
| Product | Versions Compatible and additional computed target framework versions. |
|---|---|
| .NET | net10.0 is compatible. net10.0-android was computed. net10.0-browser was computed. net10.0-ios was computed. net10.0-maccatalyst was computed. net10.0-macos was computed. net10.0-tvos was computed. net10.0-windows was computed. |
This package has no dependencies.
| Version | Downloads | Last Updated |
|---|---|---|
| 1.9.0 | 0 | 10/2/2026 |
| 1.8.2 | 110 | 9/14/2026 |
| 1.8.1 | 110 | 9/10/2026 |
| 1.8.0 | 113 | 9/9/2026 |
| 1.7.0 | 106 | 9/5/2026 |
| 1.6.0 | 183 | 8/27/2026 |
| 1.5.0 | 167 | 8/6/2026 |
| 1.4.3 | 153 | 7/7/2026 |
| 1.4.2 | 147 | 6/20/2026 |
| 1.4.1 | 148 | 6/19/2026 |
| 1.4.0 | 167 | 6/17/2026 |
| 1.3.4 | 152 | 6/15/2026 |
| 1.3.3 | 136 | 6/14/2026 |
| 1.3.2 | 141 | 6/9/2026 |
| 1.3.1 | 152 | 6/8/2026 |
| 1.3.0 | 213 | 5/8/2026 |
| 1.2.6 | 147 | 4/29/2026 |
| 1.2.5 | 159 | 4/25/2026 |
| 1.2.4 | 149 | 4/14/2026 |