Skip to main content
Version: devel View Markdown

Toolkits

A toolkit is a versioned bundle of skills, rules, and an MCP server, tied together by a workflow that tells the agent which skill to run at each step and how to leverage the MCP. Each toolkit covers one job: build a REST API pipeline, add data-quality checks, deploy a workspace, and so on. Toolkits also act as guardrails, keeping the agent from diverging from proven dlt patterns and data-engineering best practices. Head to Installation to install them.

The development cycle

Toolkits map onto a five-stage pipeline lifecycle. Each toolkit owns one stage and guides the agent through it end to end, starting from an entry skill and running the rest in sequence.

StagePurposeStep-by-step guide
IngestLoad data from sources (REST APIs, SQL databases, files) into a destination.REST API source with AI Harness
ValidateDefine column-level checks and load metrics to catch bad data early.
TransformReshape raw pipeline data into a curated model for downstream use.Explore and transform your data
DeployShip pipelines and notebooks to the dltHub platform on a schedule.Deploy with AI Harness
ObserveExplore loaded data and diagnose performance issues.Explore and transform your data

Ingest

rest-api-pipeline

Build REST API pipelines with dlt. Scope, debug, and validate data. See the worked example.

Skills (entry: /find-source)
  • /find-source: Find a dlt source for a given API or data provider.
  • /create-rest-api-pipeline: Scaffold a REST API pipeline from the discovered source.
  • /debug-pipeline: Inspect traces, load packages, and schema after a run.
  • /validate-data: Validate schema and data after a successful load.
  • /view-data: Query and explore loaded data via dataset API, ibis, and ReadableRelation.
  • /adjust-endpoint: Remove dev limits, add incremental loading, and handle rate-limits.
  • /new-endpoint: Add a new endpoint to an existing pipeline.
  • /optimize-rest-api-performance: Parallelize resources, tune page size and concurrency.

sql-database-pipeline

Connect to any SQL source, load tables to a destination, and tune performance with backends.

Skills (entry: /find-source)
  • /find-source: Find and explore a SQL database source (Postgres, MySQL, MS SQL, Oracle, SQLite, any SQLAlchemy).
  • /create-sql-database-pipeline: Scaffold a pipeline from a SQL source.
  • /debug-pipeline: Diagnose connection failures, driver issues, and failed jobs.
  • /validate-data: Validate schema and column mappings after a load.
  • /view-data: Query and explore loaded data.
  • /add-table: Add a new table or view to an existing pipeline.
  • /adjust-table: Remove dev limits, configure incremental loading and merge keys.
  • /optimize-sql-performance: Pick a faster backend, tune chunk size, parallelize tables.

filesystem-pipeline

Load files (CSV, Parquet, JSONL, or custom) from local disk, S3, GCS, Azure, or SFTP into a destination.

Skills (entry: /create-filesystem-pipeline)
  • /create-filesystem-pipeline: Load files (CSV, Parquet, JSONL, or custom) from local disk, S3, GCS, Azure, or SFTP.
  • /add-incremental-loading: Filter files by modification date, switch to merge with a primary key.
  • /optimize-filesystem-performance: Faster reader, parallel reads, narrower globs, chunked streaming.

Validate

data-quality

Inspect schema for candidates, define column-level validations and load metrics, run them on every pipeline load, and diagnose failures.

Skills (entry: /setup-data-quality)
  • /setup-data-quality: Set up data-quality workflows for a pipeline.
  • /define-data-quality-checks: Translate business rules and schema hints into checks and metrics.
  • /run-data-quality: Execute defined checks against a loaded pipeline.
  • /review-data-quality: Inspect check and metric outcomes and diagnose failures.

Transform

transformations

Transform raw dlt pipeline data into a Canonical Data Model using Kimball dimensional modeling and @dlt.hub.transformation functions. See the worked example.

Skills (entry: /annotate-sources)
  • /annotate-sources: Annotate dlt pipeline sources for transformation.
  • /create-ontology: Build a business entity graph (ontology) from annotated sources.
  • /generate-cdm: Generate a Canonical Data Model in DBML using Kimball dimensional modeling.
  • /create-transformation: Emit @dlt.hub.transformation functions that map source tables to CDM entities.
  • /debug-transformation: Diagnose transformation failures, SQL dialect errors, silently dropped columns.
  • /incremental-transformation: Switch from full-replace to incremental loading.

Deploy

dlthub-platform

Deploy dltHub workspaces and pipelines to the dltHub Platform. See the worked example.

Skills (entry: /setup-runtime)
  • /setup-runtime: Verify workspace readiness for the platform (workspace file, dlt[hub] dependencies, login state).
  • /prepare-deployment: Prepare production credentials and destinations, split dev/prod credentials.
  • /deploy-workspace: Deploy pipelines and notebooks to the platform, with optional scheduling.
  • /debug-deployment: Investigate failed runs, unexpected results, and job status.

Observe

data-exploration

Connect to a pipeline, profile tables, plan charts, and assemble marimo dashboards. See the worked example.

Skills (entry: /explore-data)
  • /explore-data: Connect to a pipeline, profile tables, plan charts, write an analysis plan.
  • /build-notebook: Assemble a marimo notebook from the analysis plan and launch it.

performance

Diagnose the bottleneck stage (extract, normalize, load) and apply parallelism, workers, memory buffers, file rotation, and batching.

Skills (entry: /optimize-performance)
  • /optimize-performance: Source-agnostic tuning: parallelism, workers, memory buffers, file rotation, batching. For source-specific tuning, see the pipeline toolkit's own optimize skill.

What's next

This demo works on codespaces. Codespaces is a development environment available for free to anyone with a Github account. You'll be asked to fork the demo repository and from there the README guides you with further steps.
The demo uses the Continue VSCode extension.

Off to codespaces!

DHelp

Ask a question

Welcome to "Codex Central", your next-gen help center, driven by OpenAI's GPT-4 model. It's more than just a forum or a FAQ hub – it's a dynamic knowledge base where coders can find AI-assisted solutions to their pressing problems. With GPT-4's powerful comprehension and predictive abilities, Codex Central provides instantaneous issue resolution, insightful debugging, and personalized guidance. Get your code running smoothly with the unparalleled support at Codex Central - coding help reimagined with AI prowess.