Build lineage-aware canonical models with agents
16,000+ companies already ingest data with dlt in production; dltHub Transformations picks up right where that pipeline left off, drafting the taxonomy, ontology, and CDM your business runs on from the schemas and data you loaded with dlt.
dltHub Transformations
Your analytics engineer co-pilot. Paste this prompt into Claude/Codex/Cursor:
Run uvx dlthub-start then /annotate-sources to build my first canonical data model from data I've already loaded with dlt
"dltHub and Snowflake deliver a simple, end-to-end pathway for financial institutions to transform raw data into governed analytics and AI-ready datasets without needing a full engineering team. Whether you're an engineer stepping in to answer business questions, an analyst building your own pipelines, or someone fluent in AI with basic Python skills, you can pull data from core banking systems, market feeds, and APIs directly into Snowflake. You deliver outcomes that once required specialised data engineering resources."

Suraj Rajan
Field CTO, Financial Services, Snowflake

Suraj Rajan
Field CTO, Financial Services, Snowflake
From annotated sources to dbt models you own
Annotate the tables you already loaded and the generator writes the canonical model and the dbt models beneath it. Contracts and quality checks run against those same definitions, so a model that drifts fails before it reaches a dashboard, and your existing dbt project keeps running next to the pipelines that feed it.
Transformations
DocsAnnotate your sources and the generator writes the canonical model and the dbt models beneath it. Your existing dbt project keeps running, orchestrated next to the pipelines that feed it.
- annotate-sources14 tables tagged
- generate-cdmCanonical model in DBML
- create-transformation9 dbt models written
- incremental-transformationSwitching from full refresh
Data apps
DocsDeploy marimo notebooks and Streamlit apps on the runtime that already loads the data, so the people asking the questions get an answer instead of a ticket.
| App | Framework | Used by | |
|---|---|---|---|
| pipeline_health | marimo | Data platform | |
| revenue_explorer | marimo | Finance | |
| source_coverage | streamlit | Analytics |
Contracts
DocsDeclare what the warehouse is allowed to receive. Schema changes are accepted, blocked or quarantined by rule, on the run that introduced them.
| Schema change | Dataset | Contract says | |
|---|---|---|---|
| Column added: discount_pct | orders | Allowed, added | |
| Type changed: amount | orders | Blocked, run halted | |
| Table added: refunds | finance | Allowed, added |
Data quality
DocsDeclare expectations once. Rows that fail are quarantined before they reach the warehouse, per dataset, table or row.
Checked per dataset, table or row
Every check and chart still traces back to the load that produced it
The lifecycle closes where operations begin. The same CDM tables plug straight into dq.checks — is_unique, is_not_null, is_in — for continuous validation: no new connection, because it's the same dataset object the transformation just wrote to. If a check fails, _dlt_load_id says exactly which load produced the row.
| Run | Started | Duration | Rows | Status |
|---|---|---|---|---|
| #4128 | 13:42 · today | 4.21s | 2,214 | Running |
| #4127 | 13:12 · today | 4.04s | 2,189 | Success |
| #4126 | 12:42 · today | 4.17s | 2,202 | Success |
| #4125 | 12:12 · today | 4.31s | 2,176 | Success |
| #4124 | 11:42 · today | 42.5s | 0 | Failed |
| #4123 | 11:12 · today | 4.08s | 2,164 | Success |
import dlt
from dlt.hub import data_quality as dq
pipeline = dlt.attach(pipeline_name="person_interactions_to_cdm")
dataset = pipeline.dataset() # same dataset object the transformation just wrote to
checks = {
"fact_event_attendance": [
dq.checks.is_unique("attendance_sk"),
dq.checks.is_not_null("person_sk"),
dq.checks.is_in("status", ["registered", "attended", "no_show"]),
]
}
dq.CheckSuite(dataset, checks=checks).checksComplete agentic workflows for every phase of data engineering
Not autocomplete, not a chatbot on a dashboard. A guided sequence of skills, commands, rules, and MCP - with guardrails agents can't skip. Maintained by dltHub, controlling the infrastructure agents and pipelines operate on.
Discover individual skills per agentic workflow
See how each workflow guides your agent - step by step, from first prompt to production deployment.
The guided entry point. Names a use case, checks the workspace, and hands off to the right toolkit in a few prompts.
Opus 5.0 · Quick Start · ~/pipelines
Take me through the full workflow with the GitHub API
The guided entry point. Names a use case, checks the workspace, and hands off to the right toolkit in a few prompts.
Opus 5.0 · Quick Start · ~/pipelines
Take me through the full workflow with the GitHub API
Blueprints for data workflows
dltHub is a composable data platform. Blueprints are its ready-made builds: each one dltHub assembled for a specific use case, end to end, from the sources you already use to a production dashboard or API.
Your coding agent
Claude Code, Cursor or Codex
the agentic layerdltHub AI harness
Agentic primitives to build, run and fix pipelines
dltHub context graph
Lineage, schema, data quality, governance, run state
Pydantic Logfire
Arize
Langfuse
LangChain
the managed infrastructure layerIngest and standardize traces into the OpenAI messages format as a training-ready dataset.
distil labs
Fine-tune a specialist model, served as a drop-in replacement via API to distil labs customers.
Frequently Asked Questions
What is dlt?
dlt (data load tool) is an open-source Python library for building data pipelines. It handles schema inference, incremental loading, nested data normalization, and works with 10,100+ sources. Apache 2.0 licensed and always free to use.
What is dltHub?
dltHub is the managed agentic platform for running dlt pipelines in production. It bundles a managed runtime (deploy with one command, no infra to patch), Python and SQL transformations orchestrated inside your pipeline, data quality checks that fail fast with actionable errors, a managed Iceberg lakehouse with the option to bring your own storage, and an MCP server so agents can analyze pipelines and datasets directly. The outcome: teams ship trustworthy data faster, without owning the infrastructure. See the full feature list in the dltHub docs.
How is dltHub different from a Claude skill or tools like Replit?
Tools like Claude skills or Replit are great for writing and running code. But they are not built for data engineering workflows end to end. dltHub gives your team complete agentic workflows that cover every phase: coding, running, deploying, and debugging pipelines, on infrastructure you control.
How is dlt different from Fivetran or a Python script that uses the request library?
dlt is the perfect match between standardization and customization. You get the automation that matters: schema inference, incremental state, normalization, and loading, while keeping the full flexibility and portability of plain Python. And with agentic dltHub workflows, your team can code, run, deploy, and debug pipelines faster.
How do I get access to dltHub?
dltHub is available now. Book a demo with our team to get set up, or see our pricing page for plans and what's included.


