The agentic data layer

dltHub is an AI-native data engineering platform for teams of humans and agents. Your team and their agents generate and manage high-quality data at scale, on infrastructure we run for you.

For

Hundreds of pipelines, one team

Any engineer on your team can ship production data, with agents doing the work on infra we run. Every run is logged and auditable.

Book a demoExplore BlueprintsAny source · Keep your warehouse · Migrate in weeks

Used by data teams at

  • Arrington Capital
  • Lightdash
  • Pro Juventute

Pick a logo to read how they did it, or browse every case study.

Blueprints for data workflows

dltHub is a composable data platform. Blueprints are its ready-made builds: each one dltHub assembled for a specific use case, end to end, from the sources you already use to a production dashboard or API.

  • Every teamBrowse every dltHub Blueprint

Your coding agent

Claude Code, Cursor or Codex

operates through
dltHubthe agentic layer

dltHub AI harness

Agentic primitives to build, run and fix pipelines

dltHub context catalog

Lineage, schema, data quality, governance, run state

grounds and runs
Trace pipelines
  • Pydantic Logfire

    Pydantic Logfire

  • Arize

    Arize

  • Langfuse

    Langfuse

  • LangChain

    LangChain

dltHubthe managed infrastructure layer

Ingest and standardize traces into the OpenAI messages format as a training-ready dataset.

API
distil labs

distil labs

Fine-tune a specialist model, served as a drop-in replacement via API to distil labs customers.

View the Agent distillation with distil labs blueprint

"What I didn't expect is how much it unblocks the team. A mid-level engineer can spin up a prototype, browse the raw data in dltHub's local DuckDB workspace, validate the SQL schema - all without pulling in a senior. That loop of prototype, inspect, fix, re-run - that's the real unlock."

Marcello Victorino

Marcello Victorino

Staff Data Engineer, Tasman Analytics

Describe it. The agent ships it. Your team runs it.

Available in dlt
Claude Code·~/crm-pipeline
>

Build a pipeline that loads CRM contacts and deals into my warehouse using dlt

The dltHub AI harness: guardrails your coding agent can't skip

Not autocomplete, not a chatbot on a dashboard. The harness is the skills, commands, rules, and MCP tools that walk Claude Code, Codex, or Cursor through building a pipeline and then keeping it running in production: scheduled runs, schema changes, failures. Maintained by dltHub and wired to the infrastructure your pipelines run on, so your data team ships without a platform team.

Start
Quick Startdlt1 skill · 1 cmd · 1 rule · MCP
Initdlt3 skills · 1 cmd · 1 rule · MCP
Ingest
REST API Pipelinedlt8 skills · 1 rule · MCP
SQL Database Pipelinedlt8 skills · 1 rule · MCP
Filesystem Pipelinedlt3 skills · 1 rule · MCP
Harden
Data QualitydltHub4 skills · 2 rules · MCP
Performancedlt1 skill · 1 rule · MCP
Transform
TransformationsdltHub6 skills · 1 rule · MCP
Explore
Data Explorationdlt2 skills · 1 rule
Operate
dltHub PlatformdltHub4 skills · 2 rules

Every skill in the harness

What your agent reaches for, step by step, from first prompt through the pipelines it keeps running.

Quick Start1/1

The guided entry point. Names a use case, checks the workspace, and hands off to the right toolkit in a few prompts.

Opus 5.0 · Quick Start · ~/pipelines

>
? for shortcuts
Ask your agentcopies with harness setup

Take me through the full workflow with the GitHub API

For regulated life sciences · GxP

GxP validation that survives your next upgrade.

dlt is code-native, test-first ingestion; dltHub runs it. Together they satisfy CSA and GAMP 5 without freezing versions or exposing you to platform churn.

  • GxP
  • CSA
  • GAMP 5
  • 21 CFR Part 11
  • EU Annex 11
  • ALCOA+
Your git tag is the validated baseline
Python in your repo. You choose when to upgrade and which version runs.
Deterministic by design
Same input, same output, re-run on every change.
ALCOA+ evidence built in
Lineage, schema contracts, and run state in the context catalog.
Read-only source access
Enforced in pipeline code for Veeva Vault and LIMS, not by policy.
Snowflake stays your orchestrator
dltHub runs the pipeline; your control plane does not move.

The validation workflow, in three commands

  1. Pin the validated baseline

    git tag validated/veeva-vault@1.4.0

    tag created · this exact code is what you validate

  2. Prove it behaves the same every run

    pytest tests/behavioural/

    214 passed · same input, same output

  3. Open the evidence trail

    dlthub show

    run history, schema contracts, lineage · the ALCOA+ record

Snowflake Industry Competency badge for Healthcare and Life Science

Snowflake Industry Competency

Recognised for healthcare and life sciences data workloads on Snowflake.

Julien Chaumond

As a simple-to-use Python library, dlt is the first tool that this new wave of people can use. By leveraging this library, we can extend the machine learning revolution into enterprise data.

Julien Chaumond

Julien Chaumond

CTO/Co-Founder at Hugging Face

Maximilian Eber

Python and machine learning under security constraints are key to our success. dlt is a lightweight yet powerful open source tool we can run together with Snowflake. Now anyone who knows Python can self-serve to fulfil their data needs.

Maximilian Eber

Maximilian Eber

CPTO & Co-Founder at Taktile

I am building our internal skills usage leaderboard so we can see how people are using AI and spread what's working. dltHub turns this into a cohesive process without messy scripts to dedupe queries or wrangle intermediate tables. And the best part is anyone on the team can use agents to easily contribute.

Nate Sesti

Nate Sesti

Co-Founder & CTO at Continue

It's easy to write AI-assisted code and get a prototype. dltHub instead is built for the inspect-validate-debug loop, which is where AI-assisted data engineering actually lives or dies. The first platform I've seen that treats that loop as a design assumption, not a feature.

Savin Goyal

Savin Goyal

CTO & Co-Founder at Outerbounds (acquired by Anaconda)

The modeling layer is where the real cost lives, and it's where consulting companies like ours have been doing the most repetitive, high-cost manual work. We are excited about the dltHub product vision for the modeling layer because the main challenge in all data projects is how the data is being interpreted.

Thomas in't Veld

Thomas in't Veld

CEO at Tasman Analytics

In dltHub, agents propose transformations in Python, dlt.hub.transformation compiles them to SQL through Ibis, and execution stays inside DuckDB. The data doesn't move into Python memory at any step — the same composition pattern Ibis and Arrow are designed around.

Wes McKinney

Wes McKinney

Builder of Apache Arrow, pandas, Ibis at Posit

Every Ops leader is trying to answer the same question: is my team getting better or worse? I always had good intuition about what data I needed, but never the resources to get it and measure it reliably. dltHub changes the equation: it's the first product I've seen built for operators and their agents, not just fully-staffed enterprise data teams.

Jacob Matson

Jacob Matson

Developer Advocate at MotherDuck

dltHub and Snowflake deliver a simple, end-to-end pathway for financial institutions to transform raw data into governed analytics and AI-ready datasets without needing a full engineering team. You can pull data from core banking systems, market feeds, and APIs directly into Snowflake.

Suraj Rajan

Suraj Rajan

Field CTO, Financial Services at Snowflake

dltHub lets my agents develop data pipelines locally, test changes quickly and cheaply in CI, and then runs them in the cloud against my largest workloads. It gives them the tools to take care of the knucklehead stuff so that I can get a good night's sleep.

Josh Wills

Josh Wills

Member of Technical Staff at DatologyAI

We weren't looking for a custom rebuild, and we didn't want to staff up a data team to run something we couldn't sustain at our size. The toolkit by dltHub took what we already had, structured the parts that needed cleaning up, and gave us back a stack we can confidently run, with Chat-BI on top that actually understands our business.

Martin Miodownik

Martin Miodownik

CTO & Co-Founder at NAVIT

We needed analytics on our GitHub Actions CI. With dltHub we wired the end-to-end pipeline up in hours, not days. Now we see where builds drag, fix the slow spots, and ship faster.

Simon Rosenberger

Simon Rosenberger

Head of Data Platform at Tower.dev

The quality of AI output depends on the frameworks and context used. This is where the dltHub AI Workbench shines. Ingestion used to sit with me and one other person. Now our analytics engineers self-serve.

Bijan Soltani

Bijan Soltani

Founder & Managing Director at Gemma Analytics

Ingestion moved from being owned by a small group with deep tool knowledge to something any Python developer on the team could author, review, and ship. The question changed from "who knows the tool?" to "what data do we need next?"

Euan Johnston

Euan Johnston

Senior Analytics Engineer at dentolo

I just wanted to reiterate how damned cool dlthub and the agentic workflows continue to be. We have a mountain of Klaviyo data, and API rate limits made extraction slow and prone to failure. The agent devised a tiered chunking strategy: 7 days, dropping to 1 day, then 1 hour on timeout, and it's gotten through some highly problematic data patterns without missing a beat. Super impressive, you guys are amazing!

Jim Barlow

Jim Barlow

Senior Analytics Engineer at Pinter

For financial services

Pipelines your auditors can actually audit.

dlt pipelines are code-native and version-controlled, and the dltHub context catalog carries the lineage. Mapped to SR 11-7, SOX ITGC, and BCBS 239.

  • SR 11-7
  • SOX ITGC
  • BCBS 239
  • SEC/FINRA 17a-4
  • DORA
Change control auditors already know
Git history, CI gates, pinned releases. No click-ops.
Reproducible for SR 11-7
Versioned, tested pipelines with the documentation model validation asks for.
End-to-end lineage for BCBS 239
Source to RAW to model-ready, traceable at every hop.
Straight into Snowflake
Core banking systems, market feeds, and APIs.
One artefact, one evidence trail
Validate the pipeline; the catalog holds the evidence.

The validation workflow, in three commands

  1. Pin the validated baseline

    git tag validated/market-feed@2.1.0

    tag created · this exact code is what auditors review

  2. Prove it behaves the same every run

    pytest tests/behavioural/

    187 passed · reproducible, deterministic output

  3. Open the evidence trail

    dlthub show

    run history, schema contracts, lineage · your audit record

Snowflake Industry Competency badge for Financial Services

Snowflake Industry Competency

Recognised for financial services data workloads on Snowflake.

Move your data stack forward without the vendor tax

Agent-led migration off your legacy vendor, 90% faster.

A months-long migration is what keeps teams tied to legacy tools they have outgrown, and what vendors count on to keep you locked in. We make it the easy part. Our engineers, armed with internal AI tooling, rebuild your Fivetran, Airbyte, and custom Python pipelines as clean dltHub code. We hand over pipelines running in production and coach your team on AI-forward data engineering.

dlt is Apache 2.0 licensed and always free to use

Use dlt as your open-source ingestion foundation and move to dltHub when you need managed runtime, observability, and governed collaboration at scale.

Apache 2.0

pip install dlt
  • Forever free OSS library
  • Apache 2.0 licensed for commercial usage
  • Code-first developer experience
  • Portable workflows across self-managed and managed runtime

dlt is the leading open-source Python library for building data pipelines using code and agents.

6M+
PyPI downloads per month
10,000+
Companies loading data into databases with dlt in production
1,000+
Companies loading into Snowflake with dlt in production

Learn agentic data engineering

Free, self-paced course. From first prompt to production deployment.

DLTHUB CONTEXT

The context your agent needs to ship any pipeline

dltHub's agentic workflows come with a REST API toolkit that taps directly into dltHub Context - a hub of deeply researched, enriched context on REST APIs across SaaS sources, databases, and destinations. Your agent pulls exactly what it needs to code any dlt pipeline, in minutes.

We already cover more than 10,100 sources, with a clear path to hundreds of thousands. From prompt to pipeline to live reports in a notebook - all in one agentic flow, with outputs tailored to data users.

Frequently Asked Questions

How is dltHub different from a Claude skill or tools like Replit?

Tools like Claude skills or Replit are great for writing and running code. But they are not built for data engineering workflows end to end.

dltHub gives your team complete agentic workflows that cover every phase: coding, running, deploying, and debugging pipelines, on infrastructure you control. Not just a skill, not just an editor, but a guided workflow from first line to production.

How is dlt different from Fivetran or a Python script that uses the request library?

dlt is the perfect match between standardization and customization. You get the automation that matters: schema inference, incremental state, normalization, and loading, while keeping the full flexibility and portability of plain Python.

And with agentic dltHub workflows, your team can code, run, deploy, and debug pipelines faster, with the reliability you can trust at every step.

What is dltHub?

dltHub is the managed platform for deploying and operating data pipelines built with dlt. It provides a runtime, observability, data quality checks, and collaboration features so teams can go from development to production with one command.

What is dlt?

dlt (data load tool) is an open-source Python library for building data pipelines. It lets you write any connector, run anywhere, and requires no backend. dlt is Apache 2.0 licensed and always free to use.