Schedule lineage-aware pipeline runs with agents, all in Python

16,000+ companies already ingest data with dlt in production; teams schedule, backfill, and chain their pipelines with dltHub.

Orchestrating pipelines with dltHub

Your pipeline ops co-pilot. Paste this prompt into Claude/Codex/Cursor:

Run uvx dlthub-start then /setup-runtime to take the dlt pipelines I already run locally and run them as production jobs on dltHub

On one scheduler

Everything an orchestrator was doing, in the runtime

Cron and event-driven triggers, follow-up chains, backfills and the logs that explain a red run. Defined in Python, deployed from the repo you already have.

Dependencies live on the job row

Docs

Schedule, profile and upstream dependencies on one row, so you learn the dependency where you are already looking.

JobTypeProfileTriggerNext run
load_salesforce_objects
crm_raw
Batchprod
At 06:00 UTCTag: extraction
in 4h
load_stripe_payouts
finance_raw
Batchprod
At 06:00 UTCTag: extraction
in 4h
build_revenue_cdm
finance_core
Batchprod
After load_salesforce_objectsAfter load_stripe_payoutsTag: transformation
—
check_revenue_freshness
finance_core
Batchprod
After build_revenue_cdmTag: quality
—
build_revenue_cdm runs only after both loads land.

Scheduling

Docs

Cron and event-driven triggers, with follow-up chains that wait for the loads above them.

JobTypeProfileTriggerNext RunLast Run
load_salesforce_objects
crm_raw
Batchprod
At 06:00 AMTag: extraction
in 8h
9/15/2026, 6:00:00 AM GMT+2
16h ago
load_stripe_payouts
finance_raw
Batchprod
At 06:00 AMTag: extraction
in 8h
9/15/2026, 6:00:00 AM GMT+2
16h ago
build_revenue_cdm
finance_core
Batchprod
After load_salesforce_objectsAfter load_stripe_payoutsTag: transformation
—
16h ago
dlthub workspace deploy
Jobs
All jobs defined in this workspace
Job nameTypeProfileStatusLast run
github_eventspipelineanalyticsSuccess2 min ago
stripe_payoutspipelinefinanceRunningnow
hubspot_contactspipelinemarketingSuccess14 min ago
rev_attribution_modeltransformfinanceSuccess1 hr ago
salesforce_objectspipelinesalesQueuedqueued
data_quality_checksverifyplatformSuccess1 hr ago
postgres_orders_cdcpipelinecommerceFailed3 hr ago
snowflake_exportexportanalyticsPaused2 days ago
Run #4128
github_events
Status
Running
Started
13:42:01 UTC
Rows loaded
2,214
Live
pipeline · github_events

Alerting

Docs

Failures, freshness breaches and recoveries reach Slack or email with the diagnosis attached.

EventPipelineSent toAt
Load failedload_oracle_biccSlack · #data-alerts06:07
Freshness breachedfinance_coreEmail · data-oncall06:15
Recovered on retryload_oracle_biccSlack · #data-alerts06:22
Recovered on retry. Nobody was paged.

See it scheduled on your own pipeline

Thirty minutes with our team. We deploy a job, set a schedule and run a backfill against a source you actually use.

dltHub harness

Load Oracle BICC into Snowflake, hourly

dlthub-router

routed to sql-database

create-sql-database-pipeline

scaffolded oracle_bicc, 14 tables

deploy-workspace

Not just ingestion

Your dbt models run on the same scheduler

Transformations are jobs like any other, triggered after the load that feeds them. That is the half most ingestion tools leave to a second system.

Transformations

Docs

Annotate your sources and the generator writes the canonical model and the dbt models beneath it. Your existing dbt project keeps running, orchestrated next to the pipelines that feed it.

Model generator
  • annotate-sources14 tables tagged
  • generate-cdmCanonical model in DBML
  • create-transformation9 dbt models written
  • incremental-transformationSwitching from full refresh
Build a revenue model from these five sourcesAgent

Data quality

Docs

Declare expectations once. Rows that fail are quarantined before they reach the warehouse, per dataset, table or row.

Checked per dataset, table or row

38 rows held back. Load shipped.

Instance sizesPreview

Docs

One argument on the job decides the machine. require={"instance": {"size": "large"}} gives that job 8 vCPU and 16 GiB. Omit it and it runs small. The size lives in the pull request, not in a console someone changed last quarter.

SizevCPUMemoryBudget
small24 GiB1×
medium48 GiB2×
large816 GiB4×
xlarge1632 GiB8×
A one-hour large run costs four hours of budget.

CI/CD

Docs

Pipelines are Python in your repository. Deploy them from your own CI with a workspace API token, reviewed and merged like the rest of your code.

Deploy from
  • GitHub
  • GitLab
Deploy job
  • checkoutorders-pipeline @ 8f21c4e
  • dlthub profile use ciworkspace token
  • dlthub deployworkspace · finance
  • dlthub job runload_orders
dlthub deploy --profile ci

Frequently Asked Questions

Do I have to replace Airflow to use dltHub?

No. dlt is just Python, so it runs anywhere Python runs. If your team already operates Airflow, the dlt adapter (PipelineTasksGroup) wraps a pipeline as a task group with retry policy, log routing, and a decompose option that turns each resource into its own Airflow task. Set allow_external_schedulers=True and dlt binds its incremental cursor to the DAG's data interval, so catchup backfills load exactly the right slice. dlt deploy ... airflow-composer generates the starting DAG for you. There are equivalent guides for Dagster, Prefect, and Kestra.

How do backfills work on dltHub Runtime?

A backfill is a refresh signal that propagates through the job graph, not a manual pause-clear-rerun. A job declared with refresh="always" originates the signal on every successful run; downstream jobs pass it through by default or stop it with refresh="block". Runtime clears each reachable job's completion state and resets its interval pointer, so the next run reprocesses from the start of its declared interval. Kick one off with dlthub job trigger "tag:backfill".

Why is there no platform-level retry setting?

Retry logic belongs where resumability is known: in the pipeline code. A blanket "retry 3x" above a non-idempotent load is how you double-ingest. dlt itself ships request retries and resumable loads, and execute={"timeout": ...} gives a grace period so a run can flush buffers and commit in-flight loads before a hard kill.

What is dlt?

dlt (data load tool) is an open-source Python library for building data pipelines. It handles schema inference, incremental loading, nested data normalization, and works with 10,100+ sources. Apache 2.0 licensed and always free to use.

What is dltHub?

dltHub is the managed agentic platform for running dlt pipelines in production. It bundles a managed runtime (deploy with one command, no infra to patch), Python and SQL transformations orchestrated inside your pipeline, data quality checks that fail fast with actionable errors, a managed Iceberg lakehouse with the option to bring your own storage, and an MCP server so agents can analyze pipelines and datasets directly. The outcome: teams ship trustworthy data faster, without owning the infrastructure. See the full feature list in the dltHub docs.

How is dltHub different from a Claude skill or tools like Replit?

Tools like Claude skills or Replit are great for writing and running code. But they are not built for data engineering workflows end to end. dltHub gives your team complete agentic workflows that cover every phase: coding, running, deploying, and debugging pipelines, on infrastructure you control.

How do I get access to dltHub?

dltHub is available now. Book a demo with our team to get set up, or see our pricing page for plans and what's included.