Schedule lineage-aware pipeline runs with agents, all in Python
16,000+ companies already ingest data with dlt in production; teams schedule, backfill, and chain their pipelines with dltHub.
Orchestrating pipelines with dltHub
Your pipeline ops co-pilot. Paste this prompt into Claude/Codex/Cursor:
Run uvx dlthub-start then /setup-runtime to take the dlt pipelines I already run locally and run them as production jobs on dltHub
On one scheduler
Everything an orchestrator was doing, in the runtime
Cron and event-driven triggers, follow-up chains, backfills and the logs that explain a red run. Defined in Python, deployed from the repo you already have.
Dependencies live on the job row
DocsSchedule, profile and upstream dependencies on one row, so you learn the dependency where you are already looking.
| Job | Type | Profile | Trigger | Next run |
|---|---|---|---|---|
load_salesforce_objects crm_raw | Batch | prod | At 06:00 UTCTag: extraction | in 4h |
load_stripe_payouts finance_raw | Batch | prod | At 06:00 UTCTag: extraction | in 4h |
build_revenue_cdm finance_core | Batch | prod | After load_salesforce_objectsAfter load_stripe_payoutsTag: transformation | — |
check_revenue_freshness finance_core | Batch | prod | After build_revenue_cdmTag: quality | — |
Scheduling
DocsCron and event-driven triggers, with follow-up chains that wait for the loads above them.
| Job | Type | Profile | Trigger | Next Run | Last Run |
|---|---|---|---|---|---|
load_salesforce_objects crm_raw | Batch | prod | At 06:00 AMTag: extraction | in 8h 9/15/2026, 6:00:00 AM GMT+2 | 16h ago |
load_stripe_payouts finance_raw | Batch | prod | At 06:00 AMTag: extraction | in 8h 9/15/2026, 6:00:00 AM GMT+2 | 16h ago |
build_revenue_cdm finance_core | Batch | prod | After load_salesforce_objectsAfter load_stripe_payoutsTag: transformation | — | 16h ago |
| Job name | Type | Profile | Status | Last run |
|---|---|---|---|---|
| github_events | pipeline | analytics | Success | 2 min ago |
| stripe_payouts | pipeline | finance | Running | now |
| hubspot_contacts | pipeline | marketing | Success | 14 min ago |
| rev_attribution_model | transform | finance | Success | 1 hr ago |
| salesforce_objects | pipeline | sales | Queued | queued |
| data_quality_checks | verify | platform | Success | 1 hr ago |
| postgres_orders_cdc | pipeline | commerce | Failed | 3 hr ago |
| snowflake_export | export | analytics | Paused | 2 days ago |
Alerting
DocsFailures, freshness breaches and recoveries reach Slack or email with the diagnosis attached.
| Event | Pipeline | Sent to | At | |
|---|---|---|---|---|
| Load failed | load_oracle_bicc | Slack · #data-alerts | 06:07 | |
| Freshness breached | finance_core | Email · data-oncall | 06:15 | |
| Recovered on retry | load_oracle_bicc | Slack · #data-alerts | 06:22 |
See it scheduled on your own pipeline
Thirty minutes with our team. We deploy a job, set a schedule and run a backfill against a source you actually use.
Load Oracle BICC into Snowflake, hourly
dlthub-router
routed to sql-database
create-sql-database-pipeline
scaffolded oracle_bicc, 14 tables
deploy-workspace
Not just ingestion
Your dbt models run on the same scheduler
Transformations are jobs like any other, triggered after the load that feeds them. That is the half most ingestion tools leave to a second system.
Transformations
DocsAnnotate your sources and the generator writes the canonical model and the dbt models beneath it. Your existing dbt project keeps running, orchestrated next to the pipelines that feed it.
- annotate-sources14 tables tagged
- generate-cdmCanonical model in DBML
- create-transformation9 dbt models written
- incremental-transformationSwitching from full refresh
Data quality
DocsDeclare expectations once. Rows that fail are quarantined before they reach the warehouse, per dataset, table or row.
Checked per dataset, table or row
Instance sizesPreview
DocsOne argument on the job decides the machine. require={"instance": {"size": "large"}} gives that job 8 vCPU and 16 GiB. Omit it and it runs small. The size lives in the pull request, not in a console someone changed last quarter.
| Size | vCPU | Memory | Budget |
|---|---|---|---|
| small | 2 | 4 GiB | 1× |
| medium | 4 | 8 GiB | 2× |
| large | 8 | 16 GiB | 4× |
| xlarge | 16 | 32 GiB | 8× |
CI/CD
DocsPipelines are Python in your repository. Deploy them from your own CI with a workspace API token, reviewed and merged like the rest of your code.
- checkoutorders-pipeline @ 8f21c4e
- dlthub profile use ciworkspace token
- dlthub deployworkspace · finance
- dlthub job runload_orders
Explore the rest of dltHub
Frequently Asked Questions
Do I have to replace Airflow to use dltHub?
No. dlt is just Python, so it runs anywhere Python runs. If your team already operates Airflow, the dlt adapter (PipelineTasksGroup) wraps a pipeline as a task group with retry policy, log routing, and a decompose option that turns each resource into its own Airflow task. Set allow_external_schedulers=True and dlt binds its incremental cursor to the DAG's data interval, so catchup backfills load exactly the right slice. dlt deploy ... airflow-composer generates the starting DAG for you. There are equivalent guides for Dagster, Prefect, and Kestra.
How do backfills work on dltHub Runtime?
A backfill is a refresh signal that propagates through the job graph, not a manual pause-clear-rerun. A job declared with refresh="always" originates the signal on every successful run; downstream jobs pass it through by default or stop it with refresh="block". Runtime clears each reachable job's completion state and resets its interval pointer, so the next run reprocesses from the start of its declared interval. Kick one off with dlthub job trigger "tag:backfill".
Why is there no platform-level retry setting?
Retry logic belongs where resumability is known: in the pipeline code. A blanket "retry 3x" above a non-idempotent load is how you double-ingest. dlt itself ships request retries and resumable loads, and execute={"timeout": ...} gives a grace period so a run can flush buffers and commit in-flight loads before a hard kill.
What is dlt?
dlt (data load tool) is an open-source Python library for building data pipelines. It handles schema inference, incremental loading, nested data normalization, and works with 10,100+ sources. Apache 2.0 licensed and always free to use.
What is dltHub?
dltHub is the managed agentic platform for running dlt pipelines in production. It bundles a managed runtime (deploy with one command, no infra to patch), Python and SQL transformations orchestrated inside your pipeline, data quality checks that fail fast with actionable errors, a managed Iceberg lakehouse with the option to bring your own storage, and an MCP server so agents can analyze pipelines and datasets directly. The outcome: teams ship trustworthy data faster, without owning the infrastructure. See the full feature list in the dltHub docs.
How is dltHub different from a Claude skill or tools like Replit?
Tools like Claude skills or Replit are great for writing and running code. But they are not built for data engineering workflows end to end. dltHub gives your team complete agentic workflows that cover every phase: coding, running, deploying, and debugging pipelines, on infrastructure you control.
How do I get access to dltHub?
dltHub is available now. Book a demo with our team to get set up, or see our pricing page for plans and what's included.


