# dlt — data load tool & dltHub > dlt is the open-source Python library for moving data from any source to any destination. It automatically infers schemas, normalizes nested JSON, handles incremental loads, and evolves types as sources change — so data engineers can ship pipelines in minutes instead of weeks. dltHub extends the library with a managed workspace for building, deploying, and operating pipelines at team scale. This page follows the [llms.txt convention](https://llmstxt.org) and is a concise, link-first summary aimed at LLMs and agentic coding tools. For a full, auto-generated index of every documentation page — kept in sync on every docs deploy — use the docs llms.txt linked below rather than hand-curating documentation links here. The marketing, blog, case study, and ecosystem sections below are generated from the dltHub CMS on every deploy/revalidate, so they stay in sync with the website automatically. ## Documentation - [dlt docs — full index (llms.txt)](https://dlthub.com/docs/llms.txt): Auto-generated, always up-to-date list of every documentation page. Start here for anything code- or API-related. - [dlt docs home](https://dlthub.com/docs/intro): Human-facing docs entry point. - [REST API tutorial](https://dlthub.com/docs/tutorial/rest-api): Build a pipeline against any REST API in minutes. - [SQL database tutorial](https://dlthub.com/docs/tutorial/sql-database): Load from any SQL database using the dlt SQL source. - [Filesystem / object storage tutorial](https://dlthub.com/docs/tutorial/filesystem): Load JSON, JSONL, CSV, and Parquet from S3, GCS, Azure, or local disk. ## Pages - [About dltHub](https://dlthub.com/about): Meet the team behind dltHub. We build the infrastructure for AI-native data engineering. Our mission is to make Python practitioners autonomous when they create and use datasets in their organizations. - [Blog](https://dlthub.com/blog): See the latest articles from dltHub. - [Case Studies | dltHub](https://dlthub.com/case-studies): How do data engineers use dlt (data load tool)? Dive into our case studies to see how data engineering teams solve problems - [Contact us](https://dlthub.com/contact): Have questions or want to explore how our paid offerings can help your organization? Our solutions engineers are ready to assist. - [Cookie Policy](https://dlthub.com/cookies) - [Book a dltHub demo](https://dlthub.com/demo): Book a live dltHub demo. See how agents build, deploy, and run dlt pipelines on dltHub, mapped to your stack. - [Data Processing Agreement | dltHub](https://dlthub.com/dpa): Data Processing Agreement between dltHub (ScaleVector GmbH) and Subscribers governing the processing of personal data under GDPR. Effective May 26, 2026. - [dltHub Education](https://dlthub.com/education): Self-paced courses and certifications to master dlt, from your first pipeline to advanced production deployments. - [dltHub Events](https://dlthub.com/events): Workshops, conferences, meetups, and webinars where the dltHub team is presenting or sponsoring. Browse upcoming events. - [The dltHub AI harness](https://dlthub.com/features/ai-harness): The dltHub AI harness equips Claude Code, Codex, and Cursor with the skills, platform tools, and run context to build and operate dlt pipelines in production. - [dlt Imprint](https://dlthub.com/imprint) - [dltHub for financial services](https://dlthub.com/industries/financial-services): Audit-ready dlt pipelines your Model Risk and SOX teams sign off on. Mapped to SR 11-7, SOX ITGC, BCBS 239, SEC/FINRA 17a-4, and DORA. - [dltHub for regulated life sciences (GxP)](https://dlthub.com/industries/life-sciences): Validated dlt pipelines that live in your SDLC. Code-native, test-first ingestion mapped to CSA, GAMP 5, 21 CFR Part 11, EU Annex 11, and ALCOA+. - [EULA](https://dlthub.com/legal/dlt-plus-eula) - [dltHub AI Source Available License](https://dlthub.com/license): The dltHub AI Source Available License governing the use of dltHub AI Workbench and related licensed materials. - [Consulting Partner Program](https://dlthub.com/partners/consulting): Join the dltHub Consulting Partner Program. Get co-marketing support, delivery resources, training, and referral opportunities across Gold, Silver, Bronze, and Affiliate tiers. - [dlt and Databricks](https://dlthub.com/partners/databricks): Move data to Databricks with dlt, the open-source Python library. Add custom sources, integrate Unity Catalog with Delta and Iceberg, and power your AI workflows. - [dltHub and Snowflake](https://dlthub.com/partners/snowflake): 1,000+ companies already load production data into Snowflake with dlt. Agents now build 91% of those pipelines. With dltHub, a single developer can go from writing pipeline code to delivering reports for the business end user in 10 minutes. - [dltHub Pricing](https://dlthub.com/pricing): dltHub pricing plans - [Privacy Policy | dltHub](https://dlthub.com/privacy-policy): Learn how dltHub collects, uses, and protects your personal information when you use our services. - [dlt: the data loading library for Python](https://dlthub.com/product/dlt): dlt is the open source Python library for moving data from any API, database, or file into any warehouse, lake, or vector store. dltHub deploys, monitors, and scales the pipelines you write. - [Migrate off your legacy data vendor 90% faster](https://dlthub.com/product/solutions-engineering): Our engineers rebuild your Fivetran, Airbyte, and hand-rolled Python pipelines as dltHub code, validate each one against your live output before cutover, and hand them over running. Your team keeps the AI-forward workflow. - [dltHub | Agentic Data Engineering Platform](https://dlthub.com/products/dlthub): dltHub is the agentic platform that deploys, monitors, and scales dlt pipelines. Complete agentic workflows for every phase of data engineering. - [dltHub Transformations | Agentic Analytics Engineering](https://dlthub.com/products/transformations): 16,000+ companies already ingest data with dlt in production; dltHub Transformations picks up right where that pipeline left off, drafting the taxonomy, ontology, and CDM your business runs on from the schemas and data you loaded with dlt. - [Terms of Use | dltHub](https://dlthub.com/terms): dltHub Terms of Use governing access to and use of the dltHub Service by enterprise subscribers. ## Blog Most recent 20 posts. Full archive: [dlthub.com/blog](https://dlthub.com/blog) · [RSS](https://dlthub.com/blog/rss.xml). - [dltHub: dlt made agents good at building pipelines. Now they're safe enough to run for your whole team.](https://dlthub.com/blog/dlthub-for-teams): dlt made agents good at building pipelines. dltHub makes their work safe to run: agentic alerts, team workspaces, and managed infrastructure for your whole team. - [Validating Trust in Data: How dltHub Delivers GxP-Ready, Auditable Ingestion into Snowflake](https://dlthub.com/blog/validating-trust-in-data-how-dlthub-delivers-gxp-ready-auditable-ingestion-into-snowflake): In regulated life sciences, "can you explain your data?" has moved from a QA question to a board-level one. dltHub turns ingestion into validated, testable code - built on dlt, run on dltHub - so your team and their agents produce data that lands in Snowflake with the lineage, tests, and evidence regulators expect. - [Transformations that know their context](https://dlthub.com/blog/transformations-that-know-their-context): dlthub ships 2 ways to run transformations - you can bring your own dbt project or use dlthub native transformations. This posts aims to explain to practitioners and technical decision makers how the two methods are different. Let’s dive in! - [Composable canonicals: from tribal data knowledge to versioned artifact](https://dlthub.com/blog/composable-canonicals): In this article, I'm exploring an AI-native architecture that aims to solve the traditional central data team problem. - [dlt vs dltHub: the five ecosystem layers dltHub adds](https://dlthub.com/blog/from-dlt-to-dlthub): dlt handles ingestion. dltHub adds the five ecosystem layers you'd otherwise build yourself: AI harness, context catalog, transformation, orchestration, managed infra. - [Moving off Airbyte: Regaining control over your data movement.](https://dlthub.com/blog/migrate-from-airbyte-to-dlthub): Airbyte promises a self serve UI, a catalog of connectors, no code. Sounds great on paper, but in practice this vision is plagued by skill barriers, reliability issues, and high cost of ownership. - [Productionize Python ETL Scripts: Migrate to dlt & dltHub](https://dlthub.com/blog/productionize-python-etl-scripts): Schema breaks, OOM deaths, duplicate loads, bus factor of one. What it takes to make hand-rolled pipelines production-grade, and what it costs to run them after. - [Migrating from Fivetran: how the move actually works](https://dlthub.com/blog/migrate-fivetran-to-dlthub): What does the move do to my bill and my renewal? And how to move? - [Consulting is becoming software.](https://dlthub.com/blog/consulting-is-becoming-software): dltHub ships four Migration Blueprints: Python scripts, dlt, Fivetran, and Airbyte to dltHub. LLM-native tooling turns multi-month, senior-only migration projects into weeks of work at a fraction of the cost. - [We are launching two dltHub Blueprints for agent spend: Agent Cost & Usage to understand it, Agent Distillation to optimize it](https://dlthub.com/blog/agent-cost-distillation): dltHub launches two Blueprints for agent spend: Agent Cost & Usage to break down what each model, person, and customer costs, and Agent Distillation (with distil labs) to replace expensive agents with cheaper specialist models. - [dltHub Blueprints: launch FAQ](https://dlthub.com/blog/blueprints-faq): Answers to the most common questions about dltHub Blueprints: why we are shipping them, why agent spend is hard to measure, how a Blueprint works, and how to get involved. - [How I Tracked the Tech Job Market as a Working Student (And What I Found)](https://dlthub.com/blog/tracking-tech-job-market-with-dlt): I analyzed 987 job posts to see what the data job market really wants. Built with dlt + Claude in one conversation, deployed to auto-update every month. - [Build vs buy is over. A connector now costs $100 a year.](https://dlthub.com/blog/tco): For the last two decades, saas connector companies told you data meant choosing between “build or buy”. Today, agentic building killed that narrative, - [Agents that remember: cognee 1.0 is out](https://dlthub.com/blog/cognee-1-0): cognee 1.0 is live: open-source memory for AI agents, now self-improving, with a Rust core and single-Postgres deployment. Cognee is a dltHub partner that uses dlt under the hood. - [What is the dltHub Context Layer?](https://dlthub.com/blog/context): For years, the thing that held a data pipeline together end to end wasn't a tool. It was you. You were the context layer. - [The rise of the Semantic engineer](https://dlthub.com/blog/the-rise-of-the-knowledge-engineer): Agents now write the pipelines, models, and dashboards. What they can't write is what your data means. Meet the data role that's emerging: the semantic engineer. - [Text-to-SQL is a definition problem: build the canonical model first](https://dlthub.com/blog/canonical-text-to-sql): Text-to-SQL doesn’t break because models can’t write SQL — it breaks because they don’t know what your data means. Write the meaning down first as a canonical knowledge layer, and use that one spec to both build the model and answer questions over it. - [The LLM got the right answer for the wrong reason](https://dlthub.com/blog/ontology-benchmark): Schema alone scored 3/10. An ontology scored 10/10. A benchmark across two datasets showing exactly where the gap is, including cases where the model gets the right answer for the wrong reason. - [Schema evolution in data pipelines: the engineer's guide](https://dlthub.com/blog/schema-evolution-guide): Schema evolution is a decision every data pipeline makes — most tools make it silently. This post discusses the five common failure modes every data pipeline sees, how dlt handles them, and how you can decide runtime policies for schema evolution with data contracts. - [dltHub Named 2026 Snowflake Startup Program Product Partner of the Year](https://dlthub.com/blog/dlthub-2026-snowflake-startup-program-product-partner-of-the-year): At Snowflake Summit 2026, dltHub was named Snowflake’s 2026 Startup Program Product Partner of the Year for helping more than 1,000 organizations bring hard-to-reach data into the Snowflake AI Data Cloud with Python-native, AI-driven pipelines. ## Case studies - [dltHub migration services give Navit production-grade data and Chat-BI, without hiring](https://dlthub.com/case-studies/navit): Navit, a ~20-person Berlin-based corporate mobility platform, applied the dltHub AI Workbench ontology toolkit to an existing first-generation pipeline. The team stayed the same size, a generalist now maintains the stack, and Chat-BI on top of the ontology reasons about the business like an analyst would. - [Tasman Analytics prototypes a client's data pipeline in a single meeting with dltHub](https://dlthub.com/case-studies/tasman-analytics): Tasman Analytics, a ~20-person data analytics consultancy, uses dltHub to prototype client connectors in real-time — scoping in minutes instead of weeks — and shift from time-and-materials to fixed-price projects. - [Powering the Energy Transition: How Vandebron Cut Data Workflow Complexity with dlt](https://dlthub.com/case-studies/vandebron): Vandebron, a Dutch green-energy provider, rebuilt its complex ingestion stack in just one week using dlt, cutting costs, code, and runtime dramatically. - [Grocery sensation Erewhon turns cultural buzz into business growth](https://dlthub.com/case-studies/erewhon): A solo data team upskills into more advanced data engineering and finds a robust, reliable solution to data ingestion, building an “enterprise-grade” data operation. - [Remerge's journey from manual processes to streamlined pipelines](https://dlthub.com/case-studies/remerge): Learn how Remerge moved away from manual spreadsheets by centralizing their data, creating a reliable single source of truth. - [Artsy moves data faster](https://dlthub.com/case-studies/artsy): Artsy transforms their 10-year-old legacy system into a streamlined, customizable solution, dramatically reducing data extraction times. - [Flatiron Health accelerates privacy-enhancing data processing](https://dlthub.com/case-studies/flatiron-health): Learn how Flatiron Health cut 50% of their cost of ingestion and transformation pipelines using dlt (data load tool). - [How insurance company Dentolo democratizes data access](https://dlthub.com/case-studies/dentolo): Dentolo transforms its data ingestion process, empowers the team with a composable data stack and democratizes data access across the organization. - [PostHog offers their users a scalable and inexpensive one-click data warehouse](https://dlthub.com/case-studies/posthog): PostHog builds a scalable, customizable data warehouse that seamlessly handles large datasets, and empowers their team to deliver a flexible and high-performing solution for users. - [How Harness transformed 14 data pipelines in 14 days](https://dlthub.com/case-studies/harness): Harness chooses dlt (data load tool) + sqlmesh to create an end-to-end next generation data platform. - [Fintech Taktile builds a compliant data platform](https://dlthub.com/case-studies/taktile): How Taktile uses dlt (data load tool) + Snowflake for custom data needs and empowers all software engineers. ## Ecosystem - [DataHub](https://dlthub.com/partner/datahub): Open-source metadata platform for the modern data stack. - [DuckDB](https://dlthub.com/partner/duckdb): In-process analytical database that's fast, lightweight, and SQL-native. Use as a dlt destination for local development or production analytics — zero infrastructure, parquet-friendly, and schema-evolving by default. - [Green Mountain Data Solutions](https://dlthub.com/partner/green-mountain-data-solutions): Green Mountain Data Solutions is a dltHub Gold consulting partner building end to end analytics on dlt and the dltHub platform. - [Hugging Face](https://dlthub.com/partner/hugging-face): The collaboration platform for the machine learning community. dltHub's native HuggingFace Hub destination lets you push training-ready datasets directly from any dlt pipeline — schema-enforced, deduplicated, and versioned. - [LanceDB](https://dlthub.com/partner/lancedb): Open-source multimodal vector database built for AI. Use dlt with LanceDB as a high-performance storage layer for vectors, images, audio, and structured data — with incremental ingestion, deduplication, and CI/CD-friendly pipelines. - [Modeo](https://dlthub.com/partner/modeo): Modeo is a French data and AI agency building Modern Data Platforms for European clients, partnered with dltHub. - [Parallel](https://dlthub.com/partner/parallel): Web research and browsing APIs for AI agents. - [Probabl](https://dlthub.com/partner/probabl): Sustained stewardship of scikit-learn and the data science stack. - [Snowflake](https://dlthub.com/partner/snowflake-ecosystem): Cloud data warehouse for the AI Data Cloud. dlt loads data into Snowflake with key-pair auth, schema normalization, and warehouse-aware staging — incremental, governed, and ready for analytics workloads. - [Temporal](https://dlthub.com/partner/temporal): Durable execution platform for AI and data workflows. - [Tower](https://dlthub.com/partner/tower): Tower is a data platform for the next generation of Python-based data and AI apps. Run any Python code on Tower, including dlt pipelines and dlt+ projects. - [Untitled Data Company](https://dlthub.com/partner/untitled-data-company): We are a boutique Data Engineering and BI consultancy and help you reduce costs and improve your data stack. We specialize in sustainable and low-cost open-source tools, such as dlt, dbt, Airflow, and Terraform and run on AWS, GCP, and on-prem. - [builders;](https://dlthub.com/partner/builders): Top engineering firm for software and data consulting. We specialize in consulting, software development, and building platforms — in days, not months. ## Agentic & LLM workflows - [dltHub AI harness — toolkits index](https://dlthub.com/ai-harness.md): Auto-generated index of the dltHub AI harness toolkits (skills, commands, MCP servers) for building and operating dlt pipelines with LLMs. - [Agent instructions (AGENTS.md)](https://dlthub.com/AGENTS.md): Starter agent guidance for a dlt workspace — install steps plus the workbench's live rules. Run `dlt ai init` to install the full version-pinned setup locally. - [Agent Skills index](https://dlthub.com/.well-known/agent-skills/index.json): Agent Skills Discovery (RFC v0.2.0) index of every dlt skill, each SKILL.md served under `/.well-known/agent-skills//SKILL.md` with a matching sha256. - [MCP Server Card](https://dlthub.com/.well-known/mcp/server-card.json): SEP-1649 discovery card for `dlt-workspace-mcp` — the stdio MCP server shipped by the `dlt` library, with transport command and install instructions. - [Cheatsheet](https://dlthub.com/cheatsheet): Quick reference for the most common dlt APIs and patterns. ## Source code - [dlt-hub/dlt on GitHub](https://github.com/dlt-hub/dlt): Core library source. - [dlt-hub/verified-sources on GitHub](https://github.com/dlt-hub/verified-sources): Community- and dltHub-verified sources. - [dlt-hub/dlthub-ai-workbench on GitHub](https://github.com/dlt-hub/dlthub-ai-workbench): Claude Code plugin marketplace for dlt (skills, commands, MCP). - [dlt on PyPI](https://pypi.org/project/dlt/): `pip install dlt`. ## Community - [Slack community](https://dlthub.com/community): Join the community Slack. - [Contact](https://dlthub.com/contact): Reach out to the dltHub team. ## Optional - [Sitemap](https://dlthub.com/sitemap.xml): Full page index. - [Blog RSS feed](https://dlthub.com/blog/rss.xml): Subscribe to posts. - [Privacy policy](https://dlthub.com/privacy-policy) - [Imprint](https://dlthub.com/imprint)