No logo available for AWS SDK for pandas to DuckDB connector icon

Load AWS SDK for pandas data to DuckDB

Build a AWS SDK for pandas to DuckDB pipeline with your coding agent. One prompt scaffolds it with the dltHub AI harness, plus the AWS SDK for pandas API base URL, auth, endpoints, and incremental loading.

SourceAWS SDK for pandasAWS SDK for pandas API DocumentationDestinationDuckDBIn-process analytical database. The default local destination for dlt pipelines.

AWS SDK for pandas (awswrangler) is a Python library that provides high-level helpers to read and write data between pandas DataFrames and various AWS services such as Amazon S3, AWS Glue, Amazon Athena, and Amazon Redshift. Everything needed to build a working AWS SDK for pandas → DuckDB pipeline is on this page: the API's base URL, authentication, endpoints, pagination and incremental field — plus a prompt that hands the whole job to your coding agent.


Build your AWS SDK for pandas to DuckDB pipeline

Paste this prompt into Claude, Codex, or Cursor. The agent does the rest.

Prompt
Run uvx dlthub-init@latest to build a pipeline from AWS SDK for pandas to DuckDB and run it on dltHub

That scaffolds a dltHub workspace and installs the dltHub AI harness — the project rules, the secrets-management skill, and the dlt MCP server your agent needs to work safely. From there it reads the AWS SDK for pandas API, proposes the endpoints to load, then writes, runs and validates the pipeline while you review rather than type. Credentials are inspected through MCP tools, so your agent never reads secrets.toml itself. How the LLM-native workflow works →

Prefer to write it yourself? Every fact the agent uses is below.


AWS SDK for pandas API at a glance

Base URLThe AWS SDK for pandas does not have a single global REST API base URL; it acts as a client wrapper for various AWS service APIs that use their respective regional service endpoints.
Example endpointPOST _search
Records found athits.hits
AuthenticationAll requests are authenticated using AWS Signature Version 4 (SigV4) via the boto3 library
PaginationCursor-based via pagination_config['StartingToken'], next cursor at page['NextToken'], page size via pagination_config['PageSize'] (default 10, max 10). In awswrangler, pagination is configured via the pagination_config dict passed to wr.timestream.query(sql, pagination_config=...). The dict keys shown are 'MaxItems', 'PageSize', and 'StartingToken'. Next-page continuation is returned by the underlying service as 'NextToken' in each page, but awswrangler may only expose it via DataFrame metadata when chunked=True; NextToken is not documented as a top-level return value in the stub signature.
API referencehttps://aws-sdk-pandas.readthedocs.io/

These values come from the AWS SDK for pandas API reference — the authoritative source if anything here looks out of date.


How do I authenticate with the AWS SDK for pandas API?

All requests utilize AWS Signature Version 4 (SigV4) handled automatically via the underlying boto3 session. Users provide credentials through the standard AWS credential resolution process (environment variables, IAM roles, or local credential files).

1. Get your credentials

AWS SDK for pandas (awswrangler) does not have its own dedicated REST API with an API key dashboard. Instead, it acts as a wrapper around Boto3. To authenticate, you must provide standard AWS credentials. These can be obtained via the AWS Management Console by creating an IAM user or role with the necessary permissions, then generating an Access Key ID and Secret Access Key. Alternatively, use IAM roles for EC2, ECS, or Lambda, or configure local environment profiles via the AWS CLI using aws configure. It relies on Boto3's credential resolution chain.

2. Add them to .dlt/secrets.toml

[sources.aws_sdk_for_pandas_source] aws_access_key_id = "YOUR_ACCESS_KEY" aws_secret_access_key = "YOUR_SECRET_KEY" aws_region = "us-east-1"

dlt reads this file automatically at runtime. With the harness, the setup-secrets skill prompts you for the values and never handles the raw credential in chat. For production, see setting up credentials with dlt.


What AWS SDK for pandas data can I load into DuckDB?

These are the AWS SDK for pandas endpoints dlt can load into DuckDB:

ResourceEndpointMethodData selectorDescription
opensearch_search_searchPOSThits.hitsExecutes a search query against an OpenSearch domain.
opensearch_bulk_bulkPOSTitemsPerforms bulk operations in an OpenSearch domain.
opensearch_index_docPOSTIndexes a document in an OpenSearch domain.
opensearch_get_doc/{id}GETRetrieves a document by ID from an OpenSearch domain.
opensearch_count_countGETcountReturns the number of documents matching a query.

How do I load only new AWS SDK for pandas records?

The AWS SDK for pandas API reference does not document a timestamp or sequence field for these endpoints, so there is nothing to advertise here as verified. Pick a field from the endpoints table above that increases with every write, then set it as the cursor_path.

{"name": "opensearch_search", "endpoint": { "path": "_search", # Replace with a field that increases on every write. "incremental": {"cursor_path": "REPLACE_ME", "initial_value": "2024-01-01T00:00:00Z"}, }}

On the first run dlt loads everything from initial_value; on every run after that it requests only what changed and appends with write_disposition="merge" if you set a primary key. See incremental loading.


What does the generated AWS SDK for pandas pipeline look like?

A standard dlt REST API pipeline — the same code you would write by hand, loading The library does not have a single REST API with fixed endpoints; it uses service-specific endpoints determined by Boto3. For operations involving specific services like Amazon OpenSearch or Amazon S3, you interact with the service's endpoint, e.g., s3 and opensearch. from the AWS SDK for pandas API into DuckDB:

import dlt from dlt.sources.rest_api import RESTAPIConfig, rest_api_resources @dlt.source def aws_sdk_for_pandas_source(boto3_session=dlt.secrets.value): config: RESTAPIConfig = { "client": { "base_url": "The AWS SDK for pandas does not have a single global REST API base URL; it acts as a client wrapper for various AWS service APIs that use their respective regional service endpoints.", "auth": {"type": "api_key", "api_key": boto3_session, "name": "boto3_session"}, }, "resources": [ {"name": "opensearch_search", "endpoint": {"path": "_search", "data_selector": "hits.hits"}}, {"name": "opensearch_bulk", "endpoint": {"path": "_bulk", "data_selector": "items"}} ], } yield from rest_api_resources(config) def load_aws_sdk_for_pandas_to_duckdb() -> None: pipeline = dlt.pipeline( pipeline_name="aws_sdk_for_pandas_pipeline", destination="duckdb", dataset_name="aws_sdk_for_pandas_data", ) load_info = pipeline.run(aws_sdk_for_pandas_source()) print(load_info) if __name__ == "__main__": load_aws_sdk_for_pandas_to_duckdb()

Run it with python aws_sdk_for_pandas_pipeline.py. The agent iterates on this until it loads cleanly — you review and approve, rather than write it from scratch.


How do I query AWS SDK for pandas data in DuckDB?

dlt creates one table per resource. Query the loaded data with Python or SQL — or ask your agent to, through the MCP server's execute_sql_query tool.

Python (pandas DataFrame):

import dlt data = dlt.pipeline("aws_sdk_for_pandas_pipeline").dataset() df = data.opensearch_search.df() print(df.head())

SQL:

SELECT * FROM aws_sdk_for_pandas_data.opensearch_search LIMIT 10;

See querying your data with dataset and exploring it in marimo notebooks.


How do I deploy the AWS SDK for pandas to DuckDB pipeline in production?

The pipeline runs locally, which is ideal for prototyping and one-off analysis. When you need it on a schedule, monitored on every load, and shared with your team, deploy the same dlt code on the dltHub platform — no infrastructure to maintain. The prompt above already ends with "run it on dltHub", so your agent can take it there directly.

  • Deploy & schedule — run the pipeline as a managed job with automatic retries.
  • Monitor — observable job queues, alerting, and load metrics for every run.
  • Transform — promote raw AWS SDK for pandas loads into governed, documented models.
  • Visualize & share — explore data in notebooks and publish live dashboards instead of static screenshots.

Book a demo →


What other destinations can I load AWS SDK for pandas data to?

dlt loads into any of these — only the destination argument changes:

DestinationExample value
PostgreSQL"postgres"
BigQuery"bigquery"
Snowflake"snowflake"
Redshift"redshift"
Databricks"databricks"
Filesystem (S3, GCS, Azure)"filesystem"

Set dlt.pipeline(destination="snowflake") and add credentials in .dlt/secrets.toml. On the dltHub platform the same pipeline runs against a managed Iceberg lakehouse. See the full destinations list.


Next steps

Was this page helpful?

Community Hub

Need more dlt context for AWS SDK for pandas to DuckDB?

Request dlt skills, commands, AGENT.md files, and AI-native context.