Load Kepler.gl data to DuckDB
Build a Kepler.gl to DuckDB pipeline with your coding agent. One prompt scaffolds it with the dltHub AI harness, plus the Kepler.gl API base URL, auth, endpoints, and incremental loading.
Kepler.gl is a data-agnostic, high-performance, open-source React component for visual exploration of large-scale geolocation data sets. Everything needed to build a working Kepler.gl → DuckDB pipeline is on this page: the API's base URL, authentication, endpoints, pagination and incremental field — plus a prompt that hands the whole job to your coding agent.
Build your Kepler.gl to DuckDB pipeline
Paste this prompt into Claude, Codex, or Cursor. The agent does the rest.
PromptRunuvx dlthub-init@latestto build a pipeline from Kepler.gl to DuckDB and run it on dltHub
That scaffolds a dltHub workspace and installs the dltHub AI harness — the project rules, the secrets-management skill, and the dlt MCP server your agent needs to work safely. From there it reads the Kepler.gl API, proposes the endpoints to load, then writes, runs and validates the pipeline while you review rather than type. Credentials are inspected through MCP tools, so your agent never reads secrets.toml itself. How the LLM-native workflow works →
Prefer to write it yourself? Every fact the agent uses is below.
Kepler.gl API at a glance
| Base URL | Not applicable as Kepler.gl is a client-side component library, not a REST service. |
| Example endpoint | GET listMaps |
| Authentication | Requires a Mapbox access token provided as a component prop |
| Pagination | Not paginated |
| API reference | https://docs.kepler.gl/docs/api-reference/get-started.md |
These values come from the Kepler.gl API reference — the authoritative source if anything here looks out of date.
How do I authenticate with the Kepler.gl API?
Kepler.gl is a React component library and does not provide a REST API for data pipelines; authentication is handled via a Mapbox access token passed as a component property to the Kepler.gl UI component.
1. Get your credentials
Kepler.gl is a frontend React component and does not have a REST API that requires authentication. However, it relies on Mapbox for basemap tiles, which requires a Mapbox Access Token. To obtain this: 1. Sign up or log in at mapbox.com. 2. Navigate to the Account Dashboard. 3. Locate the 'Access tokens' section. 4. Click 'Create a token' to generate a new key. 5. Copy the generated token for use in your application.
2. Add them to .dlt/secrets.toml
[sources.kepler_gl_source] mapbox_token = "your_mapbox_access_token_here"
dlt reads this file automatically at runtime. With the harness, the setup-secrets skill prompts you for the values and never handles the raw credential in chat. For production, see setting up credentials with dlt.
What Kepler.gl data can I load into DuckDB?
These are the Kepler.gl endpoints dlt can load into DuckDB:
| Resource | Endpoint | Method | Data selector | Description |
|---|---|---|---|---|
| cloud_maps | listMaps | N/A | N/A | Method called to load a catalog of maps saved by the user. Note: Kepler.gl is a frontend library; it does not provide a standard REST API. Cloud integration is implemented via custom Provider objects. |
| cloud_maps | downloadMap | N/A | N/A | Method called to download a specific map from storage. |
| cloud_maps | uploadMap | N/A | N/A | Method called to upload map to storage. |
| cloud_maps | login | N/A | N/A | Method called to perform user login. |
| cloud_maps | logout | N/A | N/A | Method called to logout a user. |
How do I load only new Kepler.gl records?
The Kepler.gl API reference does not document a timestamp or sequence field for these endpoints, so there is nothing to advertise here as verified. Pick a field from the endpoints table above that increases with every write, then set it as the cursor_path.
{"name": "cloud_maps", "endpoint": { "path": "listMaps", # Replace with a field that increases on every write. "incremental": {"cursor_path": "REPLACE_ME", "initial_value": "2024-01-01T00:00:00Z"}, }}
On the first run dlt loads everything from initial_value; on every run after that it requests only what changed and appends with write_disposition="merge" if you set a primary key. See incremental loading.
What does the generated Kepler.gl pipeline look like?
A standard dlt REST API pipeline — the same code you would write by hand, loading Kepler.gl does not expose REST API endpoints. Interaction is handled via library actions, most commonly addDataToMap and KeplerGlSchema.getConfigToSave. from the Kepler.gl API into DuckDB:
import dlt from dlt.sources.rest_api import RESTAPIConfig, rest_api_resources @dlt.source def kepler_gl_source(mapboxapiaccesstoken=dlt.secrets.value): config: RESTAPIConfig = { "client": { "base_url": "Not applicable as Kepler.gl is a client-side component library, not a REST service.", "auth": {"type": "api_key", "api_key": mapboxapiaccesstoken, "name": "mapboxApiAccessToken"}, }, "resources": [ {"name": "cloud_maps", "endpoint": {"path": "listMaps"}}, {"name": "cloud_maps", "endpoint": {"path": "downloadMap"}} ], } yield from rest_api_resources(config) def load_kepler_gl_to_duckdb() -> None: pipeline = dlt.pipeline( pipeline_name="kepler_gl_pipeline", destination="duckdb", dataset_name="kepler_gl_data", ) load_info = pipeline.run(kepler_gl_source()) print(load_info) if __name__ == "__main__": load_kepler_gl_to_duckdb()
Run it with python kepler_gl_pipeline.py. The agent iterates on this until it loads cleanly — you review and approve, rather than write it from scratch.
How do I query Kepler.gl data in DuckDB?
dlt creates one table per resource. Query the loaded data with Python or SQL — or ask your agent to, through the MCP server's execute_sql_query tool.
Python (pandas DataFrame):
import dlt data = dlt.pipeline("kepler_gl_pipeline").dataset() df = data.listMaps.df() print(df.head())
SQL:
SELECT * FROM kepler_gl_data.listMaps LIMIT 10;
See querying your data with dataset and exploring it in marimo notebooks.
How do I deploy the Kepler.gl to DuckDB pipeline in production?
The pipeline runs locally, which is ideal for prototyping and one-off analysis. When you need it on a schedule, monitored on every load, and shared with your team, deploy the same dlt code on the dltHub platform — no infrastructure to maintain. The prompt above already ends with "run it on dltHub", so your agent can take it there directly.
- Deploy & schedule — run the pipeline as a managed job with automatic retries.
- Monitor — observable job queues, alerting, and load metrics for every run.
- Transform — promote raw Kepler.gl loads into governed, documented models.
- Visualize & share — explore data in notebooks and publish live dashboards instead of static screenshots.
What other destinations can I load Kepler.gl data to?
dlt loads into any of these — only the destination argument changes:
| Destination | Example value |
|---|---|
| PostgreSQL | "postgres" |
| BigQuery | "bigquery" |
| Snowflake | "snowflake" |
| Redshift | "redshift" |
| Databricks | "databricks" |
| Filesystem (S3, GCS, Azure) | "filesystem" |
Set dlt.pipeline(destination="snowflake") and add credentials in .dlt/secrets.toml. On the dltHub platform the same pipeline runs against a managed Iceberg lakehouse. See the full destinations list.
Next steps
Was this page helpful?
Community Hub
Need more dlt context for Kepler.gl to DuckDB?
Request dlt skills, commands, AGENT.md files, and AI-native context.