Load Clerk data to DuckDB
Build a Clerk to DuckDB pipeline with your coding agent. One prompt scaffolds it with the dltHub AI harness, plus the Clerk API base URL, auth, endpoints, and incremental loading.
Clerk Backend API is a REST service that allows backend servers to manage Clerk resources and perform data operations outside of client-side sessions. Everything needed to build a working Clerk → DuckDB pipeline is on this page: the API's base URL, authentication, endpoints, pagination and incremental field — plus a prompt that hands the whole job to your coding agent.
Build your Clerk to DuckDB pipeline
Paste this prompt into Claude, Codex, or Cursor. The agent does the rest.
PromptRunuvx dlthub-init@latestto build a pipeline from Clerk to DuckDB and run it on dltHub
That scaffolds a dltHub workspace and installs the dltHub AI harness — the project rules, the secrets-management skill, and the dlt MCP server your agent needs to work safely. From there it reads the Clerk API, proposes the endpoints to load, then writes, runs and validates the pipeline while you review rather than type. Credentials are inspected through MCP tools, so your agent never reads secrets.toml itself. How the LLM-native workflow works →
Prefer to write it yourself? Every fact the agent uses is below.
Clerk API at a glance
| Base URL | https://api.clerk.com/v1 |
| Example endpoint | GET users |
| Records found at | data |
| Authentication | all requests require a Bearer token — sent in the Authorization header, prefixed Bearer |
| Pagination | Offset-based page size via limit (default 10, max 500). The Clerk Backend API uses offset-based pagination. Common parameters are 'limit' and 'offset'. Some specific endpoints, such as application transfers, have been noted to also support cursor-based pagination parameters 'starting_after' and 'ending_before'. |
| Incremental field | updated_at |
| API reference | https://clerk.com/docs/reference/api/overview |
These values come from the Clerk API reference — the authoritative source if anything here looks out of date.
How do I authenticate with the Clerk API?
All requests require an 'Authorization' header containing a Bearer token. The token can be a session token, an API key, or other Clerk-issued machine-to-machine tokens.
1. Get your credentials
To obtain your Clerk API credentials, log in to the Clerk Dashboard and navigate to the API keys page (typically found under the Configure > Developer section). From there, you can view and copy your Publishable Key and Secret Key, which are used to authenticate your application with Clerk. If you need machine-to-machine API keys for user or organization actions, you can enable them on this same page.
2. Add them to .dlt/secrets.toml
[sources.clerk_source] clerk_secret_key = "sk_live_..." clerk_publishable_key = "pk_live_..."
dlt reads this file automatically at runtime. With the harness, the setup-secrets skill prompts you for the values and never handles the raw credential in chat. For production, see setting up credentials with dlt.
What Clerk data can I load into DuckDB?
These are the Clerk endpoints dlt can load into DuckDB:
| Resource | Endpoint | Method | Data selector | Description |
|---|---|---|---|---|
| users | /users | GET | Retrieves a list of users. | |
| organizations | /organizations | GET | Retrieves a list of organizations. | |
| sessions | /sessions | GET | Retrieves a list of sessions. | |
| api_keys | /api_keys | GET | Retrieves a list of API keys. | |
| clients | /clients | GET | Retrieves a list of clients. |
How do I load only new Clerk records?
Clerk exposes updated_at on users, so dlt can request only the records that changed since the last run. Set it as the cursor_path and dlt tracks the high-water mark for you between runs.
{"name": "users", "endpoint": { "path": "users", "data_selector": "data", "incremental": {"cursor_path": "updated_at", "initial_value": "2024-01-01T00:00:00Z"}, }}
On the first run dlt loads everything from initial_value; on every run after that it requests only what changed and appends with write_disposition="merge" if you set a primary key. See incremental loading.
What does the generated Clerk pipeline look like?
A standard dlt REST API pipeline — the same code you would write by hand, loading GET /api_keys and POST /api_keys from the Clerk API into DuckDB:
import dlt from dlt.sources.rest_api import RESTAPIConfig, rest_api_resources @dlt.source def clerk_source(api_key=dlt.secrets.value): config: RESTAPIConfig = { "client": { "base_url": "https://api.clerk.com/v1", "auth": {"type": "bearer", "token": api_key}, }, "resources": [ {"name": "users", "endpoint": {"path": "users", "data_selector": "data"}}, {"name": "organizations", "endpoint": {"path": "organizations", "data_selector": "data"}} ], } yield from rest_api_resources(config) def load_clerk_to_duckdb() -> None: pipeline = dlt.pipeline( pipeline_name="clerk_pipeline", destination="duckdb", dataset_name="clerk_data", ) load_info = pipeline.run(clerk_source()) print(load_info) if __name__ == "__main__": load_clerk_to_duckdb()
Run it with python clerk_pipeline.py. The agent iterates on this until it loads cleanly — you review and approve, rather than write it from scratch.
How do I query Clerk data in DuckDB?
dlt creates one table per resource. Query the loaded data with Python or SQL — or ask your agent to, through the MCP server's execute_sql_query tool.
Python (pandas DataFrame):
import dlt data = dlt.pipeline("clerk_pipeline").dataset() df = data.users.df() print(df.head())
SQL:
SELECT * FROM clerk_data.users LIMIT 10;
See querying your data with dataset and exploring it in marimo notebooks.
How do I deploy the Clerk to DuckDB pipeline in production?
The pipeline runs locally, which is ideal for prototyping and one-off analysis. When you need it on a schedule, monitored on every load, and shared with your team, deploy the same dlt code on the dltHub platform — no infrastructure to maintain. The prompt above already ends with "run it on dltHub", so your agent can take it there directly.
- Deploy & schedule — run the pipeline as a managed job with automatic retries.
- Monitor — observable job queues, alerting, and load metrics for every run.
- Transform — promote raw Clerk loads into governed, documented models.
- Visualize & share — explore data in notebooks and publish live dashboards instead of static screenshots.
What other destinations can I load Clerk data to?
dlt loads into any of these — only the destination argument changes:
| Destination | Example value |
|---|---|
| PostgreSQL | "postgres" |
| BigQuery | "bigquery" |
| Snowflake | "snowflake" |
| Redshift | "redshift" |
| Databricks | "databricks" |
| Filesystem (S3, GCS, Azure) | "filesystem" |
Set dlt.pipeline(destination="snowflake") and add credentials in .dlt/secrets.toml. On the dltHub platform the same pipeline runs against a managed Iceberg lakehouse. See the full destinations list.
Next steps
Was this page helpful?
Community Hub
Need more dlt context for Clerk to DuckDB?
Request dlt skills, commands, AGENT.md files, and AI-native context.