No logo available for Steamworks to DuckDB connector icon

Load Steamworks data to DuckDB

Build a Steamworks to DuckDB pipeline with your coding agent. One prompt scaffolds it with the dltHub AI harness, plus the Steamworks API base URL, auth, endpoints, and incremental loading.

SourceSteamworksSteamworks API DocumentationDestinationDuckDBIn-process analytical database. The default local destination for dlt pipelines.

Steamworks Web API provides access to various Steamworks features, including public data and protected publisher-only services. Everything needed to build a working Steamworks → DuckDB pipeline is on this page: the API's base URL, authentication, endpoints, pagination and incremental field — plus a prompt that hands the whole job to your coding agent.


Build your Steamworks to DuckDB pipeline

Paste this prompt into Claude, Codex, or Cursor. The agent does the rest.

Prompt
Run uvx dlthub-init@latest to build a pipeline from Steamworks to DuckDB and run it on dltHub

That scaffolds a dltHub workspace and installs the dltHub AI harness — the project rules, the secrets-management skill, and the dlt MCP server your agent needs to work safely. From there it reads the Steamworks API, proposes the endpoints to load, then writes, runs and validates the pipeline while you review rather than type. Credentials are inspected through MCP tools, so your agent never reads secrets.toml itself. How the LLM-native workflow works →

Prefer to write it yourself? Every fact the agent uses is below.


Steamworks API at a glance

Base URLhttps://api.steampowered.com/ (public) or https://partner.steam-api.com/ (partner-only)
Example endpointGET IPublishedFileService/QueryFiles/v1
Records found atresponse
Authenticationuses Web API keys provided via parameter or header — sent in the x-webapi-key header
PaginationCursor-based via cursor, next cursor at next_cursor, page size via numperpage, num_per_page, max_results
Incremental fieldcursor
API referencehttps://partner.steamgames.com/doc/webapi_overview/auth

These values come from the Steamworks API reference — the authoritative source if anything here looks out of date.


How do I authenticate with the Steamworks API?

API keys can be provided either as a 'key' query parameter or by setting the 'x-webapi-key' request header. All Web API requests that contain Web API keys should be made over HTTPS.

1. Get your credentials

To obtain a publisher Web API key, you must have administrator permissions in your Steamworks account. Log in to the Steamworks dashboard and navigate to Users & Permissions, then select Manage Groups. From there, either select an existing group or create a new one, ensure the correct applications are associated with it, and select Create WebAPI Key. Follow the prompts to set desired permissions and save your changes; the key will then appear in the right-hand sidebar. For non-publisher needs, standard user keys can be generated via the Steam Community registration page.

2. Add them to .dlt/secrets.toml

[sources.steamworks_source] steam_api_key = "your_api_key_here"

dlt reads this file automatically at runtime. With the harness, the setup-secrets skill prompts you for the values and never handles the raw credential in chat. For production, see setting up credentials with dlt.


What Steamworks data can I load into DuckDB?

These are the Steamworks endpoints dlt can load into DuckDB:

ResourceEndpointMethodData selectorDescription
published_files/IPublishedFileService/QueryFiles/v1/GETresponsePerforms a search query for published files. Uses cursor or page for pagination.
app_list/IStoreService/GetAppList/v1/GETresponseReturns a list of all apps on the Steam Store. Uses last_appid for pagination.
player_bans/ISteamUser/GetPlayerBans/v1/GETplayersReturns ban status for given Steam IDs.
friend_list/ISteamUser/GetFriendList/v1/GETfriendslistReturns the friend list of a given Steam user.
owned_games/IPlayerService/GetOwnedGames/v1/GETresponseReturns a list of games owned by the user.

How do I load only new Steamworks records?

Steamworks exposes cursor on IPublishedFileService/QueryFiles/v1, so dlt can request only the records that changed since the last run. Set it as the cursor_path and dlt tracks the high-water mark for you between runs.

{"name": "published_files", "endpoint": { "path": "IPublishedFileService/QueryFiles/v1", "data_selector": "response", "incremental": {"cursor_path": "cursor", "initial_value": "2024-01-01T00:00:00Z"}, }}

On the first run dlt loads everything from initial_value; on every run after that it requests only what changed and appends with write_disposition="merge" if you set a primary key. See incremental loading.


What does the generated Steamworks pipeline look like?

A standard dlt REST API pipeline — the same code you would write by hand, loading api.steampowered.com (public API) and partner.steam-api.com (partner-only secure server) from the Steamworks API into DuckDB:

import dlt from dlt.sources.rest_api import RESTAPIConfig, rest_api_resources @dlt.source def steamworks_source(key=dlt.secrets.value): config: RESTAPIConfig = { "client": { "base_url": "https://api.steampowered.com/ (public) or https://partner.steam-api.com/ (partner-only)", "auth": {"type": "api_key", "api_key": key, "name": "x-webapi-key", "location": "header"}, }, "resources": [ {"name": "published_files", "endpoint": {"path": "IPublishedFileService/QueryFiles/v1", "data_selector": "response"}}, {"name": "app_list", "endpoint": {"path": "IStoreService/GetAppList/v1", "data_selector": "response"}} ], } yield from rest_api_resources(config) def load_steamworks_to_duckdb() -> None: pipeline = dlt.pipeline( pipeline_name="steamworks_pipeline", destination="duckdb", dataset_name="steamworks_data", ) load_info = pipeline.run(steamworks_source()) print(load_info) if __name__ == "__main__": load_steamworks_to_duckdb()

Run it with python steamworks_pipeline.py. The agent iterates on this until it loads cleanly — you review and approve, rather than write it from scratch.


How do I query Steamworks data in DuckDB?

dlt creates one table per resource. Query the loaded data with Python or SQL — or ask your agent to, through the MCP server's execute_sql_query tool.

Python (pandas DataFrame):

import dlt data = dlt.pipeline("steamworks_pipeline").dataset() df = data.published_files.df() print(df.head())

SQL:

SELECT * FROM steamworks_data.published_files LIMIT 10;

See querying your data with dataset and exploring it in marimo notebooks.


How do I deploy the Steamworks to DuckDB pipeline in production?

The pipeline runs locally, which is ideal for prototyping and one-off analysis. When you need it on a schedule, monitored on every load, and shared with your team, deploy the same dlt code on the dltHub platform — no infrastructure to maintain. The prompt above already ends with "run it on dltHub", so your agent can take it there directly.

  • Deploy & schedule — run the pipeline as a managed job with automatic retries.
  • Monitor — observable job queues, alerting, and load metrics for every run.
  • Transform — promote raw Steamworks loads into governed, documented models.
  • Visualize & share — explore data in notebooks and publish live dashboards instead of static screenshots.

Book a demo →


What other destinations can I load Steamworks data to?

dlt loads into any of these — only the destination argument changes:

DestinationExample value
PostgreSQL"postgres"
BigQuery"bigquery"
Snowflake"snowflake"
Redshift"redshift"
Databricks"databricks"
Filesystem (S3, GCS, Azure)"filesystem"

Set dlt.pipeline(destination="snowflake") and add credentials in .dlt/secrets.toml. On the dltHub platform the same pipeline runs against a managed Iceberg lakehouse. See the full destinations list.


Next steps

Was this page helpful?

Community Hub

Need more dlt context for Steamworks to DuckDB?

Request dlt skills, commands, AGENT.md files, and AI-native context.