Moving from local to production
In previous how-to guides, you used the local stack to create and run your pipeline. This saved you
the headache of setting up a cloud account, credentials, and often also money. Our choice for a local
"warehouse" is duckdb, which is fast, feature-rich, and works everywhere. However, at some point, you might want
to move to production or share the results with your colleagues. The local duckdb file is not
sufficient for that! Let's move a dataset for the chess.com API we have already to
BigQuery:
1. Replace the "destination" argument with "bigquery"
import dlt
if __name__ == "__main__":
pipeline = dlt.pipeline(
pipeline_name="chess_pipeline",
destination='bigquery',
dataset_name="games_data"
)
# get data for a few famous players
data = chess_source(
data=['magnuscarlsen', 'rpragchess'],
start_month="2022/11",
end_month="2022/12"
)
load_info = pipeline.run(data)
And that's it regarding the code modifications! If you run the script, dlt will create an identical
dataset to what you had in duckdb but in BigQuery.
2. Enable access to BigQuery and obtain credentials
Please follow these steps to enable dlt to write data
to BigQuery.
3. Add credentials to secrets.toml
Please add the following section to your secrets.toml file, using the credentials obtained from the
previous step:
[destination.bigquery]
location = "US"
[destination.bigquery.credentials]
project_id = "project_id" # please set me up!
private_key = "private_key" # please set me up!
client_email = "client_email" # please set me up!
4. Run the pipeline again
python chess_pipeline.py
Head on to the next section if you see exceptions!
5. Troubleshoot exceptions
Credentials missing: ConfigFieldMissingException
You'll see this exception if dlt cannot find your BigQuery credentials. In the exception below, all
of them ('project_id', 'private_key', 'client_email') are missing. The exception also gives you the
list of all lookups for configuration performed -
here we explain how to read such a list.
dlt.common.configuration.exceptions.ConfigFieldMissingException: Following fields are missing: ['project_id', 'private_key', 'client_email'] in configuration with spec GcpServiceAccountCredentials
for field "project_id" config providers and keys were tried in the following order:
In Environment Variables key CHESS__DESTINATION__BIGQUERY__CREDENTIALS__PROJECT_ID was not found.
In Environment Variables key CHESS__DESTINATION__CREDENTIALS__PROJECT_ID was not found.
The most common cases for the exception:
- The secrets are not in
secrets.tomlat all. - They are placed in the wrong section. For example, the fragment below will not work:
[destination.bigquery] # 'credentials' missed
project_id = "project_id"
- You run the pipeline script from a different folder from which it is saved. For example,
python chess_demo/chess_pipeline.pywill run the script from thechess_demofolder but the current working directory is the folder above. This preventsdltfrom findingchess_demo/.dlt/secrets.tomland filling in credentials.