Load YouTube API Scraper data to DuckDB
Build a YouTube API Scraper to DuckDB pipeline with your coding agent. One prompt scaffolds it with the dltHub AI harness, plus the YouTube API Scraper API base URL, auth, endpoints, and incremental loading.
The YouTube Data API v3 allows developers to access and modify YouTube content, including videos, playlists, and channel information. Everything needed to build a working YouTube API Scraper → DuckDB pipeline is on this page: the API's base URL, authentication, endpoints, pagination and incremental field — plus a prompt that hands the whole job to your coding agent.
Build your YouTube API Scraper to DuckDB pipeline
Paste this prompt into Claude, Codex, or Cursor. The agent does the rest.
PromptRunuvx dlthub-init@latestto build a pipeline from YouTube API Scraper to DuckDB and run it on dltHub
That scaffolds a dltHub workspace and installs the dltHub AI harness — the project rules, the secrets-management skill, and the dlt MCP server your agent needs to work safely. From there it reads the YouTube API Scraper API, proposes the endpoints to load, then writes, runs and validates the pipeline while you review rather than type. Credentials are inspected through MCP tools, so your agent never reads secrets.toml itself. How the LLM-native workflow works →
Prefer to write it yourself? Every fact the agent uses is below.
YouTube API Scraper API at a glance
| Base URL | https://www.googleapis.com/youtube/v3 |
| Example endpoint | GET videos |
| Records found at | items |
| Authentication | requests require either an API key or an OAuth 2.0 token — sent in the Authorization header, prefixed Bearer |
| Pagination | Cursor-based via pageToken, next cursor at nextPageToken, page size via maxResults (default 5, max 50) |
| Record id | id |
| API reference | https://developers.google.com/youtube/v3/docs |
These values come from the YouTube API Scraper API reference — the authoritative source if anything here looks out of date.
How do I authenticate with the YouTube API Scraper API?
The API supports either an API key provided via a query parameter or an OAuth 2.0 access token provided via an Authorization header or query parameter. OAuth 2.0 tokens are mandatory for modifying data or accessing private user information.
1. Get your credentials
- Navigate to the Google Cloud Console (console.cloud.google.com).
- Create or select a project.
- In the search bar, search for "YouTube Data API v3" and enable it for your project.
- Navigate to the APIs & Services > Credentials page.
- Click "+ CREATE CREDENTIALS" and select "API key".
- Copy the generated API key for use in your pipeline. For production, it is recommended to click "Edit API key" and apply application or API restrictions.
2. Add them to .dlt/secrets.toml
[sources.youtube_api_scraper_source] api_key = "your_actual_api_key_here"
dlt reads this file automatically at runtime. With the harness, the setup-secrets skill prompts you for the values and never handles the raw credential in chat. For production, see setting up credentials with dlt.
What YouTube API Scraper data can I load into DuckDB?
These are the YouTube API Scraper endpoints dlt can load into DuckDB:
| Resource | Endpoint | Method | Data selector | Description |
|---|---|---|---|---|
| videos | GET /videos | GET | items | Returns a list of videos matching request parameters. |
| playlists | GET /playlists | GET | items | Returns a list of playlists matching request parameters. |
| channels | GET /channels | GET | items | Returns a list of channels matching request parameters. |
| playlist_items | GET /playlistItems | GET | items | Returns a collection of playlist items matching request parameters. |
| search | GET /search | GET | items | Returns a list of search results matching request parameters. |
How do I load only new YouTube API Scraper records?
The YouTube API Scraper API reference does not document a timestamp or sequence field for these endpoints, so there is nothing to advertise here as verified. Pick a field from the endpoints table above that increases with every write, then set it as the cursor_path.
{"name": "videos", "endpoint": { "path": "videos", # Replace with a field that increases on every write. "incremental": {"cursor_path": "REPLACE_ME", "initial_value": "2024-01-01T00:00:00Z"}, }}
On the first run dlt loads everything from initial_value; on every run after that it requests only what changed and appends with write_disposition="merge" if you set a primary key. See incremental loading.
What does the generated YouTube API Scraper pipeline look like?
A standard dlt REST API pipeline — the same code you would write by hand, loading videos and search from the YouTube API Scraper API into DuckDB:
import dlt from dlt.sources.rest_api import RESTAPIConfig, rest_api_resources @dlt.source def youtube_api_scraper_source(api_key=dlt.secrets.value): config: RESTAPIConfig = { "client": { "base_url": "https://www.googleapis.com/youtube/v3", "auth": {"type": "bearer", "token": api_key}, }, "resources": [ {"name": "videos", "endpoint": {"path": "videos", "data_selector": "items"}}, {"name": "playlists", "endpoint": {"path": "playlists", "data_selector": "items"}} ], } yield from rest_api_resources(config) def load_youtube_api_scraper_to_duckdb() -> None: pipeline = dlt.pipeline( pipeline_name="youtube_api_scraper_pipeline", destination="duckdb", dataset_name="youtube_api_scraper_data", ) load_info = pipeline.run(youtube_api_scraper_source()) print(load_info) if __name__ == "__main__": load_youtube_api_scraper_to_duckdb()
Run it with python youtube_api_scraper_pipeline.py. The agent iterates on this until it loads cleanly — you review and approve, rather than write it from scratch.
How do I query YouTube API Scraper data in DuckDB?
dlt creates one table per resource. Query the loaded data with Python or SQL — or ask your agent to, through the MCP server's execute_sql_query tool.
Python (pandas DataFrame):
import dlt data = dlt.pipeline("youtube_api_scraper_pipeline").dataset() df = data.videos.df() print(df.head())
SQL:
SELECT * FROM youtube_api_scraper_data.videos LIMIT 10;
See querying your data with dataset and exploring it in marimo notebooks.
How do I deploy the YouTube API Scraper to DuckDB pipeline in production?
The pipeline runs locally, which is ideal for prototyping and one-off analysis. When you need it on a schedule, monitored on every load, and shared with your team, deploy the same dlt code on the dltHub platform — no infrastructure to maintain. The prompt above already ends with "run it on dltHub", so your agent can take it there directly.
- Deploy & schedule — run the pipeline as a managed job with automatic retries.
- Monitor — observable job queues, alerting, and load metrics for every run.
- Transform — promote raw YouTube API Scraper loads into governed, documented models.
- Visualize & share — explore data in notebooks and publish live dashboards instead of static screenshots.
What other destinations can I load YouTube API Scraper data to?
dlt loads into any of these — only the destination argument changes:
| Destination | Example value |
|---|---|
| PostgreSQL | "postgres" |
| BigQuery | "bigquery" |
| Snowflake | "snowflake" |
| Redshift | "redshift" |
| Databricks | "databricks" |
| Filesystem (S3, GCS, Azure) | "filesystem" |
Set dlt.pipeline(destination="snowflake") and add credentials in .dlt/secrets.toml. On the dltHub platform the same pipeline runs against a managed Iceberg lakehouse. See the full destinations list.
Next steps
Was this page helpful?
Community Hub
Need more dlt context for YouTube API Scraper to DuckDB?
Request dlt skills, commands, AGENT.md files, and AI-native context.