Skip to content
Back to skills

Etl Pipeline

ASecurity

Build an ETL/data pipeline that extracts, transforms, and loads data idempotently

  • 3 stars
  • 0 votes
  • 0 copies
  • 2 views
  • Added September 3, 2026
ai-agentsgoapidatabase

Works with

  • api

Security analysis

A100/100

Scanned September 3, 2026

npx -y skills add black141312/ada --skill etl-pipeline --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Etl Pipeline?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for Etl Pipeline
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/black141312-etl-pipeline/badge)](https://www.skillsdirectory.com/skills/black141312-etl-pipeline)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
SKILL.md
---
name: etl-pipeline
description: Build an ETL/data pipeline that extracts, transforms, and loads data idempotently
category: data-ml
---

# ETL Pipeline

Reach for this when moving data between systems (files, APIs, databases, warehouses) on a schedule or one-shot, and you need it to be re-runnable without duplicating or corrupting data.

1. Pin down source and sink contracts: schema, primary keys, volume, update frequency, and whether the source is append-only or mutable.
2. Split the job into explicit extract, transform, and load stages so each can be tested and re-run in isolation.
3. Make loads idempotent: upsert on a natural/surrogate key, or stage-then-swap, so a re-run never double-writes.
4. Process incrementally using a watermark (updated_at, sequence id, or partition) and persist the high-water mark after a successful load.
5. Validate row counts and key invariants between stages; fail loud and stop before loading bad data downstream.
6. Add structured logging and a dead-letter path for bad records, then wire retries with backoff on transient failures.
7. Make the run parameterized (date range, env) and schedulable via cron/Airflow/Prefect rather than hardcoded.

## Rules
- Never mutate the source; treat raw extracts as immutable and transform into a separate layer.
- Idempotency is non-negotiable — assume every run can be retried or run twice.
- Keep transforms pure and deterministic; isolate I/O so logic is unit-testable without the network.
- Checkpoint the watermark only after the load commits, never before.
- Surface partial failures explicitly; a green exit code must mean the data is actually correct.

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…