Skip to content
Back to skills

Data Ai

ASecurity

Data Engineering skill - ETL, pipelines, warehousing, streaming

  • 3 stars
  • 0 votes
  • 0 copies
  • 0 views
  • Added September 8, 2026
databasespythongosqldatabase

Security analysis

A100/100

Pro scans all 9 files and shows the line behind each finding

Scanned September 8, 2026

npx -y skills add pluginagentmarketplace/custom-plugin-linux --skill data-ai --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Data Ai?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for Data Ai
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/pluginagentmarketplace-data-ai-4922624b/badge)](https://www.skillsdirectory.com/skills/pluginagentmarketplace-data-ai-4922624b)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
SKILL.md
---
name: data
description: Data Engineering skill - ETL, pipelines, warehousing, streaming
version: "1.0.0"
sasmp_version: "1.3.0"

input_schema:
  type: object
  properties:
    task: { type: string, enum: [pipeline, etl, warehouse, streaming, quality] }
    scale: { type: string, enum: [small, medium, large] }
  required: [task]

output_schema:
  type: object
  properties:
    architecture: { type: string }
    code: { type: string }
    tools: { type: array }

retry_config:
  max_attempts: 3
  backoff: exponential

timeout_ms: 60000
---

# Data Engineering Skill

## PURPOSE
Data pipelines, ETL processes, and data platform design.

## CORE COMPETENCIES
```
Data Processing:
├── Batch (Spark, Pandas)
├── Streaming (Kafka, Flink)
├── ETL/ELT patterns
└── Data quality

Storage:
├── Data lakes (S3, GCS)
├── Warehouses (Snowflake, BigQuery)
├── Databases (PostgreSQL, DuckDB)
└── File formats (Parquet, Delta)

Orchestration:
├── Airflow
├── Dagster
├── Prefect
└── dbt
```

## CODE PATTERNS

### Spark ETL
```python
from pyspark.sql import SparkSession
from pyspark.sql.functions import col, when

spark = SparkSession.builder.appName("ETL").getOrCreate()

# Extract
df = spark.read.parquet("s3://data/raw/")

# Transform
transformed = df.filter(col("status") == "active") \
    .withColumn("category", when(col("value") > 100, "high").otherwise("low")) \
    .dropDuplicates(["id"])

# Load
transformed.write.mode("overwrite").partitionBy("date").parquet("s3://data/processed/")
```

### Airflow DAG
```python
from airflow import DAG
from airflow.operators.python import PythonOperator
from datetime import datetime, timedelta

default_args = {
    "retries": 3,
    "retry_delay": timedelta(minutes=5),
}

with DAG("etl_pipeline", default_args=default_args, schedule_interval="@daily") as dag:
    extract = PythonOperator(task_id="extract", python_callable=extract_data)
    transform = PythonOperator(task_id="transform", python_callable=transform_data)
    load = PythonOperator(task_id="load", python_callable=load_data)

    extract >> transform >> load
```

## TROUBLESHOOTING

| Issue | Cause | Solution |
|-------|-------|----------|
| OOM in Spark | Partition skew | Repartition, salting |
| Slow queries | No partition pruning | Add partition filters |
| Data quality | Missing validation | Add checks, dbt tests |

Files in this skill

  • SKILL.md731 B
  • ai-SKILL.md2.6 KB
  • assets/config.yaml690 B
  • assets/schema.json1.1 KB
  • data-SKILL.md2.4 KB
  • ml-SKILL.md2.9 KB
  • references/GUIDE.md1.8 KB
  • references/PATTERNS.md1.5 KB
  • scripts/validate.py3.7 KB

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…