Skip to content
Back to skills

Data Engineering

ASecurity

Method for scraping, transforming, validating and loading data through NDJSON. Use when building a scraper or ETL step, validating an NDJSON feed, or importing data into a CMS or database.

  • 80 stars
  • 0 votes
  • 0 copies
  • 0 views
  • Added May 28, 2026
ai-agentsapidatabase

Works with

  • api

Security analysis

A100/100

Pro scans all 2 files and shows the line behind each finding

Scanned October 5, 2026

npx -y skills add monkilabs/opencastle --skill data-engineering --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Data Engineering?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for Data Engineering
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/monkilabs-data-engineering/badge)](https://www.skillsdirectory.com/skills/monkilabs-data-engineering)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
SKILL.md
---
name: data-engineering
description: "Method for scraping, transforming, validating and loading data through NDJSON. Use when building a scraper or ETL step, validating an NDJSON feed, or importing data into a CMS or database."
---

# Data Engineering

The project's sources, record schema, scripts and commands live in `.opencastle/stack/data-pipeline-config.md`. Use those names; when a script it lists does not exist, write it rather than guessing at one. Starter scraper and validator: [REFERENCE.md](./REFERENCE.md).

## Scraper

Headless browser (Playwright, or Puppeteer Cluster for volume) with retries and a per-page timeout, so one slow page does not stall the run.

## NDJSON Output

One record per line, valid JSON, fields as the project's schema names them. Every record carries a source-unique ID (e.g. `source` + `sourceId`) to deduplicate and import on; keep text in its original encoding.

## Pipeline

1. **Scrape** a sample of 50–200 records first; check it has the schema's required fields. Missing fields → fix the extractor selectors and re-run the sample.
2. **Validate** every line (JSON parse + schema): 0 parse errors, all required fields. Isolate failing lines and inspect their source HTML.
3. **Dry-run** the import against a staging target: counts within ±5% of expectation, no duplicates. Otherwise reset staging and adjust the dedupe key.
4. **Snapshot** the target (timestamped export) before writing.
5. **Import** with idempotent upserts keyed on the source ID; restore the snapshot on failure.

Files in this skill

  • REFERENCE.md1.7 KB
  • SKILL.md2.7 KB

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…