Skip to content
Back to skills

Data Prep

ASecurity

"Guides AdelaiDet COCO/PIC/LVIS/text dataset layout, semantic-mask

  • 247 stars
  • 0 votes
  • 0 copies
  • 5 views
  • Added September 8, 2026
tools

Security analysis

A100/100

Pro scans all 6 files and shows the line behind each finding

Scanned September 8, 2026

npx -y skills add VectorSpaceLab/AREX-Skill --skill data-prep --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Data Prep?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for Data Prep
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/vectorspacelab-data-prep/badge)](https://www.skillsdirectory.com/skills/vectorspacelab-data-prep)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
SKILL.md
---
name: "data-prep"
description: "Guides AdelaiDet COCO/PIC/LVIS/text dataset layout, semantic-mask
  generation, dataset registration, mapper expectations, and MEInst mask
  components."
disable-model-invocation: true
metadata:
  disco-role: operating
license: NOASSERTION
---

# data-prep

Use this sub-skill when a task asks how to prepare datasets for AdelaiDet, register custom data, convert PIC/COCO/LVIS annotations, generate thing semantic masks, inspect mapper expectations, or create MEInst mask components.

## Use this route for

- Dataset directory layout and required annotation files.
- COCO instance annotations, PIC person conversion, LVIS/COCO semantic masks, and `thing_train2017`-style outputs.
- Custom dataset registration and `DATASETS.TRAIN` / `DATASETS.TEST` config overrides.
- Text dataset annotation requirements before BAText/ABCNet training.
- MEInst mask encoding/component prerequisites.
- Visual sanity checks before training.

## Do not use this route for

- Environment/build/import failures. Use `../setup-build/SKILL.md`.
- Training command composition after data is ready. Use `../train-eval/SKILL.md`.
- Text lexicon/evaluator semantics. Use `../text-spotting/SKILL.md`.
- Demo rendering only. Use `../demo-visualize/SKILL.md`.
- Checkpoint or ONNX conversion. Use `../export-convert/SKILL.md`.

## Read first

- `references/dataset-preparation.md` for conversion/preparation recipes.
- `references/data-formats.md` for expected annotation shapes and config connections.
- `../../references/model-overview.md` to map model family to data prerequisites.

## Skill-owned scripts

- `scripts/prepare_thing_semantic.py` — self-contained COCO/PIC-style thing semantic mask generator with explicit input/output paths.
- `scripts/gen_pic_person_coco.py` — safer PIC instance/semantic masks to COCO person JSON conversion.
- `scripts/meinst_mask_encoding.py` — helper for planning/checking MEInst mask-component generation inputs.

## Typical workflow

1. Identify the model family and its dataset expectations.
2. Validate the raw dataset layout and annotation JSON.
3. Generate derived artifacts only when the selected config needs them.
4. Visualize a small sample with `../demo-visualize/scripts/visualize_dataset.py`.
5. Return to `train-eval` for launch.

## Decision points

- If a task mentions scene text, Bezier control points, dictionaries, or lexicons, load `text-spotting` too.
- If a config references missing `thing_*` semantic masks, use `scripts/prepare_thing_semantic.py`.
- If FCPose/PIC person data is needed, use `scripts/gen_pic_person_coco.py` or follow the PIC layout reference.
- If MEInst asks for components/PCA artifacts, use the MEInst notes before training.

Files in this skill

  • SKILL.md2.7 KB
  • references/data-formats.md2.2 KB
  • references/dataset-preparation.md2.7 KB
  • scripts/gen_pic_person_coco.py4.6 KB
  • scripts/meinst_mask_encoding.py2 KB
  • scripts/prepare_thing_semantic.py4 KB

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…