Skip to content
Back to skills

Golden Dataset

ASecurity

Golden dataset curation, backup/restore, validation with schema checks, duplicate detection, and coverage analysis

  • 6 stars
  • 0 votes
  • 0 copies
  • 1 view
  • Added September 6, 2026
ai-agentspythongosqltestinggitdevopsci/cddocumentation

Works with

  • cli

Security analysis

A100/100

Pro scans all 13 files and shows the line behind each finding

Scanned September 6, 2026

npx -y skills add ArieGoldkin/claude-forge --skill golden-dataset --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Golden Dataset?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for Golden Dataset
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/ariegoldkin-golden-dataset/badge)](https://www.skillsdirectory.com/skills/ariegoldkin-golden-dataset)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
SKILL.md
---
name: golden-dataset
description: "Golden dataset curation, backup/restore, validation with schema checks, duplicate detection, and coverage analysis"
effort: medium
paths:
  - "**/*golden*"
  - "**/*dataset*"
  - "**/*evaluation*"
  - "tests/data/**"
---

# Golden Dataset

**Curate, manage, and validate high-quality test datasets for AI/ML systems**

## Overview

A **golden dataset** is a curated collection of high-quality examples used for regression testing, retrieval evaluation, model benchmarking, and reproducibility. This skill covers the full lifecycle: curating new entries with quality analysis, managing backup/restore operations, and validating data integrity.

### Example Golden Dataset Metrics

| Metric | Value |
|--------|-------|
| Documents | 98 completed |
| Chunks | 415 embedded segments |
| Test queries | 203 with expected results |
| Pass rate | 91.6% retrieval quality |

**Purpose:** Test hybrid search (vector + BM25 + RRF), validate metadata boosting, detect retrieval regressions, benchmark embedding models.

---

## Curation

Quality criteria, workflows, and multi-agent analysis patterns for evaluating and adding documents to the golden dataset.

**Key areas:**
- **Content type classification** -- article, tutorial, research paper, documentation, video transcript, code repository
- **Difficulty stratification** -- trivial, easy, medium, hard, adversarial (based on semantic complexity)
- **Quality dimensions** -- accuracy (0.25), coherence (0.20), depth (0.25), relevance (0.30)
- **Multi-agent pipeline** -- parallel evaluation with Quality Evaluator, Difficulty Classifier, Domain Tagger, Query Generator
- **Duplicate prevention** -- URL check + semantic similarity > 80% threshold

**Detailed patterns:** [references/curation.md](${CLAUDE_SKILL_DIR}/references/curation.md)
**Multi-agent pipeline:** [references/multi-agent-pipeline.md](${CLAUDE_SKILL_DIR}/references/multi-agent-pipeline.md)

---

## Management

Backup, restore, and lifecycle operations for protecting golden dataset integrity.

**Key areas:**
- **Data integrity contracts** -- real canonical URLs required, no placeholders
- **JSON backup strategy** -- version-controlled, human-readable, portable; embeddings excluded and regenerated on restore
- **Restore process** -- load JSON, validate structure, create analyses/chunks, regenerate embeddings, verify integrity
- **CLI usage** -- `poetry run python scripts/backup_golden_dataset.py backup|verify|restore`
- **CI/CD automation** -- GitLab CI pipeline integration for scheduled backups

**Detailed patterns:** [references/management.md](${CLAUDE_SKILL_DIR}/references/management.md)
**Backup/restore code:** [references/backup-restore.md](${CLAUDE_SKILL_DIR}/references/backup-restore.md)
**CI/CD automation:** [references/ci-cd-automation.md](${CLAUDE_SKILL_DIR}/references/ci-cd-automation.md)
**Disaster recovery:** [references/disaster-recovery.md](${CLAUDE_SKILL_DIR}/references/disaster-recovery.md)
**Validation contracts:** [references/validation-contracts.md](${CLAUDE_SKILL_DIR}/references/validation-contracts.md)

---

## Validation

Schema checks, duplicate detection, and coverage analysis for ensuring every entry meets quality standards.

**Key areas:**
- **Schema validation** -- document schema v2.0 (id, title, source_url, content_type, sections) and query schema (id, query, difficulty, expected_chunks, min_score)
- **Validation rules** -- no placeholder URLs, unique IDs, referential integrity, content quality, difficulty distribution
- **Duplicate detection** -- semantic similarity via embeddings + URL normalization
- **Coverage analysis** -- content type balance, difficulty distribution, domain spread

**Detailed patterns:** [references/validation.md](${CLAUDE_SKILL_DIR}/references/validation.md)
**Validation rules:** [references/validation-rules.md](${CLAUDE_SKILL_DIR}/references/validation-rules.md)
**Duplicate detection:** [references/duplicate-detection.md](${CLAUDE_SKILL_DIR}/references/duplicate-detection.md)
**Validation workflows:** [references/validation-workflows.md](${CLAUDE_SKILL_DIR}/references/validation-workflows.md)

---

## Key Decisions

| Decision | Choice | Rationale |
|----------|--------|-----------|
| Backup format | JSON (not SQL dump) | Version-controllable, human-readable, portable across DB versions |
| Embeddings in backup | Excluded | Regenerated on restore with current model; keeps backups small |
| URL contract | Real canonical URLs only | Enables re-fetch, source validation, audit trail |
| Quality threshold | Minimum 0.70 score | Balance between quality and coverage |
| Minimum metadata | 2 tags, 3 test queries | Ensures searchability and testability |
| Difficulty levels | 5 tiers (trivial to adversarial) | Comprehensive retrieval evaluation at all complexity levels |
| Duplicate threshold | 80% semantic similarity | Prevents near-duplicates while allowing related content |

---

## Example Implementation Files

| File | Purpose |
|------|---------|
| `scripts/backup_golden_dataset.py` | Main backup script |
| `data/golden_dataset_backup.json` | JSON backup (version controlled) |
| `data/golden_dataset_metadata.json` | Quick stats |
| `documents_expanded.json` | Expanded document data |
| `queries.json` | Test queries with expected results |
| `source_url_map.json` | URL deduplication index |

---

## Related Skills

- `langfuse-observability` -- tracing patterns for curation pipeline
- `pgvector-search` -- retrieval evaluation and duplicate detection
- `ai-native-development` -- embedding generation for restore
- `devops-deployment` -- CI/CD backup automation

---

**Version:** 1.0.0 (December 2025)
**Issue:** #599

Files in this skill

  • SKILL.md5.6 KB
  • checklists/backup-restore-checklist.md14.3 KB
  • references/backup-restore.md13.7 KB
  • references/ci-cd-automation.md2.2 KB
  • references/curation.md5.9 KB
  • references/disaster-recovery.md2.4 KB
  • references/duplicate-detection.md2.8 KB
  • references/management.md6.5 KB
  • references/multi-agent-pipeline.md12.9 KB
  • references/validation-contracts.md11.5 KB
  • references/validation-rules.md4.4 KB
  • references/validation-workflows.md7.4 KB
  • references/validation.md3.2 KB

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…