Skip to content
Back to skills

Dedupe

ASecurity

Deduplication reference — exact matching, fuzzy matching, hash-based dedup, bloom filters, and data quality. Use when removing duplicate records, files, or data entries.

  • 12 stars
  • 0 votes
  • 0 copies
  • 1 view
  • Added September 7, 2026
toolsgobashsqlgitdatabase

Works with

  • cli

Security analysis

A100/100

Pro scans all 2 files and shows the line behind each finding

Scanned September 7, 2026

npx -y skills add bytesagain/ai-skills --skill dedupe --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Dedupe?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for Dedupe
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/bytesagain-dedupe/badge)](https://www.skillsdirectory.com/skills/bytesagain-dedupe)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
SKILL.md
---
name: "dedupe"
version: "1.0.0"
description: "Deduplication reference — exact matching, fuzzy matching, hash-based dedup, bloom filters, and data quality. Use when removing duplicate records, files, or data entries."
author: "BytesAgain"
homepage: "https://bytesagain.com"
source: "https://github.com/bytesagain/ai-skills"
tags: [dedupe, deduplication, data-quality, hash, fuzzy-match, etl, atomic]
category: "atomic"
---

# Dedupe — Data Deduplication Reference

Quick-reference skill for deduplication strategies, algorithms, and data quality patterns.

## When to Use

- Removing duplicate rows from datasets or databases
- Deduplicating files in storage systems
- Implementing fuzzy matching for near-duplicate detection
- Choosing between exact and probabilistic dedup methods
- Building ETL pipelines with deduplication stages

## Commands

### `intro`

```bash
scripts/script.sh intro
```

Overview of deduplication — types, strategies, and tradeoffs.

### `exact`

```bash
scripts/script.sh exact
```

Exact deduplication — hash-based, key-based, and sorting approaches.

### `fuzzy`

```bash
scripts/script.sh fuzzy
```

Fuzzy deduplication — similarity measures, blocking, and record linkage.

### `files`

```bash
scripts/script.sh files
```

File-level deduplication — fdupes, jdupes, rdfind, and storage dedup.

### `algorithms`

```bash
scripts/script.sh algorithms
```

Dedup algorithms — bloom filters, HyperLogLog, MinHash, SimHash.

### `sql`

```bash
scripts/script.sh sql
```

SQL deduplication patterns — ROW_NUMBER, DISTINCT, GROUP BY strategies.

### `cli`

```bash
scripts/script.sh cli
```

Command-line dedup tools — sort, uniq, awk, and stream processing.

### `checklist`

```bash
scripts/script.sh checklist
```

Deduplication quality checklist and validation steps.

### `help`

```bash
scripts/script.sh help
```

### `version`

```bash
scripts/script.sh version
```

## Configuration

| Variable | Description |
|----------|-------------|
| `DEDUPE_DIR` | Data directory (default: ~/.dedupe/) |

---

*Powered by BytesAgain | bytesagain.com | hello@bytesagain.com*

Files in this skill

  • SKILL.md2.1 KB
  • scripts/script.sh16.5 KB

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…