Skip to content
Back to skills

Bio Fasta Database Curator

ASecurity

Curate and validate FASTA or FAA databases. Use when standardizing headers, merging or deduplicating references, converting GenBank, or preparing BLAST, MMseqs2, or HMM inputs.

  • 9 stars
  • 0 votes
  • 0 copies
  • 0 views
  • Added September 27, 2026
ai-agentspythonbashdatabase

Works with

  • mcp

Security analysis

A100/100

Pro scans all 4 files and shows the line behind each finding

Scanned October 2, 2026

npx -y skills add fmschulz/omics-skills --skill bio-fasta-database-curator --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Bio Fasta Database Curator?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for Bio Fasta Database Curator
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/fmschulz-bio-fasta-database-curator/badge)](https://www.skillsdirectory.com/skills/fmschulz-bio-fasta-database-curator)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
SKILL.md
---
name: bio-fasta-database-curator
description: Curate and validate FASTA or FAA databases. Use when standardizing headers, merging or deduplicating references, converting GenBank, or preparing BLAST, MMseqs2, or HMM inputs.
---

# FASTA Database Curator

Curate FASTA/FAA reference databases: standardize headers, merge files, remove duplicates, and validate the result for BLAST, DIAMOND, MMseqs2, or HMMER. Every transformation leaves a header mapping and a deduplication report.

Supplementary version-grounded tool notes: [tools.md](tools.md).

## Header Format Standards

### Recommended Format

Use pipe-separated fields with a consistent source prefix:

```
>PREFIX|ACCESSION|DESCRIPTION
SEQUENCE...
```

**Examples:**
```
>VP|Mavirus_MCP|Major_capsid_protein_[Virophage]
>PLV|NC_021333_1|Polinton-like_virus_hypothetical_protein
>NCLDV|YP_009173877.1|DNA_polymerase_[Marseilleviridae]
```

### What the bundled script does

`scripts/curate_fasta.py` keeps the whole raw header, replaces each whitespace run and each character outside `A-Za-z0-9_.|:-` with `_`, collapses repeated `_`, and prepends `PREFIX|` when `--prefix` is set:

```
>seq 1 alpha protein          ->  >REF|seq_1_alpha_protein
>NC_021333.1 hypothetical     ->  >REF|NC_021333.1_hypothetical
```

Split accession and description into separate pipe fields with a custom rule, and keep the mapping table for it as well.

BLAST `makeblastdb -parse_seqids` reads `|` as NCBI seq-id syntax and fails on custom prefixes such as `VP|...` ("Could not construct seq-id"). Build pipe-delimited databases without `-parse_seqids`. DIAMOND, MMseqs2, and HMMER keep the full first word of the header.

## Quick Reference

| Task | Action |
|------|--------|
| Inspect database | Count records, sample headers, check whitespace and length distribution before changing anything. |
| Standardize headers | Define deterministic transformation rules and preserve original-to-new ID mapping. |
| Merge or deduplicate | Decide whether duplicates are removed by ID, sequence, or both, then report what changed. |
| Validate output | Re-count records, verify FASTA syntax, and write database statistics. |
| Run the bundled curator | `uv run --script skills/bio-fasta-database-curator/scripts/curate_fasta.py input.fasta --output curated.fasta --prefix REF --deduplicate both` |
| Build search databases | Commands for `makeblastdb`, `diamond makedb`, `mmseqs createdb`, and `hmmpress` in [tools.md](tools.md). |

## Instructions

Use `scripts/curate_fasta.py` for routine FASTA curation. It parses raw headers
before any library can truncate them, uppercases sequences, uses SHA-256
sequence digests, refuses empty inputs, empty records, and existing outputs,
and writes a header mapping and a JSON deduplication report. A retained record
whose curated ID repeats an earlier one gets a `|record_N` suffix. Keep the
snippets below for custom transformations only.

### Step 1: Analyze Input Database

First, understand what you're working with:

```bash
# Count sequences
grep -c "^>" database.fasta

# Sample headers (first 20)
grep "^>" database.fasta | head -20

# Check for problematic characters
grep "^>" database.fasta | grep -E "[\t ]" | head -10

# Sequence length distribution
awk '/^>/ {if (seq) print length(seq); seq=""} !/^>/ {seq=seq$0} END {print length(seq)}' database.fasta | sort -n | uniq -c
```

### Step 2: Curate with the Bundled Script

The script standardizes headers, merges multiple inputs, deduplicates, and writes the mapping and report in one pass:

```bash
uv run --script skills/bio-fasta-database-curator/scripts/curate_fasta.py \
  input1.fasta input2.fasta \
  --output curated.fasta \
  --mapping header_mapping.tsv \
  --report dedup_report.json \
  --prefix REF \
  --deduplicate both        # none | id | sequence | both
```

Without `--mapping` and `--report`, the script writes `curated.fasta.mapping.tsv` and `curated.fasta.report.json` next to the output.

Write custom Biopython transformations only when a rule falls outside the script's flags, and keep the original-to-new ID mapping in that case too.

### Step 3: Generate Statistics and Validate

Use SeqKit (versions and more commands in [tools.md](tools.md)) to parse every record, check the alphabet, and summarize; count prefixes from the headers:

```bash
seqkit stats -a -T curated.fasta              # counts, length distribution, sequence type, GC(%)
seqkit seq -t protein --validate-seq curated.fasta > /dev/null   # nonzero exit on an invalid residue; -t dna for nucleotides
seqkit grep -nrp " " curated.fasta | head     # must return nothing: no whitespace in headers
grep '^>' curated.fasta | cut -c2- | cut -d'|' -f1 | sort | uniq -c   # records per prefix
```

## Input Requirements

- One or more FASTA, FAA, FNA, FFN, or GenBank files.
- Desired header convention, prefix policy, and duplicate-removal rule.
- Taxonomy labels or accession metadata when headers need biological grouping.
- Downstream tool constraints, such as BLAST, DIAMOND, MMseqs2, HMMER, or pyhmmer header behavior.

## Output

- Curated FASTA/FAA database with stable identifiers.
- Header mapping table from original IDs to curated IDs.
- Deduplication report with retained and removed records.
- Summary statistics for record count, length distribution, sequence alphabet, prefix/taxonomy counts, and GC content when nucleotide sequences are used.
- Validation notes documenting any skipped transformations or unresolved IDs.

## Quality Gates

- [ ] Every output header is unique and contains no whitespace.
- [ ] Original-to-curated ID mapping is written before destructive transformations.
- [ ] Duplicate policy is explicit: by ID, by sequence, or by both.
- [ ] FASTA parser can read the curated database end-to-end.
- [ ] Record counts before and after curation match the deduplication and filtering report.

## Format Conversions

### GenBank to FASTA

```python
from Bio import SeqIO


def genbank_to_fasta(input_gb: str, output_fasta: str):
    """Convert GenBank format to FASTA."""
    records = SeqIO.parse(input_gb, "genbank")
    count = SeqIO.write(records, output_fasta, "fasta")
    return count
```

### Multi-line to Single-line FASTA

```bash
seqkit seq -w 0 multi.fasta > single.fasta
```

### Extract CDS from GenBank

```python
def extract_cds_proteins(input_gb: str, output_faa: str):
    """Extract CDS translations from GenBank file."""
    with open(output_faa, 'w') as out:
        for record in SeqIO.parse(input_gb, "genbank"):
            for feature in record.features:
                if feature.type == "CDS":
                    if "translation" in feature.qualifiers:
                        protein = feature.qualifiers["translation"][0]
                        locus = feature.qualifiers.get("locus_tag", ["unknown"])[0]
                        product = feature.qualifiers.get("product", ["unknown"])[0]
                        product = "_".join(product.split())  # no whitespace in IDs
                        out.write(f">{locus}|{product}\n{protein}\n")
```

## Best Practices

- Define a taxonomy prefix scheme and stick to it (e.g. `VP|` virophages, `PLV|` polinton-like viruses, `NCLDV|` NCLDVs, `MIRUS|` Mirus viruses).
- Keep the header mapping and deduplication report with the database so every transformation stays auditable.
- Re-run statistics after processing and compare against the pre-curation counts.

## Examples

```bash
uv run --script skills/bio-fasta-database-curator/scripts/curate_fasta.py \
  fixtures/mixed-headers.fasta \
  --output curated.fasta \
  --prefix REF \
  --deduplicate both
```

```
User: "Standardize the headers in virophage_raw.fasta and remove duplicates"

1. Analyze input: count records, sample headers, find whitespace and duplicate accessions.
2. Define rules: `VP|` prefix, whitespace to underscores, deduplicate by ID and sequence.
3. Run curate_fasta.py with --prefix VP --deduplicate both.
4. Validate: record counts match the report (input = retained + removed),
   no whitespace in headers, every record parses and is non-empty.
5. Report statistics from seqkit stats -a and the deduplication report.
```

## Troubleshooting

### Whitespace in Headers
**Problem:** BLAST/MMseqs2 truncate at first whitespace
**Solution:** Replace spaces with underscores or pipes

### Duplicate IDs
**Problem:** Same accession from different sources
**Solution:** Add source prefix to disambiguate

### Invalid Characters
**Problem:** Non-standard amino acid codes
**Solution:** Replace with X or remove sequences

### Mixed Case Sequences
**Problem:** Inconsistent case in sequences
**Solution:** Standardize to uppercase

Files in this skill

  • SKILL.md8.8 KB
  • fixtures/mixed-headers.fasta92 B
  • scripts/curate_fasta.py7.3 KB
  • tools.md2.9 KB

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…