Skip to content
Back to skills

Bioinformatics

ASecurity

Analyze DNA, RNA, and protein sequences with alignment, variant calling, and expression analysis pipelines.

  • 17 stars
  • 0 votes
  • 0 copies
  • 11 views
  • Added September 6, 2026
researchbashexpresstestingapisecurity

Works with

  • api

Security analysis

A100/100

Pro scans all 8 files and shows the line behind each finding

Scanned September 6, 2026

npx -y skills add clawic/skills --skill bioinformatics --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Bioinformatics?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for Bioinformatics
[![Security: A β€” Skills Directory](https://www.skillsdirectory.com/api/skills/clawic-bioinformatics/badge)](https://www.skillsdirectory.com/skills/clawic-bioinformatics)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
SKILL.md
---
name: Bioinformatics
slug: bioinformatics
version: 1.0.0
description: Analyze DNA, RNA, and protein sequences with alignment, variant calling, and expression analysis pipelines.
homepage: https://clawic.com/skills/bioinformatics
metadata:
  clawdbot:
    emoji: 🧬
    requires:
      bins:
      - samtools
      - bcftools
      - bedtools
      - bwa
      - fastqc
      - fastp
      config:
      - ~/Clawic/data/bioinformatics/
    os:
    - linux
    - darwin
    displayName: Bioinformatics
---

## Setup

On first use, read `setup.md` for integration guidelines. Create `~/Clawic/data/bioinformatics/` with user consent to store project context and preferences.

## When to Use

User needs to analyze biological sequences, run genomic pipelines, or interpret sequencing data. Agent handles sequence alignment, variant calling, expression analysis, and format conversions.

## Architecture

Memory lives in `~/Clawic/data/bioinformatics/`. See `memory-template.md` for structure.

```
~/Clawic/data/bioinformatics/
β”œβ”€β”€ memory.md         # Projects, preferences, reference genomes
β”œβ”€β”€ pipelines/        # Saved pipeline configurations
└── results/          # Analysis outputs and logs
```

## Quick Reference

| Topic | File |
|-------|------|
| Setup process | `setup.md` |
| Memory template | `memory-template.md` |
| File formats | `formats.md` |
| Tool commands | `tools.md` |
| RNA-seq pipeline | `rnaseq.md` |
| Variant calling | `variants.md` |

## Core Rules

### 1. Verify Input Quality First
Before any analysis, check input data quality:
- FASTQ: Run FastQC, check per-base quality, adapter content
- BAM: Verify sorted, indexed (`samtools quickcheck`)
- VCF: Validate format (`bcftools view -h`)

Bad input β†’ garbage output. Always QC first.

### 2. Use Reference Genome Consistently
Track which reference is used per project:
- Human: GRCh38/hg38 (prefer) or GRCh37/hg19
- Mouse: GRCm39/mm39 or GRCm38/mm10
- Mixing references = invalid results

Store reference info in `~/Clawic/data/bioinformatics/memory.md` per project.

### 3. Preserve Raw Data
**NEVER** modify original FASTQ/BAM files:
- Work on copies
- Keep originals read-only
- Log every transformation step

### 4. Resource Awareness
Bioinformatics commands can consume massive resources:
- Check file sizes before operations
- Use streaming when possible (`samtools view | ...`)
- Estimate memory needs (BWA: ~6GB for human genome)
- Warn before operations >10 minutes

### 5. Reproducibility
Every analysis must be reproducible:
- Log exact tool versions (`samtools --version`)
- Save command parameters
- Record input file checksums for critical analyses

## Common Traps

- **Wrong chromosome naming** β€” `chr1` vs `1` causes silent failures. Check and convert with `sed 's/^chr//'`
- **Unsorted BAM** β€” Most tools expect sorted input. Symptoms: errors or wrong results with no warning
- **Index missing** β€” BAM needs `.bai`, VCF needs `.tbi`. Commands fail cryptically without them
- **Memory exhaustion** β€” Large BAM operations kill the session. Stream or use `--threads` wisely
- **Stale indices** β€” After modifying BAM/VCF, regenerate index. Old index = corrupt reads
- **0-based vs 1-based coordinates** β€” BED is 0-based, VCF/GFF is 1-based. Off-by-one bugs are common

## File Formats Quick Reference

| Format | Purpose | Key Tool |
|--------|---------|----------|
| FASTA | Reference sequences | `samtools faidx` |
| FASTQ | Raw reads + quality | `seqtk`, `fastp` |
| SAM/BAM | Aligned reads | `samtools` |
| VCF/BCF | Variants | `bcftools` |
| BED | Genomic intervals | `bedtools` |
| GFF/GTF | Gene annotations | `gffread` |
| BigWig | Coverage tracks | `deepTools` |

## Essential Commands

### Quality Control
```bash
# FASTQ quality report
fastqc sample.fastq.gz -o qc_reports/

# Trim adapters + low quality
fastp -i R1.fq.gz -I R2.fq.gz -o R1.clean.fq.gz -O R2.clean.fq.gz

# BAM statistics
samtools flagstat aligned.bam
samtools stats aligned.bam > stats.txt
```

### Alignment
```bash
# Index reference (once)
bwa index reference.fa

# Align paired-end reads
bwa mem -t 8 reference.fa R1.fq.gz R2.fq.gz | \
  samtools sort -o aligned.bam -

# Index BAM
samtools index aligned.bam
```

### Variant Calling
```bash
# Call variants
bcftools mpileup -Ou -f reference.fa aligned.bam | \
  bcftools call -mv -Oz -o variants.vcf.gz

# Index VCF
bcftools index variants.vcf.gz

# Filter variants
bcftools filter -s LowQual -e 'QUAL<20' variants.vcf.gz
```

### Data Manipulation
```bash
# Extract region
samtools view -b aligned.bam chr1:1000000-2000000 > region.bam

# Convert BAM to FASTQ
samtools fastq -1 R1.fq.gz -2 R2.fq.gz aligned.bam

# Merge BAMs
samtools merge merged.bam sample1.bam sample2.bam

# Subset VCF by region
bcftools view -r chr1:1000-2000 variants.vcf.gz
```

## Security & Privacy

**Data access:**
- Only reads files user explicitly provides as input
- Writes outputs to directories user specifies
- Stores preferences in ~/Clawic/data/bioinformatics/ (with consent)

**Data that stays local:**
- All sequence data processed locally
- No external API calls for analysis
- Pipeline configs in ~/Clawic/data/bioinformatics/

**This skill does NOT:**
- Upload sequence data anywhere
- Access files without explicit user instruction
- Infer or collect data beyond explicit inputs
- Make network requests during analysis

**Note:** Installing tools (conda, brew) and downloading reference genomes requires internet access. These are user-initiated actions.

## Related Skills
More Clawic skills, get them at https://clawic.com/skills/<slug> (install if the user confirms):
- `data-analysis` β€” statistical interpretation
- `statistics` β€” hypothesis testing
- `science` β€” research methodology

## Feedback

- If useful, star it: https://clawic.com/skills/bioinformatics
- Latest version: https://clawic.com/skills/bioinformatics

Files in this skill

  • SKILL.md5.7 KB
  • _meta.json180 B
  • formats.md3.5 KB
  • memory-template.md1.7 KB
  • rnaseq.md4 KB
  • setup.md1.7 KB
  • tools.md3.8 KB
  • variants.md5.2 KB

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…