Skip to content
Back to skills

Alterlab Gget

ASecurity

Run fast one-liner queries to 20+ bioinformatics databases from the gget CLI or Python — gene info (Ensembl), BLAST, AlphaFold structures, Enrichr enrichment, and more. Use for quick interactive lookups of genes, sequences, structures, or pathways — for batch processing or advanced BLAST use biopython, for multi-database Python workflows use bioservices. Part of the AlterLab Academic Skills suite.

  • 68 stars
  • 0 votes
  • 0 copies
  • 1 view
  • Added May 27, 2026
data-aipythonbashexpressgitapidatabasedocumentation

Works with

  • cli
  • api

Security analysis

A92/100
  • mediumInstalls packages at runtime which could introduce malicious dependencies

Pro scans all 7 files and shows the line behind each finding

Scanned September 23, 2026

npx -y skills add AlterLab-IEU/AlterLab-Academic-Skills --skill alterlab-gget --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Alterlab Gget?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for Alterlab Gget
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/alterlab-ieu-alterlab-gget/badge)](https://www.skillsdirectory.com/skills/alterlab-ieu-alterlab-gget)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
SKILL.md
---
name: alterlab-gget
description: "Run fast one-liner queries to 20+ bioinformatics databases from the gget CLI or Python — gene info (Ensembl), BLAST, AlphaFold structures, Enrichr enrichment, and more. Use for quick interactive lookups of genes, sequences, structures, or pathways — for batch processing or advanced BLAST use biopython, for multi-database Python workflows use bioservices. Part of the AlterLab Academic Skills suite."
license: MIT
allowed-tools: Read Write Edit Bash(python:*) Bash(uv:*)
compatibility: "Install with `uv pip install gget` (0.30.8 as of 2026-09; requires Python >= 3.12). Core modules need no API key or account. cosmic needs a COSMIC account; gpt needs an OpenAI key; alphafold, cellxgene, elm, gpt and cbio need a one-time `gget setup <module>`."
metadata:
    skill-author: AlterLab
    version: "1.1.0"
    last_updated: "2026-09-23"
---

# gget

## Overview

gget is a command-line bioinformatics tool and Python package providing unified access to 20+ genomic databases and analysis methods. Query gene information, sequence analysis, protein structures, expression data, and disease associations through a consistent interface. All gget modules work both as command-line tools and as Python functions.

**Project home:** development moved to the scverse organisation (`github.com/scverse/gget`); the manual stays at `pachterlab.github.io/gget`.

**Important**: The databases queried by gget are continuously updated, which sometimes changes their structure. gget modules are tested automatically on a biweekly basis and updated to match new database structures when necessary.

## Installation

Install gget in a clean virtual environment to avoid conflicts:

```bash
# Install (or upgrade) into a clean environment
uv pip install --upgrade gget

# In Python/Jupyter
import gget
```

## Quick Start

Basic usage pattern for all modules:

```bash
# Command-line
gget <module> [arguments] [options]

# Python
gget.module(arguments, options)
```

Most modules return:
- **Command-line**: JSON (default) or CSV with `-csv` flag
- **Python**: DataFrame or dictionary

Common flags across modules:
- `-o/--out`: Save results to file
- `-q/--quiet`: Suppress progress information
- `-csv`: Return CSV format (command-line only)

## Module Catalog

Pick a module, then see `references/module_examples.md` for worked CLI + Python
examples and `references/module_reference.md` for the full parameter table.

| Module | Purpose | Queried source |
|--------|---------|----------------|
| `ref` | Reference genome download links/metadata | Ensembl |
| `search` | Find genes by name/description | Ensembl |
| `info` | Gene/transcript metadata (~1000 IDs max) | Ensembl, UniProt, NCBI |
| `seq` | Nucleotide/amino-acid sequences (FASTA) | Ensembl |
| `blast` | BLAST against standard databases | NCBI BLAST |
| `blat` | Genomic position of a sequence | UCSC BLAT |
| `muscle` | Multiple sequence alignment | Muscle5 (local) |
| `diamond` | Fast local protein/translated alignment | DIAMOND (local) |
| `pdb` | Experimental protein structures + metadata | RCSB PDB |
| `alphafold` | Predict 3D protein structure (setup req.) | AlphaFold2 (local) |
| `elm` | Eukaryotic linear motifs (setup req.) | ELM |
| `archs4` | Correlated genes / tissue expression | ARCHS4 |
| `cellxgene` | Single-cell RNA-seq (setup req.) | CZ CELLxGENE Census |
| `enrichr` | Ontology/pathway enrichment | Enrichr |
| `bgee` | Orthologs and expression | Bgee |
| `opentargets` | Disease/drug associations | OpenTargets |
| `cbio` | Cancer genomics heatmaps | cBioPortal |
| `cosmic` | Somatic cancer mutations (license/account) | COSMIC |
| `mutate` | Generate mutated sequences | local |
| `virus` | Download filtered virus genome datasets | NCBI Virus |
| `g2p` | Residue-level structural/functional annotations | Genomics 2 Proteins portal |
| `gene_expression` | Mean/variance of normalized expression per partition | 8cubeDB |
| `psi_block` | ψ_block block-level specificity scores | 8cubeDB |
| `specificity` | Gene-level ψ / ζ specificity statistics | 8cubeDB |
| `gpt` | Natural-language text generation (setup req.) | OpenAI API |
| `setup` | Install third-party deps for a module | local |

`cbio` is exposed in Python as `gget.cbio_search()` and `gget.cbio_plot()`.

**Setup-required modules** (`gget setup <module>` before first use):
`alphafold` (~4GB params, needs `uv pip install openmm` first), `cellxgene`,
`elm`, `gpt`, and `cbio`.

## When to Use This Skill

- **Quick interactive lookup** (gene info, BLAST, one structure, one enrichment) →
  use gget directly; see `references/module_examples.md`.
- **Batch processing / advanced BLAST** → use the **biopython** skill.
- **Multi-database Python workflows** → use the **bioservices** skill.

### Does NOT Trigger

| Scenario | Use Instead |
|----------|-------------|
| Local BLAST+ database builds and large CLI searches | `alterlab-blast` |
| Scripted Entrez/SeqIO pipelines and file parsing | `alterlab-biopython` |
| One workflow spanning many web services in Python | `alterlab-bioservices` |
| Serious CELLxGENE Census querying beyond a one-liner | `alterlab-cellxgene` |
| Running AlphaFold properly (complexes, confidence analysis) | `alterlab-alphafold` |
- **Chaining several gget modules into a pipeline** → see `references/workflows.md`
  and the ready-made `scripts/` (gene_analysis, batch_sequence_analysis,
  enrichment_pipeline).

## Best Practices (essentials)

- Use `--limit` to bound large queries; save with `-o/--out` for reproducibility.
- Gene symbols are **case-sensitive** in cellxgene ('PAX7' vs 'Pax7').
- Run `gget setup` before first use of alphafold, cellxgene, elm, gpt.
- Process max ~1000 Ensembl IDs at once with `gget info`.
- Database structures change; keep gget updated: `uv pip install --upgrade gget`.
- Use virtual environments to avoid dependency conflicts.

## Output Formats

- **Command-line**: JSON default; `-csv` for CSV; FASTA (`seq`, `mutate`);
  PDB (`pdb`, `alphafold`); PNG (`cbio plot`).
- **Python**: DataFrame/dict default; `json=True` for JSON; `save=True` or
  `out="filename"` to write; AnnData for `cellxgene`.

## References

- `references/module_examples.md` — worked CLI + Python examples for every module
- `references/module_reference.md` — full parameter tables for all modules
- `references/database_info.md` — queried databases and their update frequencies
- `references/workflows.md` — extended multi-module workflow examples

For additional help:
- Official documentation: https://pachterlab.github.io/gget/
- GitHub issues: https://github.com/pachterlab/gget/issues
- Citation: Luebbert, L. & Pachter, L. (2023). Efficient querying of genomic reference databases with gget. Bioinformatics. https://doi.org/10.1093/bioinformatics/btac836

Part of the AlterLab Academic Skills suite.

Files in this skill

  • SKILL.md24.6 KB
  • references/database_info.md10 KB
  • references/module_reference.md17.2 KB
  • references/workflows.md25.2 KB
  • scripts/batch_sequence_analysis.py5.8 KB
  • scripts/enrichment_pipeline.py7 KB
  • scripts/gene_analysis.py5.6 KB

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…