Skip to content
Back to skills

Genotex Benchmark Guide

ASecurity

Benchmark for LLM agents on gene expression data analysis

  • 3,639 stars
  • 0 votes
  • 0 copies
  • 2 views
  • Added June 6, 2026
researchpythongoexpresstestinggitdatabase

Security analysis

A100/100

Scanned June 6, 2026

npx -y skills add brycewang-stanford/Auto-Empirical-Research-Skills --skill genotex-benchmark-guide --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Genotex Benchmark Guide?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for Genotex Benchmark Guide
[![Security: A β€” Skills Directory](https://www.skillsdirectory.com/api/skills/brycewang-stanford-genotex-benchmark-guide/badge)](https://www.skillsdirectory.com/skills/brycewang-stanford-genotex-benchmark-guide)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
SKILL.md
---
name: genotex-benchmark-guide
description: "Benchmark for LLM agents on gene expression data analysis"
metadata:
  openclaw:
    emoji: "🧫"
    category: "domains"
    subcategory: "biomedical"
    keywords: ["GenoTEX", "gene expression", "benchmark", "LLM agent", "bioinformatics", "GEO"]
    source: "https://github.com/Liu-Hy/GenoTEX"
---

# GenoTEX Benchmark Guide

## Overview

GenoTEX is a benchmark for evaluating LLM-based agents on gene expression data analysis tasks. It provides curated datasets from GEO (Gene Expression Omnibus) with ground-truth analysis pipelines, testing agents on data preprocessing, differential expression, enrichment analysis, and biological interpretation. Published at MLCB 2025 as an oral presentation.

## Benchmark Structure

```
GenoTEX Benchmark
β”œβ”€β”€ Data Collection
β”‚   └── Curated GEO datasets with ground truth
β”œβ”€β”€ Task Categories
β”‚   β”œβ”€β”€ Data preprocessing (QC, normalization)
β”‚   β”œβ”€β”€ Differential expression analysis
β”‚   β”œβ”€β”€ Gene set enrichment analysis
β”‚   β”œβ”€β”€ Clustering and classification
β”‚   └── Biological interpretation
β”œβ”€β”€ Evaluation
β”‚   β”œβ”€β”€ Code correctness (executes without error)
β”‚   β”œβ”€β”€ Statistical validity (appropriate tests)
β”‚   β”œβ”€β”€ Result accuracy (vs ground truth)
β”‚   └── Interpretation quality (biological insight)
└── Baselines
    β”œβ”€β”€ GPT-4 agent
    β”œβ”€β”€ Claude agent
    └── Domain-specific fine-tuned models
```

## Usage

```python
from genotex import GenoTEXBenchmark

bench = GenoTEXBenchmark()

# List available tasks
tasks = bench.list_tasks()
for task in tasks[:5]:
    print(f"Task: {task.id}")
    print(f"  Dataset: {task.geo_accession}")
    print(f"  Category: {task.category}")
    print(f"  Difficulty: {task.difficulty}")

# Get a specific task
task = bench.get_task("GSE12345_DEG")
print(f"Description: {task.description}")
print(f"Input files: {task.input_files}")
print(f"Expected output: {task.expected_output_type}")
```

## Running Evaluations

```python
# Evaluate an agent on GenoTEX
from genotex import evaluate_agent

results = evaluate_agent(
    agent_fn=my_agent_function,
    tasks="all",            # or specific task IDs
    timeout_per_task=300,   # seconds
)

print(f"Tasks completed: {results.completed}/{results.total}")
print(f"Code correctness: {results.code_correct_rate:.1%}")
print(f"Statistical validity: {results.stats_valid_rate:.1%}")
print(f"Result accuracy: {results.accuracy:.3f}")
```

## Task Examples

```python
# Example: Differential Expression Analysis
task = {
    "id": "GSE12345_DEG",
    "description": "Identify differentially expressed genes "
                   "between treatment and control groups in "
                   "this RNA-seq dataset.",
    "input": "GSE12345_counts.csv",  # Raw count matrix
    "metadata": "GSE12345_metadata.csv",  # Sample info
    "expected": {
        "method": "DESeq2 or limma-voom",
        "output": "DEG table with log2FC, p-value, adj.p",
        "ground_truth": "GSE12345_deg_truth.csv",
    },
}

# Example: Gene Set Enrichment
task = {
    "id": "GSE12345_GSEA",
    "description": "Perform gene set enrichment analysis on "
                   "the DEGs and identify enriched pathways.",
    "input": "GSE12345_deg_results.csv",
    "expected": {
        "method": "fgsea, clusterProfiler, or enrichR",
        "output": "Enriched pathways with NES and FDR",
    },
}
```

## Use Cases

1. **Agent evaluation**: Test bioinformatics agents on real tasks
2. **Method comparison**: Compare LLM agents on genomics
3. **Benchmark development**: Extend with new GEO datasets
4. **Teaching**: Standard tasks for bioinformatics education
5. **Tool development**: Test new analysis pipelines

## References

- [GenoTEX GitHub](https://github.com/Liu-Hy/GenoTEX)
- [GEO Database](https://www.ncbi.nlm.nih.gov/geo/)
- [MLCB 2025](https://mlcb.github.io/)

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…