Skip to content
Back to skills

Scientific Llm Benchmarks

ASecurity

A comprehensive reference of benchmarks for evaluating large language models on scientific reasoning and discovery.

  • 47 stars
  • 0 votes
  • 0 copies
  • 0 views
  • Added September 11, 2026
researchpythonbashreactgit

Security analysis

A100/100

Pro scans all 5 files and shows the line behind each finding

Scanned September 11, 2026

npx -y skills add akillness/jeo-skills --skill scientific-llm-benchmarks --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Scientific Llm Benchmarks?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for Scientific Llm Benchmarks
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/akillness-scientific-llm-benchmarks/badge)](https://www.skillsdirectory.com/skills/akillness-scientific-llm-benchmarks)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
SKILL.md
---
name: scientific-llm-benchmarks
description: "A comprehensive reference of benchmarks for evaluating large language models on scientific reasoning and discovery."
compatibility: ">"
allowed-tools: Bash Read Write Edit Glob Grep WebFetch
metadata:
  tags: scientific-llm-benchmarks, llm-benchmarks, science-llms, evaluation, awesome-list, reasoning, discovery, stem
  version: 1.0.0
  source: "https://github.com/subinium/Awesome-Scientific-LLM-Benchmarks"
  license: MIT
---

# Scientific LLM Benchmarks

This skill provides references to benchmarks used for evaluating large language models on scientific reasoning and discovery. The data comes from the Awesome-Scientific-LLM-Benchmarks repository.

## Contents
The complete benchmark list is stored locally within this skill:
- **References List:** `references/benchmarks.md`
- **Data (YAML format):** `data/benchmarks.yaml` 

## Benchmark Domains Covered
- **General / Multi-domain Science:** Cross-disciplinary STEM reasoning benchmarks.
- **Mathematics:** Arithmetic, competition, olympiad, and frontier / formal-proof mathematics.
- **Physics and Astronomy:** Physics olympiad, graduate physics, computational physics, and astronomy.
- **Chemistry:** Molecular property, reaction, retrosynthesis, safety, and chemical knowledge.
- **Materials Science:** Crystals, materials property prediction, and materials-science knowledge.
- **Biology and Life Sciences:** Genomics, proteins, bioinformatics agents, protocols, and research biology.
- **Agentic Science and AI Research:** LLM agents that write research code, run data analyses, attempt autonomous discovery, and conduct ML/AI research.

## Helper Scripts
Also included are python scripts inside `scripts/`:
- `generate_readme.py`: Regenerates the markdown tables and list from `data/benchmarks.yaml`.
- `fetch_examples.py`: Fetches real sample rows from HuggingFace dataset URLs specified in the dataset metadata.

Files in this skill

  • SKILL.md1.9 KB
  • data/benchmarks.yaml85.7 KB
  • references/benchmarks.md54.4 KB
  • scripts/fetch_examples.py4.1 KB
  • scripts/generate_readme.py7.9 KB

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…