Skip to content
Back to skills

Model Benchmarker

ASecurity

Use when: benchmark model latency, throughput, or quality.

  • 5 stars
  • 0 votes
  • 0 copies
  • 1 view
  • Added October 4, 2026
ai-agentspythongobashapiperformance

Works with

  • terminal
  • cli
  • api

Security analysis

A100/100

Scanned October 4, 2026

npx -y skills add openamer/openamer --skill model-benchmarker --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Model Benchmarker?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for Model Benchmarker
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/openamer-model-benchmarker/badge)](https://www.skillsdirectory.com/skills/openamer-model-benchmarker)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
SKILL.md
---
name: model-benchmarker
description: "Use when: benchmark model latency, throughput, or quality."
category: system
triggers:
  - benchmark
  - latenzen
  - model performance
  - throughput test
  - provider comparison
---

# Model Benchmarker

Tests Model-Performance über mehrere Provider (OpenRouter, lokale Ollama-Instanzen).

## Tests

| Test | Beschreibung | Runs | Metrik |
|------|-------------|------|--------|
| **Latenz** | Zeit bis erster Token (kleiner Prompt) | 10 | Median in Sekunden |
| **Durchsatz** | Tokens pro Sekunde (4K Prompt) | 5 | Median tok/s |
| **Qualität** | Antwort auf 4 Testfragen bewertet | 1 | Score 0-100% |

## CLI-Verwendung

```bash
# Einmaliger Test
python model-benchmarker.py --run openrouter/deepseek-v4-flash
python model-benchmarker.py --run "local/qwen3.5:9b"

# Alle Provider testen (Default-Modell pro Provider)
python model-benchmarker.py --all

# Vergleichstabelle aller getesteten Modelle
python model-benchmarker.py --compare
python model-benchmarker.py --compare --json   # JSON-Ausgabe

# Trend-Historie
python model-benchmarker.py --history openrouter
python model-benchmarker.py --history local
```

## Provider-Konfiguration

Liest automatisch aus `config.yaml`:
- `model.default` → OpenRouter-Modell
- `custom_providers` → Lokale Ollama-Instanzen und deren Modelle
- API-Keys aus `.env` (`OPENROUTER_API_KEY`)

## Ergebnisverzeichnis

Alle Ergebnisse in `~/.model-benchmarks/results/<provider>-<model>.json`
mit vollständiger Historie (max 50 Einträge).

## Cron-Job

Wöchentlicher `--all` Run ist als Cron-Job eingerichtet:
- Führt alle Provider-Modelle durch
- Sendet Ergebnisse an Terminal (lokal)
- Kein Agent nötig (`no_agent: true`)

## Abhängigkeiten

Keine — nur Python-Standardbibliothek (urllib, json, statistics, etc.)

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…