Skip to content
Back to skills

Benchmark Model Runtime

ASecurity

Benchmark model runtimes across latency, throughput, memory, energy, load time, size, and stability. Use when comparing Core AI, Core ML, MLX, ExecuTorch, PyTorch, quantization, devices, or packaging.

  • 7 stars
  • 0 votes
  • 0 copies
  • 1 view
  • Added September 5, 2026
ai-agentsbackendperformance

Security analysis

A100/100

Pro scans all 3 files and shows the line behind each finding

Scanned September 5, 2026

npx -y skills add gaelic-ghost/socket --skill benchmark-model-runtime --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Benchmark Model Runtime?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for Benchmark Model Runtime
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/gaelic-ghost-benchmark-model-runtime/badge)](https://www.skillsdirectory.com/skills/gaelic-ghost-benchmark-model-runtime)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
SKILL.md
---
name: benchmark-model-runtime
description: Benchmark model runtimes across latency, throughput, memory, energy, load time, size, and stability. Use when comparing Core AI, Core ML, MLX, ExecuTorch, PyTorch, quantization, devices, or packaging.
---

# Benchmark Model Runtime

## Define A Fair Workload

Pin the exact model artifact, tokenizer/template, runtime and version, device and OS, precision, cache policy, batch size, prompt-length buckets, generated-token target, sampling settings, and measurement tool. Compare numerical or behavioral parity before performance.

## Workflow

1. Verify each artifact produces acceptable outputs on the same small parity set.
2. Separate cold load, warm load, prompt processing, time to first token, decode throughput, and end-to-end latency.
3. Measure peak and steady memory; include model, cache, runtime, and process overhead consistently.
4. Stabilize device power, charging, background load, and thermal state. Record deviations instead of silently rerunning only slow samples.
5. Warm up separately, then run enough measured repetitions to report median and tail percentiles.
6. Sweep representative prompt lengths, output lengths, and batch/concurrency levels.
7. Record failures, fallback execution, recompilation, memory pressure, and thermal throttling.
8. Measure energy with an appropriate system tool when the decision depends on battery or sustained deployment.
9. Retain raw samples and summarize them with units, sample counts, and uncertainty.

## Apple Runtime Checks

- Confirm delegated operator coverage and fallback behavior rather than assuming the named backend ran the whole graph.
- Distinguish compilation/conversion time from load and inference time.
- For stateful generation, include key-value cache initialization, update, and memory growth.
- Treat ExecuTorch MLX results as revision-specific while the upstream backend remains experimental.

## References

Use `references/runtime-benchmarking.md` for metric definitions and reporting requirements.

Files in this skill

  • SKILL.md2 KB
  • agents/openai.yaml304 B
  • references/runtime-benchmarking.md1.1 KB

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…