Skip to content
Back to skills

Evaluating Machine Learning Models

ASecurity

Evaluate trained machine learning models with the right metrics and comparison logic. Use for benchmark review, threshold selection, calibration, validation, and model comparison; not for feature engineering or leakage auditing.

  • 3,126 stars
  • 0 votes
  • 0 copies
  • 0 views
  • Added May 29, 2026
ai-agentsgobashtesting

Security analysis

A100/100

Pro scans all 8 files and shows the line behind each finding

Scanned May 29, 2026

npx -y skills add foryourhealth111-pixel/Vibe-Skills --skill evaluating-machine-learning-models --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Evaluating Machine Learning Models?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for Evaluating Machine Learning Models
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/foryourhealth111-pixel-evaluating-machine-learning-models/badge)](https://www.skillsdirectory.com/skills/foryourhealth111-pixel-evaluating-machine-learning-models)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
SKILL.md
---
name: evaluating-machine-learning-models
description: |
  Evaluate trained machine learning models with the right metrics and comparison logic.
  Use for benchmark review, threshold selection, calibration, validation, and model comparison; not for feature engineering or leakage auditing.
allowed-tools: Read, Write, Edit, Grep, Glob, Bash(cmd:*)
version: 1.0.0
author: Jeremy Longshore <jeremy@intentsolutions.io>
license: MIT
---
# Model Evaluation Suite

Use this skill when the model exists and the question is whether it is good enough.

## Overview

This skill focuses on choosing and interpreting the right evaluation metrics for the problem, then comparing candidate models or thresholds.

## When to Use This Skill

- Comparing candidate models with consistent metrics
- Reviewing precision/recall/F1/AUC, regression error, calibration, or ranking quality
- Stress-testing validation strategy before deployment or publication

## Not For / Boundaries

- Building the training pipeline itself: use `scikit-learn` for classical modeling or `ml-pipeline-workflow` for end-to-end workflow ownership
- Engineering features: use `preprocessing-data-with-automated-pipelines`
- Checking train/test contamination: use `ml-data-leakage-guard`

## Typical Outputs

- Metric suite recommendations
- Model comparison tables
- Notes on threshold tradeoffs, calibration, and validation weaknesses

## Related Skills

- `scikit-learn` for class-level error breakdowns and confusion matrices
- `scientific-reporting` when the evaluation must become a deliverable

Files in this skill

  • SKILL.md1.5 KB
  • assets/README.md359 B
  • assets/visualization_script.py5.5 KB
  • references/README.md397 B
  • scripts/README.md407 B
  • scripts/data_loader.py2.8 KB
  • scripts/evaluate_model.py2.8 KB
  • scripts/metrics_calculator.py2.8 KB

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…