Skip to content
Back to skills

Evaluation And Certification

ASecurity

"Compute ART robustness/privacy metrics, run evaluation objects,

  • 247 stars
  • 0 votes
  • 0 copies
  • 2 views
  • Added September 8, 2026
toolspythonbashrailsbackendsecurity

Works with

  • cli

Security analysis

A100/100

Pro scans all 5 files and shows the line behind each finding

Scanned September 8, 2026

npx -y skills add VectorSpaceLab/AREX-Skill --skill evaluation-and-certification --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Evaluation And Certification?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for Evaluation And Certification
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/vectorspacelab-evaluation-and-certification/badge)](https://www.skillsdirectory.com/skills/vectorspacelab-evaluation-and-certification)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
SKILL.md
---
name: evaluation-and-certification
description: "Compute ART robustness/privacy metrics, run evaluation objects,
  route SummaryWriter logging, and perform gradient checks and
  certification/verification workflows."
disable-model-invocation: true
metadata:
  disco-role: operating
license: MIT
---

# evaluation-and-certification

Use this sub-skill when the model is already wrapped and you need to measure robustness, privacy leakage, gradient health, or certification/verification results.

## Use this for

- ART robustness metrics: `adversarial_accuracy`, `empirical_robustness`, `loss_sensitivity`, `clever_u`, `wasserstein_distance`.
- Privacy leakage metrics and threshold helpers: `PDTP`, `SHAPr`, `ComparisonType`, and the ROC helpers in `art.metrics.privacy`.
- Gradient checks: `loss_gradient_check`.
- Evaluation objects: `SecurityCurve`, `GreatScorePyTorch`.
- Tree robustness verification and certified/tree-specific workflows.
- Certification wrappers: randomized smoothing, de-randomized smoothing, DeepZ, and interval/IBP classifiers.
- SummaryWriter/TensorBoard routing for evaluation objects or attack/certification telemetry.

## Start here

1. Read [references/metrics-evaluations-certification.md](references/metrics-evaluations-certification.md) to pick the metric, evaluation object, or certifier and the correct import path.
2. If you need TensorBoard output, follow the SummaryWriter routing table there before enabling logging.
3. Run the bundled smoke script with `--help` first, then choose the tree mode you need:

   ```bash
   python scripts/smoke_metrics_tree.py --help
   python scripts/smoke_metrics_tree.py --tree-mode verify
   python scripts/smoke_pytorch_adv_accuracy.py --attack fgm --json
   ```

   If a tiny tree fixture is too brittle in your environment, rerun with `--tree-mode signature`. Use `smoke_pytorch_adv_accuracy.py` when the user needs an integrated CPU PyTorch wrapper + bounded attack + adversarial-accuracy sanity check.
4. Keep attack generation in `../evasion-and-preprocessing/SKILL.md` and estimator construction in `../estimators-and-models/SKILL.md`.

## Route away from this sub-skill

- Attack crafting, perturbation budgets, preprocessing defences, and adversarial training -> `../evasion-and-preprocessing/SKILL.md`
- Estimator wrappers, `clip_values`, label shape fixes, and gradient-enabled model setup -> `../estimators-and-models/SKILL.md`
- Poisoning, backdoors, extraction, and attack setup for privacy/inference workflows -> `../poisoning-inference-extraction/SKILL.md`
- Package install, optional dependency, and device/backend readiness -> `../setup-and-backends/SKILL.md`

## Guardrails

- Do not run original repo tests, examples, notebooks, or maintainer scripts.
- Keep smoke inputs synthetic, tiny, deterministic, and CPU-only.
- Use `SecurityCurve` as a post-attack evaluation; if a weak attack leaves accuracy high, strengthen the attack rather than assuming the model is robust.
- Tree verification only makes sense for normalized `[0, 1]` data and tree classifiers that expose `get_trees()`.

Files in this skill

  • SKILL.md3 KB
  • references/metrics-evaluations-certification.md6.5 KB
  • references/troubleshooting.md4.7 KB
  • scripts/smoke_metrics_tree.py12.3 KB
  • scripts/smoke_pytorch_adv_accuracy.py5.2 KB

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…