Skip to content
Back to skills

Datasets And Metrics

ASecurity

"Use AIF360 legacy dataset containers and fairness metric classes

  • 247 stars
  • 0 votes
  • 0 copies
  • 2 views
  • Added September 8, 2026
devopspythongoawsapi

Works with

  • api

Security analysis

A100/100

Pro scans all 6 files and shows the line behind each finding

Scanned September 8, 2026

npx -y skills add VectorSpaceLab/AREX-Skill --skill datasets-and-metrics --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Datasets And Metrics?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for Datasets And Metrics
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/vectorspacelab-datasets-and-metrics/badge)](https://www.skillsdirectory.com/skills/vectorspacelab-datasets-and-metrics)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
SKILL.md
---
name: datasets-and-metrics
description: "Use AIF360 legacy dataset containers and fairness metric classes
  for tabular protected-group analysis."
disable-model-invocation: true
metadata:
  disco-role: operating
license: Apache 2.0
---

# AIF360 datasets and metrics router

Use this sub-skill when the task is about AIF360's legacy dataset objects and metric classes: constructing `StructuredDataset`, `BinaryLabelDataset`, `StandardDataset`, or `RegressionDataset` instances; loading legacy raw-dataset wrappers; defining privileged/unprivileged protected groups; and computing dataset, classification, regression, sample-distortion, or MDSS classification metrics.

## Read first

- [Data formats](references/data-formats.md): dataset object conventions, in-memory pandas construction, built-in dataset wrapper caveats, and group/label encoding rules.
- [API reference](references/api-reference.md): constructor signatures, metric class selection, method groups, and optional OT metric caveats.
- [Workflows](references/workflows.md): runnable patterns for synthetic datasets, prediction metric reports, raw wrappers, sample distortion, and regression metrics.
- [Troubleshooting](references/troubleshooting.md): import/install warnings, optional dependencies, raw-data failures, group/schema errors, and workflow-specific metric failures.

## Fast routing

1. **In-memory legacy dataset**: read [data formats](references/data-formats.md#in-memory-binarylabeldataset-from-pandas) and build a numeric, NA-free pandas `DataFrame`; pass `label_names`, `protected_attribute_names`, and explicit `favorable_label`/`unfavorable_label` when labels are not `1.0`/`0.0`.
2. **Built-in wrappers**: read [built-in dataset wrappers](references/data-formats.md#built-in-legacy-dataset-wrappers) before calling `AdultDataset`, `GermanDataset`, `CompasDataset`, `BankDataset`, `MEPSDataset19/20/21`, or `LawSchoolGPADataset`; legacy wrappers are not safe smoke tests because public raw files may be absent or network-bound.
3. **Metric choice**: use `BinaryLabelDatasetMetric` for one true dataset, `ClassificationMetric` for true-vs-predicted `BinaryLabelDataset` pairs, `SampleDistortionMetric` for original-vs-distorted `StructuredDataset` pairs, `RegressionDatasetMetric` for ranked regression datasets, and `MDSSClassificationMetric` only when a classification metric object must score a known group.
4. **Group definitions**: create `privileged_groups` and `unprivileged_groups` as lists of dictionaries keyed by `dataset.protected_attribute_names`; the dictionary values must match the encoded numeric protected attributes.
5. **Smoke check**: run the bundled no-data script with `python scripts/metric_report_smoke.py --pretty` from this sub-skill directory, or pass its path to a Python interpreter from another working directory.

## Route away when appropriate

- If the task asks for bias mitigation `fit`, `transform`, `predict`, postprocessing thresholds, or algorithm selection after metric diagnosis, route to [mitigation-algorithms](../mitigation-algorithms/SKILL.md).
- If the task asks for the preferred pandas/scikit-learn interface, `aif360.sklearn.datasets.fetch_*`, protected attributes in pandas indexes, sklearn scorers, or sklearn pipelines, route to [sklearn-interface](../sklearn-interface/SKILL.md).
- If the task asks for subgroup search, FACTS, bias scanning beyond `MDSSClassificationMetric.score_groups`, or metric text/JSON explainers, route to [detectors-and-explainers](../detectors-and-explainers/SKILL.md).

## Minimal operating checklist

- Confirm `aif360` imports and note that base dataset/metric workflows are CPU-only.
- Keep raw benchmark data and network access out of smoke tests; use synthetic `BinaryLabelDataset` examples first.
- Check `dataset.protected_attribute_names`, `dataset.privileged_protected_attributes`, and `dataset.unprivileged_protected_attributes` before constructing metric group dictionaries.
- For `ClassificationMetric`, make the predicted dataset by deep-copying the true dataset and changing only `labels` and optionally `scores`; otherwise equality validation will fail.
- Mark optional-dependency metrics such as optimal transport as optional/unverified unless the required extra has been installed and tested in the current task environment.

Files in this skill

  • SKILL.md4.2 KB
  • references/api-reference.md10.6 KB
  • references/data-formats.md8.3 KB
  • references/troubleshooting.md9.6 KB
  • references/workflows.md8.6 KB
  • scripts/metric_report_smoke.py6 KB

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…