Skip to content
Back to skills

Asnm Tun Eval

ASecurity

Evaluates classifier robustness against tunneling and non-payload-based adversarial obfuscations in network traffic. It tests whether models trained on direct attacks can detect obfuscated variants and how training data augmentation with obfuscated samples improves detection. Use when the user wants to benchmark on ASNM-TUN, or asks about evaluating this task. Reports F1-measure.

  • 3 stars
  • 0 votes
  • 0 copies
  • 3 views
  • Added September 11, 2026
researchpythontestinggitperformance

Security analysis

A100/100

Scanned September 11, 2026

npx -y skills add qhjqhj00/research-skills-pool --skill asnm-tun-eval --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Asnm Tun Eval?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for Asnm Tun Eval
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/qhjqhj00-asnm-tun-eval/badge)](https://www.skillsdirectory.com/skills/qhjqhj00-asnm-tun-eval)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
SKILL.md
---
name: asnm-tun-eval
description: Evaluates classifier robustness against tunneling and non-payload-based adversarial obfuscations in network traffic. It tests whether models trained on direct attacks can detect obfuscated variants and how training data augmentation with obfuscated samples improves detection. Use when the user wants to benchmark on ASNM-TUN, or asks about evaluating this task. Reports F1-measure.
metadata:
  skill_kind: dataset_eval
  source_arxiv: 1910.10528
  bibtex_key: homoliak2019asnm
  confidence: high
---

# asnm-tun-eval

> ASNM Datasets: A Collection of Network Traffic Features for Testing of Adversarial Classifiers and Network Intrusion Detectors — Homoliak et al. (2019) (arXiv:1910.10528, 2019)

## What this evaluates

Evaluates classifier robustness against tunneling and non-payload-based adversarial obfuscations in network traffic. It tests whether models trained on direct attacks can detect obfuscated variants and how training data augmentation with obfuscated samples improves detection.

## Datasets

- **ASNM-TUN** — total 394; splits: train (-1), test (-1)

## Metrics

- `F1-measure` **(primary)** — range: percent
  - Harmonic mean of precision and recall: 2 * (precision * recall) / (precision + recall). Used as the headline metric for classifier performance.
- `Recall` — range: percent
  - True positive rate: TP / (TP + FN). Reported as average recall across classes.
- `Accuracy` — range: percent
  - Ratio of correctly classified instances to total instances: (TP + TN) / (TP + FP + TN + FN).

## Input / output format

**Input**: Aggregated bidirectional TCP flow features (ASNM features), including metadata such as packet counts, byte counts, inter-arrival times, and protocol flags. Specific features are selected via Forward Feature Selection (FFS).

**Output**: Binary class label: 'Legitimate' or 'Attack' (or 'Obfuscated Attack' / 'All Attacks' depending on the experimental setup).

## Scoring recipe

```python
def compute_metrics(y_true, y_pred):
    tp = sum(t == 1 and p == 1 for t, p in zip(y_true, y_pred))
    fp = sum(t == 0 and p == 1 for t, p in zip(y_true, y_pred))
    fn = sum(t == 1 and p == 0 for t, p in zip(y_true, y_pred))
    tn = sum(t == 0 and p == 0 for t, p in zip(y_true, y_pred))
    prec = tp / (tp + fp) if (tp + fp) > 0 else 0
    rec = tp / (tp + fn) if (tp + fn) > 0 else 0
    f1 = 2 * prec * rec / (prec + rec) if (prec + rec) > 0 else 0
    acc = (tp + tn) / (tp + fp + fn + tn)
    return {'precision': prec, 'recall': rec, 'f1': f1, 'accuracy': acc}
```

## Common pitfalls

- Forward Feature Selection (FFS) is applied to the dataset before or during cross-validation. If applied to the entire dataset prior to splitting, it causes data leakage and inflates performance metrics.
- The datasets use aggregated flow-level metadata rather than raw packet payloads. Models trained on these features will not generalize to payload-based intrusion detection systems.
- Class imbalance is mitigated via stratified sampling in folds, but baseline accuracy is extremely high (>99%), which can mask poor detection rates for minority attack classes.

## Evidence (verbatim from paper)

> we used 5-fold cross-validation and forward feature selection (FFS) on top of the Naive Bayes classifier with kernel functions for the estimation of density distribution, which represents a non-parametric estimation method. In FFS, we accepted one iteration without improvement as we wanted to avoid the selection process to get stuck in local extremes. The maximal number of selected features was limited to 20 (although it was never reached). We used the binary label of the dataset (i.e., label_2), and we obtained $F_{1}$ -measure over $90\%$ and an average recall of both classes equal to $92\%$.

## Citation

```bibtex
@misc{homoliak2019asnm,
  title={ASNM Datasets: A Collection of Network Traffic Features for Testing of Adversarial Classifiers and Network Intrusion Detectors},
  author={Homoliak et al. (2019)},
  year={2019},
  note={arXiv:1910.10528}
}
```

- arXiv: 1910.10528

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…