Skip to content
Back to skills

Adversarial Rc Eval

ASecurity

This evaluation probes a model's ability to answer reading comprehension questions under adversarial conditions, specifically testing generalization across datasets constructed by progressively stronger language models. It measures how well models trained on challenging, model-in-the-loop generated questions can handle both adversarial and standard benchmarks. Use when the user wants to benchmark on SQuAD, BiDAF-adversarial, BERT-adversarial, RoBERTa-adversarial, DROP, Natural Questions, or a...

  • 3 stars
  • 0 votes
  • 0 copies
  • 0 views
  • Added September 11, 2026
researchpythongotestingperformance

Security analysis

A100/100

Scanned September 11, 2026

npx -y skills add qhjqhj00/research-skills-pool --skill adversarial-rc-eval --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Adversarial Rc Eval?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for Adversarial Rc Eval
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/qhjqhj00-adversarial-rc-eval/badge)](https://www.skillsdirectory.com/skills/qhjqhj00-adversarial-rc-eval)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
SKILL.md
---
name: adversarial-rc-eval
description: This evaluation probes a model's ability to answer reading comprehension questions under adversarial conditions, specifically testing generalization across datasets constructed by progressively stronger language models. It measures how well models trained on challenging, model-in-the-loop generated questions can handle both adversarial and standard benchmarks. Use when the user wants to benchmark on SQuAD, BiDAF-adversarial, BERT-adversarial, RoBERTa-adversarial, DROP, Natural Questions, or asks about evaluating this task. Reports F1.
metadata:
  skill_kind: dataset_eval
  source_arxiv: 2002.00293
  bibtex_key: bartolo2020beattheai
  confidence: high
---

# adversarial-rc-eval

> Beat the AI: Investigating Adversarial Human Annotation for Reading Comprehension — Bartolo et al. (2020) (arXiv:2002.00293, 2020)

## What this evaluates

This evaluation probes a model's ability to answer reading comprehension questions under adversarial conditions, specifically testing generalization across datasets constructed by progressively stronger language models. It measures how well models trained on challenging, model-in-the-loop generated questions can handle both adversarial and standard benchmarks.

## Datasets

- **SQuAD** — total ?; splits: train (-1), test (-1)
- **BiDAF-adversarial** — total ?; splits: train (-1), test (-1)
- **BERT-adversarial** — total ?; splits: train (-1), test (-1)
- **RoBERTa-adversarial** — total ?; splits: train (-1), test (-1)
- **DROP** — total ?; splits: test (-1)
- **Natural Questions** — total ?; splits: test (-1)

## Metrics

- `F1` **(primary)** — range: [0, 100] (percent)
  - Token-level F1 score between the predicted answer string and the gold answer string, calculated as the harmonic mean of precision and recall over word tokens.
- `EM` — range: [0, 100] (percent)
  - Exact match accuracy; returns 1 if the predicted answer string exactly matches the gold answer string, else 0.

## Input / output format

**Input**: A context passage (paragraph) and a natural language question requiring an answer extracted from or inferred from the passage.

**Output**: A text span or string representing the predicted answer.

## Scoring recipe

```python
def score(predictions, golds):
    em = sum(1 for p, g in zip(predictions, golds) if p == g) / len(predictions)
    f1s = []
    for p, g in zip(predictions, golds):
        p_tok, g_tok = set(p.lower().split()), set(g.lower().split())
        if not p_tok or not g_tok: f1s.append(0.0)
        else:
            common = p_tok & g_tok
            prec, rec = len(common)/len(p_tok), len(common)/len(g_tok)
            f1s.append(2*prec*rec/(prec+rec))
    return sum(f1s)/len(f1s), em
```

## Common pitfalls

- Random initialization and mini-batch ordering during training significantly impact adversarial annotation consistency; retrained models often achieve non-zero EM on their own adversarial datasets (Table 5).
- Evaluations are averaged over 10 runs with different random seeds, and results report mean ± standard deviation, not single-run scores.
- Adversarial datasets are model-specific; training on data generated by a weaker model does not guarantee performance on datasets generated by stronger models.

## Evidence (verbatim from paper)

> First, we observe – as expected given our annotation constraints – that model performance is 0.0EM on datasets created with the same respective model in the annotation loop. We observe however that retrained models do not reliably perform as poorly on those samples. For example, BERT reaches 19.7EM, whereas the original model used during annotation provides no correct answer with 0.0EM.

## Citation

```bibtex
@misc{bartolo2020beattheai,
  title={Beat the AI: Investigating Adversarial Human Annotation for Reading Comprehension},
  author={Bartolo et al. (2020)},
  year={2020},
  note={arXiv:2002.00293}
}
```

- arXiv: 2002.00293

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…