Skip to content
Back to skills

Bartscore

ASecurity

BARTScore evaluates the quality of generated text by treating evaluation as a conditional text generation task. It measures the likelihood of a hypothesis given a source text, or a reference given a hypothesis, using pre-trained sequence-to-sequence models. This approach enables unsupervised, multi-perspective assessment of fluency, factuality, and informativeness without relying on human annotations or simple n-gram overlap. Use when the user has predictions and gold and needs to compute Spe...

  • 3 stars
  • 0 votes
  • 0 copies
  • 1 view
  • Added September 11, 2026
researchpythongoperformance

Security analysis

A100/100

Scanned September 11, 2026

npx -y skills add qhjqhj00/research-skills-pool --skill bartscore --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Bartscore?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for Bartscore
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/qhjqhj00-bartscore/badge)](https://www.skillsdirectory.com/skills/qhjqhj00-bartscore)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
SKILL.md
---
name: bartscore
description: BARTScore evaluates the quality of generated text by treating evaluation as a conditional text generation task. It measures the likelihood of a hypothesis given a source text, or a reference given a hypothesis, using pre-trained sequence-to-sequence models. This approach enables unsupervised, multi-perspective assessment of fluency, factuality, and informativeness without relying on human annotations or simple n-gram overlap. Use when the user has predictions and gold and needs to compute Spearman Correlation.
metadata:
  skill_kind: metric
  source_arxiv: 2106.11520
  bibtex_key: yuan2021bartscore
  confidence: high
---

# bartscore

> BARTScore: Evaluating Generated Text as Text Generation — Yuan et al. (2021) (arXiv:2106.11520, 2021)

## What this evaluates

BARTScore evaluates the quality of generated text by treating evaluation as a conditional text generation task. It measures the likelihood of a hypothesis given a source text, or a reference given a hypothesis, using pre-trained sequence-to-sequence models. This approach enables unsupervised, multi-perspective assessment of fluency, factuality, and informativeness without relying on human annotations or simple n-gram overlap.

## Datasets

- **WMT19** — total ?; splits: test (-1)
- **REALSumm** — total ?; splits: test (-1)
- **SummEval** — total ?; splits: test (-1)
- **NeR18** — total ?; splits: test (-1)
- **Rank19** — total 373; splits: test (373)
- **QAGS** — total 474; splits: test (474)
- **BAGEL** — total 202; splits: test (202)
- **SFHOT** — total 398; splits: test (398)
- **SFRES** — total 581; splits: test (581)

## Metrics

- `Spearman Correlation` **(primary)** — range: [-1, 1]
  - Measures the monotonic relationship between two ranked variables. Computed as the Pearson correlation between the rank of automated metric scores and the rank of human judgment scores.
- `Kendall's Tau` — range: [-1, 1]
  - Measures the ordinal association between two measured quantities by counting concordant and discordant pairs.
- `Pearson Correlation` — range: [-1, 1]
  - Measures the linear correlation between two sets of data.
- `Accuracy` — range: [0, 1]
  - Measures the percentage of correct ranking between factual texts and non-factual texts.

## Input / output format

**Input**: Source text (s), hypothesis/generation (h), and optionally reference text (r). For prompting variants, seed phrases are prepended to the decoder input.

**Output**: A continuous scalar score representing the log-likelihood or averaged generation probability of the hypothesis/reference pair.

## Scoring recipe

```python
def compute_bartscore(source, hypothesis, reference, model, prompts):
    scores = []
    for prompt in prompts:
        # Choose direction based on task perspective: s->h or h->r
        input_text = f"{prompt} {source}"
        target_text = hypothesis
        log_prob = model.log_prob(target_text, input_text)
        scores.append(log_prob)
    return np.mean(scores)
```

## Common pitfalls

- Using the same prompt strategy across all tasks ignores task-specific optimal prompts (e.g., 'Such as' for MT vs. ensembling for SUM/D2T).
- Fine-tuning on paraphrase data (ParaBank2) improves general quality metrics but degrades factuality performance, as summaries and documents are not strict paraphrases.
- Selecting the wrong BARTScore usage mode (s→h vs h→r) for a given evaluation perspective leads to suboptimal correlation with human judgments.

## Evidence (verbatim from paper)

> Spearman Correlation [73] assesses the monotonic relationships between two variables. Kendall's Tau [27] measures the ordinal association between two measured quantities. Accuracy, in our experiments, measures the percentage of correct ranking between factual texts and non-factual texts. We follow previous works in the choices of measures for different datasets to make a fair comparison.

## Citation

```bibtex
@misc{yuan2021bartscore,
  title={BARTScore: Evaluating Generated Text as Text Generation},
  author={Yuan et al. (2021)},
  year={2021},
  note={arXiv:2106.11520}
}
```

- arXiv: 2106.11520

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…