Skip to content
Back to skills

Arcdeck Eval

ASecurity

Evaluates the ability of LLM/VLM systems to generate high-quality, narrative-coherent presentation slides from academic papers. It probes content coverage, rhetorical structure preservation, textual fluency, and visual layout quality compared to human-authored references. Use when the user wants to benchmark on ArcBench, or asks about evaluating this task. Reports VLM-based Q/A Quiz Accuracy.

  • 3 stars
  • 0 votes
  • 0 copies
  • 0 views
  • Added September 11, 2026
researchpythongo

Security analysis

A100/100

Scanned September 11, 2026

npx -y skills add qhjqhj00/research-skills-pool --skill arcdeck-eval --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Arcdeck Eval?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for Arcdeck Eval
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/qhjqhj00-arcdeck-eval/badge)](https://www.skillsdirectory.com/skills/qhjqhj00-arcdeck-eval)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
SKILL.md
---
name: arcdeck-eval
description: Evaluates the ability of LLM/VLM systems to generate high-quality, narrative-coherent presentation slides from academic papers. It probes content coverage, rhetorical structure preservation, textual fluency, and visual layout quality compared to human-authored references. Use when the user wants to benchmark on ArcBench, or asks about evaluating this task. Reports VLM-based Q/A Quiz Accuracy.
metadata:
  skill_kind: dataset_eval
  source_arxiv: 2604.11969
  bibtex_key: ozden2026arcdeck
  confidence: high
---

# arcdeck-eval

> Narrative-Driven Paper-to-Slide Generation via ArcDeck — Ozden et al. (2026) (arXiv:2604.11969, 2026)

## What this evaluates

Evaluates the ability of LLM/VLM systems to generate high-quality, narrative-coherent presentation slides from academic papers. It probes content coverage, rhetorical structure preservation, textual fluency, and visual layout quality compared to human-authored references.

## Datasets

- **ArcBench** — total 100; splits: test (100)

## Metrics

- `VLM-based Q/A Quiz Accuracy` **(primary)** — range: percent
  - Percentage of correctly answered multiple-choice questions generated from the source paper and answered using only the generated slides. Calculated as (correct / 25) * 100 across four categories: Story, Visuals, Hard, and Depth.
- `VLM-as-Judge Score` — range: [1, 100]
  - Score from 1 to 100 per dimension (Text Quality, Narrative Flow, Visual Layout, Visual Thematic), derived from a 10-item checklist where each satisfied criterion contributes to the final score.
- `ROUGE-L` — range: [0, 1]
  - Longest common subsequence overlap between generated slide text and the source paper, measuring sequence-level content coverage.
- `Perplexity` — range: other
  - Linguistic fluency score computed using LLaMA-3-8B, representing the model's uncertainty over the generated slide text.
- `VLM Pairwise Preference Win Rate` — range: percent
  - Percentage of times a generated deck is preferred over a baseline or author-prepared deck by a VLM judge, averaged over 11 randomized runs.

## Input / output format

**Input**: Source paper (text & figures) and optionally generated slide decks or multiple-choice questions. For pairwise tests, two slide decks plus the source paper.

**Output**: Generated slides in 13.33×7.5-inch layout (HTML/CSS converted to PDF), or multiple-choice answers, or judge scores (1–100), or pairwise preference selections.

## Scoring recipe

```python
# VLM Quiz Accuracy
correct = sum(1 for q, a in zip(questions, answers) if a == q.gold)
accuracy = (correct / 25) * 100

# VLM-as-Judge Score
judge_score = sum(1 for criterion in checklist if judge.satisfies(criterion)) * 10

# ROUGE-L
rouge_l = rouge_l_score(generated_text, source_text)

# Pairwise Win Rate
wins = sum(1 for run in range(11) if judge.prefers(deck_A, deck_B, source))
win_rate = (wins / 11) * 100
```

## Common pitfalls

- Single-model judge bias is mitigated by using both closed (GPT-5) and open-source (Qwen3-VL) judges.
- Slide stylistic variations are controlled by evaluating all methods under fixed themes to ensure preference reflects content/structure.
- Quiz questions are generated by one VLM from the source paper and answered by a separate VLM from the slides, introducing potential generation/answering model mismatch.
- Pairwise evaluations are repeated 11 times with randomized order to account for position bias and stochastic judge behavior.

## Evidence (verbatim from paper)

> We complement the VLM-based evaluations with two standard text metrics reported in Tab.5: ROUGE-L, measuring sequence overlap between the generated slide text and the source paper to quantify content coverage, and Perplexity (computed via LLaMA-3-8B), assessing linguistic fluency of the slide text.

## Citation

```bibtex
@misc{ozden2026arcdeck,
  title={Narrative-Driven Paper-to-Slide Generation via ArcDeck},
  author={Ozden et al. (2026)},
  year={2026},
  note={arXiv:2604.11969}
}
```

- arXiv: 2604.11969

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…