Skip to content
Back to skills

Analyze Comparison Tests

ASecurity

Collects comparison test run artifacts and answers the user's questions based on the trajectories of each run. WHEN TO USE: collect comparison test artifacts

  • 253 stars
  • 0 votes
  • 0 copies
  • 1 view
  • Added September 3, 2026
ai-agentsbashgit

Security analysis

A100/100

Pro scans all 2 files and shows the line behind each finding

Scanned September 3, 2026

npx -y skills add microsoft/GitHub-Copilot-for-Azure --skill analyze-comparison-tests --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Analyze Comparison Tests?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for Analyze Comparison Tests
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/microsoft-analyze-comparison-tests/badge)](https://www.skillsdirectory.com/skills/microsoft-analyze-comparison-tests)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
SKILL.md
---
name: analyze-comparison-tests
description: "Collects comparison test run artifacts and answers the user's questions based on the trajectories of each run. WHEN TO USE: collect comparison test artifacts"
license: MIT
metadata:
  author: Microsoft
  version: "1.0.0"
---

# Steps

1. Collect run artifacts

Execute the collect-artifacts script to download the test run artifacts.

The user must provide an JSON file to correlate each comparison test run with the GitHub Actions run. The script expects one input argument as the path to this JSON file. The JSON input is supposed to be the JSON output when queuing the comparison test runs using the `npm run compare:run` command.

```bash
cd tests/
npm run compare:collect -- input.json
```

The collect-artifacts script will download the test run artifacts to a directory named `comparison-artifacts` in the current working directory. Before executing the script, check if there is already such an directory. If so, skip executing the script and proceed to step 2.

2. Extract insights

The downloaded artifacts will have the following folder structure:

```text
comparison-artifacts/
├── <branch-name>/
│   ├── <stimulus-name-1>/
│   │   ├── <model>-with-skill/
│   │   │   ├── agent-metadata-<date-string-1>.md
│   │   │   ├── agent-metadata-<date-string-2>.md
│   │   │   └── ...
│   │   └── <model>-without-skill/
│   │       ├── agent-metadata-<date-string-1>.md
│   │       ├── agent-metadata-<date-string-2>.md
│   │       └── ...
│   └── <stimulus-name-2>/
│       ├── <model>-with-skill/
│       │   └── agent-metadata-*.md
│       └── <model>-without-skill/
│           └── agent-metadata-*.md
└── <branch-name-2>/
    └── ...
```

Each `<branch-name>/<stimulus-name>/<model>-with-skill` or `<branch-name>/<stimulus-name>/<model>-without-skill` directory contains the test run trajectories for that stimulus and model on that branch, with or without skills. Each trajectory is a markdown file that records user prompts, tool call requests, tool execution results, assistant responses that happened during the run. It also contains statistics such as token usage and turns. Based on the trajectories, answer the user's questions for each test run. Generate a report following the [report-template](./references/report-template.md) to show your answers.

Files in this skill

  • SKILL.md2.4 KB
  • references/report-template.md217 B

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…