Skip to content
Back to skills

Evaluate Agent And Model Workflows With Evalscope

ASecurity

Run repeatable EvalScope benchmark suites for LLM, VLM, RAG, and agent workflows, then inspect traces and reports before changing models or prompts.

  • 36 stars
  • 0 votes
  • 0 copies
  • 1 view
  • Added September 2, 2026
ai-agentspythongobashreactdockergitapibackendsecuritydocumentation

Works with

  • api
  • mcp

Security analysis

A92/100
  • mediumInstalls packages at runtime which could introduce malicious dependencies

Pro shows the line behind each finding and how to fix it

Scanned September 2, 2026

npx -y skills add agentskillexchange/skills --skill evaluate-agent-and-model-workflows-with-evalscope --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Evaluate Agent And Model Workflows With Evalscope?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for Evaluate Agent And Model Workflows With Evalscope
[![Security: A β€” Skills Directory](https://www.skillsdirectory.com/api/skills/agentskillexchange-evaluate-agent-and-model-workflows-with-evalscope/badge)](https://www.skillsdirectory.com/skills/agentskillexchange-evaluate-agent-and-model-workflows-with-evalscope)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
SKILL.md
---
name: "Evaluate agent and model workflows with EvalScope"
slug: "evaluate-agent-and-model-workflows-with-evalscope"
description: "Run repeatable EvalScope benchmark suites for LLM, VLM, RAG, and agent workflows, then inspect traces and reports before changing models or prompts."
github_stars: 2953
verification: "security_reviewed"
source: "https://github.com/modelscope/evalscope"
author: "ModelScope Community"
publisher_type: "organization"
category: "Monitoring & Alerts"
framework: "Multi-Framework"
tool_ecosystem:
  github_repo: "modelscope/evalscope"
  github_stars: 2953
---

# Evaluate agent and model workflows with EvalScope

Run repeatable EvalScope benchmark suites for LLM, VLM, RAG, and agent workflows, then inspect traces and reports before changing models or prompts.

## Prerequisites

Python 3.10+, pip, evalscope package, model endpoint or local model backend, API credentials when evaluating hosted models, selected benchmark datasets

## Installation

Use the upstream install or setup path that matches your environment:
- pip install evalscope
- pip install 'evalscope[service]'

Requirements and caveats from upstream:
- <img src="https://img.shields.io/badge/python-%E2%89%A53.10-5be.svg">
- **πŸ€– Agent Evaluation Mode**: Drives benchmarks (e.g. GSM8K, AIME, SWE-bench Agentic) inside a controlled multi-turn AgentLoop with pluggable strategies, tools and Docker sandbox; full per-sample Agent Trace is recorde...
- πŸ”₯ **[2026.05.26]** Added the [GAIA](https://evalscope.readthedocs.io/en/latest/third_party/gaia.html) agent benchmark (multi-turn ReAct + bash in a Docker sandbox, official rule-based scorer) and generic [MCP server](...

Basic usage or getting-started notes:
- πŸ”₯ **[2026.05.19]** Added support for [SWE-bench_Pro](https://evalscope.readthedocs.io/en/latest/third_party/swe_bench_pro.html) and [τ³-bench](https://evalscope.readthedocs.io/en/latest/third_party/tau3_bench.html): S...
- πŸ”₯ **[2026.05.15]** Introduced **Agent Evaluation Mode**: any benchmark based on DefaultDataAdapter (GSM8K, AIME, IFEval, etc.) can now be driven through a multi-turn AgentLoop with pluggable strategies (function_calli...
- πŸ”₯ **[2026.05.08]** Partnered with [LightSeek](https://lightseek.org/) to launch [TokenSpeed](https://lightseek.org/blog/lightseek-tokenspeed.html), a speed-of-light LLM inference engine for agentic workloads. EvalScop...

- Source: https://github.com/modelscope/evalscope
- Extracted from upstream docs: https://raw.githubusercontent.com/modelscope/evalscope/HEAD/README.md

## Documentation

- https://evalscope.readthedocs.io/en/latest/

## Source

- [Agent Skill Exchange](https://agentskillexchange.com/skills/evaluate-agent-and-model-workflows-with-evalscope/)

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…