Skip to content
Back to skills

Benchmark Enterprise Rag Agents With Enterpriserag Bench

ASecurity

Use EnterpriseRAG-Bench to evaluate an enterprise RAG or knowledge-agent system against a realistic synthetic company corpus with answer, recall, and comparative scoring.

  • 41 stars
  • 0 votes
  • 0 copies
  • 1 view
  • Added September 22, 2026
ai-agentspythongogitapisecuritydocumentation

Works with

  • api

Security analysis

A92/100
  • mediumInstalls packages at runtime which could introduce malicious dependencies

Pro shows the line behind each finding and how to fix it

Scanned September 22, 2026

npx -y skills add agentskillexchange/skills --skill benchmark-enterprise-rag-agents-with-enterpriserag-bench --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Benchmark Enterprise Rag Agents With Enterpriserag Bench?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for Benchmark Enterprise Rag Agents With Enterpriserag Bench
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/agentskillexchange-benchmark-enterprise-rag-agents-with-enterpriserag/badge)](https://www.skillsdirectory.com/skills/agentskillexchange-benchmark-enterprise-rag-agents-with-enterpriserag)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
SKILL.md
---
name: "Benchmark enterprise RAG agents with EnterpriseRAG-Bench"
slug: "benchmark-enterprise-rag-agents-with-enterpriserag-bench"
description: "Use EnterpriseRAG-Bench to evaluate an enterprise RAG or knowledge-agent system against a realistic synthetic company corpus with answer, recall, and comparative scoring."
github_stars: 562
verification: "security_reviewed"
source: "https://github.com/onyx-dot-app/EnterpriseRAG-Bench"
author: "Onyx"
publisher_type: "organization"
category: "Security & Verification"
framework: "Multi-Framework"
tool_ecosystem:
  github_repo: "onyx-dot-app/EnterpriseRAG-Bench"
  github_stars: 562
---

# Benchmark enterprise RAG agents with EnterpriseRAG-Bench

Use EnterpriseRAG-Bench to evaluate an enterprise RAG or knowledge-agent system against a realistic synthetic company corpus with answer, recall, and comparative scoring.

## Prerequisites

Python 3.10+, EnterpriseRAG-Bench dataset and questions, a RAG or knowledge-agent system under test, OpenAI or Anthropic compatible LLM credentials for evaluation

## Installation

Install or set up from the source-backed instructions:

Clone https://github.com/onyx-dot-app/EnterpriseRAG-Bench, install Python dependencies with pip install -r requirements.txt, set LLM_PROVIDER and LLM_API_KEY, download the benchmark dataset from the latest GitHub release or Hugging Face, write system outputs to answer_evaluation/answers.jsonl, then run python -m src.scripts.answer_evaluation.metrics_based_eval --answers-file answer_evaluation/answers.jsonl.

- Source: https://github.com/onyx-dot-app/EnterpriseRAG-Bench

## Documentation

- https://github.com/onyx-dot-app/EnterpriseRAG-Bench/blob/main/quickstart.md

## Source

- [Agent Skill Exchange](https://agentskillexchange.com/skills/benchmark-enterprise-rag-agents-with-enterpriserag-bench/)

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…