Skip to content
Back to skills

Ai Agent Observability Evaluation

ASecurity

Use when measuring, evaluating, replaying, evidencing, or tracking success for AI agent tasks, steps, traces, and outcomes.

  • 28 stars
  • 0 votes
  • 0 copies
  • 2 views
  • Added May 27, 2026
developmentgosecurity

Security analysis

A100/100

Pro scans all 15 files and shows the line behind each finding

Scanned September 20, 2026

npx -y skills add peterbamuhigire/skills-web-dev --skill ai-agent-observability-evaluation --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Ai Agent Observability Evaluation?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for Ai Agent Observability Evaluation
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/peterbamuhigire-ai-agent-observability-evaluation/badge)](https://www.skillsdirectory.com/skills/peterbamuhigire-ai-agent-observability-evaluation)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
SKILL.md
---
name: ai-agent-observability-evaluation
description: Use when measuring, evaluating, replaying, evidencing, or tracking success for AI agent tasks, steps, traces, and outcomes.
metadata:
  portable: true
  compatible_with:
  - claude-code
  - codex
---

# AI Agent Observability Evaluation

## Operating contract

## Inputs

| Input | Required | Purpose |
|---|---|---|
| Domain evidence | yes | task contract, trace schema, evaluation dataset, success thresholds, privacy rules, and sampling window |

## Outputs

- Produce: trace and outcome scorecard, failure taxonomy, replay evidence, regressions, and release recommendation.

## Capability and permission boundaries

Default to read-only analysis. Read only scoped records; redact secrets and regulated data. Writes, execution, network calls, production configuration, customer communication, billing changes, and delegation require explicit authority and an identified owner. Never widen tenant, time-window, or system scope implicitly.

## Degraded mode

When required telemetry, evidence, execution, network access, or write authority is unavailable, return a partial result with each unassessed item labelled, preserve the safest existing state, and state the evidence or approval needed to continue. Never convert missing evidence into a pass.

## Decision rules

| Condition | Action |
|---|---|
| Scope, owner, or threshold is missing | Stop the affected decision and request it |
| Evidence is incomplete but read-only analysis is safe | Produce a qualified partial result and gap list |
| A mutation exceeds authority or tenant boundary | Block it and route for approval |
| Evidence meets the stated threshold | Issue the output with provenance and owner |

## Anti-Patterns

- Treating absent evidence as success. Fix: mark the check unassessed and name the missing source.
- Expanding one tenant or workflow to all tenants. Fix: enforce supplied scope at every query and action.
- Performing a production write during analysis. Fix: emit a reviewed change plan until authority is explicit.
- Reporting a metric without population, window, or source. Fix: attach all three.
- Hiding a failed threshold inside an average. Fix: report failure slices and the remediation owner.

Acknowledgement: Shared by Peter Bamuhigire, techguypeter.com, +256 784 464178.

<!-- dual-compat-start -->

## Use When

- Design agent evals, task-success metrics, traces, replay, evidence capture, and regression gates.
- Instrument agent runs with events, step logs, artifacts, and customer-visible status.
- Turn agent execution evidence into QA, audit, incident, or product-quality signals.

## Do Not Use When

- The work is not AI-specific or agentic-AI-specific.
- A narrower retained AI parent skill fits the request better.

## Required Inputs

- Product, tenant, user, data, risk, and operational context relevant to the AI workflow.
- Target artifact: design, implementation plan, audit, test strategy, UX flow, commercial policy, or runbook.
- Constraints from security, privacy, reliability, billing, support, and compliance stakeholders when relevant.

## Workflow

1. Read this SKILL.md first.
2. Load [references/routing.md](references/routing.md) to select the absorbed child reference that matches the task.
3. Load only the selected child reference files needed for the current request.
4. Produce execution-oriented output with assumptions, risks, evidence, and next actions where relevant.

## Quality Standards

- Keep routing explicit: name which reference files were used when the work depends on absorbed material.
- Preserve tenant isolation, auditability, cost controls, safety gates, and operational evidence when they matter.
- Prefer concrete contracts, checklists, tables, schemas, runbooks, and decision records over broad summaries.

## Anti-Patterns

- Loading every absorbed reference by default.
- Treating AI-specific billing, compliance, safety, or UX concerns as generic SaaS work without checking AI failure modes.
- Hiding retired skill names; old slugs must remain discoverable through [references/routing.md](references/routing.md).

## Outputs

- A concrete deliverable matched to the request: architecture, implementation plan, audit, policy, runbook, UX flow, test strategy, or operating model.
- The selected consolidated reference files and any assumptions, risks, evidence requirements, or follow-up actions that affect execution.
## References

- [references/routing.md](references/routing.md) maps retired child skill slugs to their consolidated reference folders.

## Consolidated Child References

- Load [references/routing.md](references/routing.md) to map retired AI child skill slugs to their reference modules.
## Evidence Produced

| Category | Artifact | Format | Example |
| --- | --- | --- | --- |
| Correctness | Agent evaluation and replay report | Markdown plus run artefacts | task-success, unsafe-action, cost, latency, replay, and regression results |

<!-- dual-compat-end -->

Files in this skill

  • SKILL.md2.7 KB
  • references/ai-agent-eval/entrypoint.md14.4 KB
  • references/ai-agent-eval/references/golden-tasks-construction.md6.3 KB
  • references/ai-agent-eval/references/replay-based-eval.md6.6 KB
  • references/ai-agent-evidence-automation/entrypoint.md11.7 KB
  • references/ai-agent-evidence-automation/references/auditor-portal-design.md8.6 KB
  • references/ai-agent-evidence-automation/references/evidence-catalogue.md7.3 KB
  • references/ai-agent-evidence-automation/references/evidence-pack-format.md9.5 KB
  • references/ai-agent-observability-and-replay/entrypoint.md18.8 KB
  • references/ai-agent-observability-and-replay/references/trace-schema-agent.md6.4 KB
  • references/ai-agent-task-success-tracking/entrypoint.md11.9 KB
  • references/ai-agent-task-success-tracking/references/dispute-resolution.md5.7 KB
  • references/ai-agent-task-success-tracking/references/judge-cascade-pipeline.md8.9 KB
  • references/ai-agent-task-success-tracking/references/success-contract-spec.md6.2 KB
  • references/routing.md1 KB

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…