Skip to content
Back to skills

Enterprise Agent Ops

ASecurity

Use when operate long-lived agent workloads with observability, security boundaries, and lifecycle management. Triggers on \"enterprise-agent-ops\", \"enterprise agent ops\".

  • 2 stars
  • 0 votes
  • 0 copies
  • 1 view
  • Added September 19, 2026
ai-agentsci/cdsecurity

Works with

  • cli

Security analysis

A100/100

Scanned September 19, 2026

npx -y skills add majinmagros/magros.ai-skills --skill enterprise-agent-ops --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Enterprise Agent Ops?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for Enterprise Agent Ops
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/majinmagros-enterprise-agent-ops/badge)](https://www.skillsdirectory.com/skills/majinmagros-enterprise-agent-ops)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
SKILL.md
---
name: enterprise-agent-ops
description: "Use when operate long-lived agent workloads with observability, security boundaries, and lifecycle management. Triggers on \"enterprise-agent-ops\", \"enterprise agent ops\"."
metadata:
  origin: ECC
---

# Enterprise Agent Ops

Use this skill for cloud-hosted or continuously running agent systems that need operational controls beyond single CLI sessions.

## When to Use

- "Run this agent 24/7 in production"
- "Agent fleet needs observability and kill switches"
- "Rollout/rollback plan for agent version"
- "Track cost per successful agent task"
- "Failure spike in the agent service"

## Example

```yaml
agent_service:
  timeout_s: 300
  max_retries: 2
  kill_switch: env.KILL_AGENT_V2
  audit: high-risk-actions-only
```

## Operational Domains

1. runtime lifecycle (start, pause, stop, restart)
2. observability (logs, metrics, traces)
3. safety controls (scopes, permissions, kill switches)
4. change management (rollout, rollback, audit)

## Baseline Controls

- immutable deployment artifacts
- least-privilege credentials
- environment-level secret injection
- hard timeout and retry budgets
- audit log for high-risk actions

## Metrics to Track

- success rate
- mean retries per task
- time to recovery
- cost per successful task
- failure class distribution

## Incident Pattern

When failure spikes:
1. freeze new rollout
2. capture representative traces
3. isolate failing route
4. patch with smallest safe change
5. run regression + security checks
6. resume gradually

## Deployment Integrations

This skill pairs with:
- PM2 workflows
- systemd services
- container orchestrators
- CI/CD gates

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…