Skip to content
Back to skills

Cost Trend

ASecurity

Read every docs/benchmarks/runs/*.json and surface drift in win rate, latency, escalation rate, and LLM-baseline cost over time

  • 73,733 stars
  • 0 votes
  • 0 copies
  • 2 views
  • Added May 27, 2026
data-aibashnode

Security analysis

A100/100

Scanned May 27, 2026

npx -y skills add ruvnet/ruflo --skill cost-trend --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Cost Trend?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for Cost Trend
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/ruvnet-cost-trend/badge)](https://www.skillsdirectory.com/skills/ruvnet-cost-trend)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
SKILL.md
---
name: cost-trend
description: Read every docs/benchmarks/runs/*.json and surface drift in win rate, latency, escalation rate, and LLM-baseline cost over time
argument-hint: ""
allowed-tools: Bash
---

# Cost Trend

The smoke gate is binary (`winRate ≥ 0.80` → pass/fail). The corpus benchmarks captured over time form a curve — and curves catch regressions the gate misses (win rate slowly creeping from 100% to 85% is "still passing" by smoke but a real degradation).

This skill reads every persisted run in `docs/benchmarks/runs/*.json` and reports first→last deltas plus a per-run series, flagging regressions in win rate or latency.

## When to use

- Before a release — check that the speedup hasn't drifted.
- After expanding the corpus — verify older runs still hit the same win rate on the new corpus *they* reflected.
- After upgrading `agent-booster` — surface latency / strategy changes.

## Steps

1. **Run the trend script** from the project root:

   ```bash
   node plugins/ruflo-cost-tracker/scripts/trend.mjs
   ```

   Optional env:
   - `TREND_FORMAT=json` — emit JSON instead of markdown
   - `TREND_LIMIT=10` — consider only the most recent N runs

2. **Inspect the drift summary** — first vs last on win rate, avg latency, p99, escalation rate, speedup vs Gemini.

3. **Inspect the per-run series** — one row per run, including Sonnet 4.6 + Opus 4.7 baseline latencies if those were enabled (`BENCH_ANTHROPIC=1` at run time).

4. **Regression flags** — the script emits `> ⚠ Regression` callouts when:
   - Win rate dropped between first and last run
   - Avg latency rose ≥ 1.5× from first run

## Cross-references

- `cost-benchmark` — the producer of the run JSONs this skill consumes
- `bench/booster-corpus.json` — the corpus version is recorded in each run, so trends across corpus versions remain interpretable
- `docs/benchmarks/runs/latest.json` — the most-recent run; smoke step 23 gates on `winRate ≥ 0.80` from this file

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…