Skip to content
Back to skills

Cocometer

ASecurity

Compare CocoMeter accuracy and cost results with correctness-first ordering. Usage: $meter compare <before.json> <after.json>.

  • 724 stars
  • 0 votes
  • 0 copies
  • 1 view
  • Added September 5, 2026
ai-agents

Security analysis

A100/100

Pro scans all 20 files and shows the line behind each finding

Scanned September 5, 2026

npx -y skills add Snowflake-Labs/cocoplus --skill cocometer --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Cocometer?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for Cocometer
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/snowflake-labs-cocometer-2abb4772/badge)](https://www.skillsdirectory.com/skills/snowflake-labs-cocometer-2abb4772)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
SKILL.md
---
name: "meter-compare"
description: "Compare CocoMeter accuracy and cost results with correctness-first ordering. Usage: $meter compare <before.json> <after.json>."
version: "1.0.0"
author: "CocoPlus"
tags:
  - cocoplus
  - cocometer
  - comparison
---

Use this skill for `$meter compare`.

## Objective

Compare two CocoMeter or harness benchmark result files while respecting the configured optimization objective. The default is correctness-first: accuracy delta is the primary signal and token/cost delta is secondary.

CocoMeter session summaries may include three named cost rows:
- `execution_cost`: functional stage work.
- `coordination_cost`: advisor calls, handoff context assembly, artifact transfer, and status sync.
- `landing_cost`: reserve-window spend for final evaluation, console update, and handoff.

When `[meter] track_acrr = true`, session summaries may also include:
- `acrr_this_session`: average Agent Cognitive Redundancy Ratio across completed CocoFlow runs in the session.
- `acrr_runs[]`: the last run-level records with `complexity_tier_estimated`, `model_tier_configured`, `model_tier_actual`, `model_tier_used`, `model_drift`, `escalations_taken`, and `acrr`.

When `[meter] meter_reconciliation_enabled = true`, finalized session summaries may include `reconciliation_status`, `metering_gap_fraction`, `duplicates_found`, `model_tier_configured`, `model_tier_actual`, and `model_drift`. Treat transcript-derived totals as authoritative for billing comparisons when `reconciliation_status` is `gap_corrected`.

Treat ACRR as a calibration diagnostic, not an accuracy score. ACRR near 1.0 indicates that the initial complexity tier was sufficient. Consistently high ACRR indicates the task class is being underestimated before dispatch.

## Workflow

1. Parse `$meter compare <before.json> <after.json>`.
2. Read both files. Accept fields named `accuracy`, `score`, `task_accuracy`, `cost`, `tokens`, or `credits`.
3. Read `cocoplus.toml [meter]`:
   - `optimization_objective`, default `correctness-first`
   - `accuracy_equivalence_band`, default `0.02`
   - `cost_first_acknowledged`, default `false`
   - `track_acrr`, default `true`
4. Compute accuracy delta before cost delta.
5. If objective is `cost-first` but `cost_first_acknowledged` is not true, treat the run as correctness-first and warn.
6. If both files include ACRR, report ACRR trend after accuracy and before detailed cost rows. A lower ACRR with unchanged accuracy means the harness is better calibrated; a higher ACRR with unchanged accuracy means the run succeeded but over-expanded.
7. If either file includes reconciliation or model-drift metadata, report it after ACRR and before cost rows.

## Output Order

Always lead with accuracy unless cost-first is both configured and acknowledged:

```text
Accuracy: +3.2%
Cost: +12.0% (accepted because accuracy improved)
ACRR: 1.0 → 1.2 (watch for rising complexity miscalibration)
Verdict: improvement under correctness-first objective
```

If cost falls but accuracy regresses:

```text
Accuracy: -1.5%
Cost: -18.0%
Verdict: regression under correctness-first objective
WARNING: cost reduction accompanied by accuracy regression.
```

## Exit Criteria

- [ ] Accuracy delta appears before cost delta by default.
- [ ] Cost reduction is never labeled as improvement when accuracy regresses.
- [ ] Equivalent accuracy uses the configured band.
- [ ] Missing or malformed input files fail with a clear diagnostic.
- [ ] Coordination and landing costs remain separate from execution cost when present.
- [ ] ACRR appears as a calibration diagnostic when present and is not treated as a replacement for accuracy.
- [ ] Transcript reconciliation and model drift appear when present, and corrected transcript totals lead runtime estimates for billing comparison.

## Anti-Rationalization

| Shortcut / Temptation | Why It Fails |
|-----------------------|--------------|
| Treat the skill as complete because the file exists | Skill contracts must describe observable behavior and verification, not just command names. |
| Skip artifact and safety checks for a small command | Small commands still mutate state or guide execution; preserve the same gates. |

Files in this skill

  • archetype-classifier.skill.md1.5 KB
  • chargeback-refresh.skill.md1.5 KB
  • cost-center-resolver.skill.md1.5 KB
  • invoice-generator.skill.md1.5 KB
  • meter-accuracy.skill.md2.3 KB
  • meter-benchmark.skill.md1.3 KB
  • meter-chargeback.skill.md3.2 KB
  • meter-compare.skill.md4.1 KB
  • meter-estimate.skill.md3.4 KB
  • meter-history.skill.md1.9 KB
  • meter-invoice.skill.md1.5 KB
  • meter-off.skill.md2.3 KB
  • meter-on.skill.md2.3 KB
  • meter-reconcile.skill.md4.1 KB
  • meter-refresh.skill.md1.5 KB
  • meter-report.skill.md3.5 KB
  • meter-status.skill.md1.5 KB
  • meter-sync.skill.md2 KB
  • meter-verify.skill.md1.8 KB

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…