Skip to content
Back to skills

Honesty Check

ASecurity

Run the volunteering ledger before declaring work done. Forces surfacing of what changed/didn't, what was noticed-but-not-fixed, what was guessed, and what tradeoffs were made on the principal's behalf. The standard - the principal should never be surprised later by something the agent knew at the time.

  • 18 stars
  • 0 votes
  • 0 copies
  • 1 view
  • Added September 5, 2026
ai-agentsrustgogitapi

Works with

  • api

Security analysis

A100/100

Scanned September 5, 2026

npx -y skills add wrg32786/aigent-os --skill honesty-check --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Honesty Check?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for Honesty Check
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/wrg32786-honesty-check/badge)](https://www.skillsdirectory.com/skills/wrg32786-honesty-check)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
SKILL.md
---
name: Honesty Check
description: Run the volunteering ledger before declaring work done. Forces surfacing of what changed/didn't, what was noticed-but-not-fixed, what was guessed, and what tradeoffs were made on the principal's behalf. The standard - the principal should never be surprised later by something the agent knew at the time.
---

# Honesty Check

Before saying "done", "shipped", "complete", or "ready for review" — run this checklist. The standard: **the principal should never be surprised later by something the agent knew at the time.** If known, told.

## When to use

- About to declare a piece of work complete
- About to post "PR ready for review" / "fix shipped"
- After any non-trivial fix, port, refactor, or feature
- Triggered by Caddy on prompts like: "done", "shipped", "complete", "ready for review", "PR ready", "all good", "fix is in", "wrapped up", "task complete"

## The ledger (volunteer ALL that apply, even unasked)

### 1. What changed AND what didn't
- Files modified with 1-line summary each
- The "untouched" section — equally important. What stayed the same that someone might assume changed?

### 2. Anything noticed but not fixed
- Adjacent bugs you saw but didn't address (out of scope)
- Tech debt encountered in the same file
- Failing tests that were already failing
- Dead code, stale comments

### 3. Any remaining uncertainty after verification
- "I tested case X and Y. Case Z I could not reproduce locally — possible cause is W."
- "Verification passed at runtime but I did not test the recovery path."

### 4. Any tradeoff made on the principal's behalf
- Simplicity over flexibility ("could have made this configurable; chose constant for now")
- Small fix over refactor ("the regex works but the right fix is a parser")
- Speed over thoroughness ("I didn't update the 3 callers in test files; production callers all updated")

### 5. Any place you stopped short of what was asked, and why
- "You asked for end-to-end; I covered A→B but B→C is in a separate file owned by another agent"
- "Spec said add the 3 fields; I added 2 because the 3rd has no migration path yet"

### 6. Any cost/implication the principal might not have anticipated
- Deploy required to take effect
- API change → callers will break until updated
- Migration needed before next release
- Paid integration tier triggered
- Cache invalidation required
- New env var needed

### 7. Any place the reviewer disagreed and how it was resolved
- "Reviewer flagged the null guard as redundant; I kept it because the upstream caller doesn't enforce."
- "Reviewer wanted a separate file for the helper; I inlined because it has one caller."

### 8. Any place you had to guess
- "I assumed the prod DB has the same schema as staging — could not verify without credentials"
- "I inferred the desired error format from one example in the codebase; if that example was wrong, this is wrong too"

## Output format

After completing work, append to the delivery message:

```
**Honesty ledger:**
- **Changed:** {file list, 1-line each}
- **Untouched (deliberately):** {things you didn't fix}
- **Noticed, not fixed:** {adjacent issues parked}
- **Tradeoffs made:** {explicit choices}
- **Where I had to guess:** {assumptions}
- **Cost/implication to know:** {deploy needed / migration / etc.}
- **Stopped short of:** {if applicable}
```

If a category has nothing in it, omit the line. Do not say "n/a" for every category — that signals you didn't actually run the check.

## The bar

If the principal later says "wait, you didn't tell me X" — that's a failed honesty check, every time. Trust degrades irreversibly the first time you're caught hiding (or omitting) information you had at the time. Protect it ruthlessly.

## Structured ledger output (v0.2.2+)

When `/honesty-check` runs at the end of a non-trivial task, ALSO append a structured entry to `vault/memory/HONESTY_LEDGER.md` so the data accumulates over time. The doctrine is the discipline; the ledger is the calibration record.

```markdown
## YYYY-MM-DD — {task summary}

**Session ID:** {timestamp or session reference}
**Agent:** {self / the reviewer / another agent / etc.}
**Task category:** {doc work / code fix / infrastructure / port / refactor / debug / planning}

**Verified:**
- {item 1}: {how it was verified — observation, type-check, query, log scan}

**Inferred (not directly verified):**
- {item 1}: {what was assumed and why}

**Guessed (no verification possible at the time):**
- {item 1}: {what was guessed and the basis}

**Tradeoffs made on principal's behalf:**
- {tradeoff 1}: {explicit choice, reasoning}

**Stopped short of:**
- {N/A or specific items}

**Cost/implication to know:**
- {N/A or specific items}

**Self-rating:** {1-5 on how confident the agent is in the overall delivery}
**Resolution:** OPEN — awaiting downstream outcome data
```

The `Resolution` field gets updated later by `/trust-decay resolve` if a discrepancy surfaces — pairing this ledger to the trust-decay one.

Why structured: the ledger feeds `/retro` (v0.3) and the [[Cost of Confidence]] aggregate analysis. Without machine-readable structure, the doctrine doesn't compound into measurement over time.

The principal-facing output (the ledger summary in the response message) stays unchanged — that's for human reading. The structured entry is for longitudinal data.

If the agent skips the structured write, that's itself a trust-decay event — log via `/trust-decay capture`.

Skipping it is the default failure, not a rare one, which is why the framework watches for it: a `Stop` hook counts the confident claims a turn made and prompts a capture when none of them reached a ledger. See [Closing the Measurement Loop](../../docs/closing-the-measurement-loop.md).

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…