Skip to content
Back to skills

Best Practice

ASecurity

Applies Anthropic's documented best practices for building with Claude to the task at hand — model and effort selection, prompt construction, subagent delegation, tool design, evaluation, and guardrails. Use when the user invokes /best-practice, or asks for work done "by best practices", "the way Anthropic recommends", "as the official documentation recommends", or uses the Russian equivalents "по лучшим практикам", "сделай как рекомендует Anthropic", "как в официальной документации". Also us...

  • 2 stars
  • 0 votes
  • 0 copies
  • 0 views
  • Added September 19, 2026
ai-agentspythonrustgobashrailsapifrontendperformancedocumentation

Works with

  • api
  • mcp

Security analysis

A100/100

Pro scans all 9 files and shows the line behind each finding

Scanned September 19, 2026

npx -y skills add starsinc1708/best-practice-skill --skill best-practice --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Best Practice?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for Best Practice
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/starsinc1708-best-practice/badge)](https://www.skillsdirectory.com/skills/starsinc1708-best-practice)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
SKILL.md
---
name: best-practice
description: Applies Anthropic's documented best practices for building with Claude to the task at hand — model and effort selection, prompt construction, subagent delegation, tool design, evaluation, and guardrails. Use when the user invokes /best-practice, or asks for work done "by best practices", "the way Anthropic recommends", "as the official documentation recommends", or uses the Russian equivalents "по лучшим практикам", "сделай как рекомендует Anthropic", "как в официальной документации". Also use when the user asks how to prompt Claude, choose a model, set effort, write a subagent prompt, design evaluations, or harden an agent against prompt injection.
---

# Best Practice

Do the requested work the way Anthropic's documentation says to do it — and make the reasoning visible, so the user learns the rule rather than just receiving the output.

This skill has two modes. Detect which one applies before doing anything else.

**Execution mode** — the user wants a task done, and wants it done properly.
Trigger: `/best-practice <task>`, or a task request carrying a quality qualifier ("по лучшим практикам", "the way Anthropic recommends", "properly", "как в документации").
Behavior: run the workflow below, then produce the deliverable.

**Advisory mode** — the user is asking what the best practice *is*.
Trigger: "how should I prompt for X", "which model for Y", "how do I write a subagent prompt", "is this prompt good".
Behavior: skip to the relevant reference file, answer with the rule and its source, skip the deliverable.

---

## Workflow (execution mode)

Copy this checklist into your working notes and check items off as you go.

```
Best Practice Progress:
- [ ] Step 1: Classify the task
- [ ] Step 2: Define success criteria
- [ ] Step 3: Choose model and effort
- [ ] Step 4: Construct the prompt
- [ ] Step 5: Decide on delegation
- [ ] Step 6: Execute
- [ ] Step 7: Verify against the criteria from Step 2
- [ ] Step 8: Report what was applied
```

### Step 1: Classify the task

Put the request in exactly one bucket. The bucket determines which reference files you read.

| Bucket | Signals | Read |
|---|---|---|
| Generation | write, draft, build, create, design | `reference/prompting.md` |
| Classification | route, categorize, label, moderate, triage | `reference/evaluation.md`, `reference/prompting.md` |
| Extraction / summarization | summarize, extract, pull out, condense | `reference/prompting.md` |
| Agentic / multi-step | automate, orchestrate, agent, pipeline, long-running | `reference/agent-architecture.md`, `reference/tool-use.md` |
| Tool or integration design | tool, API, MCP, connector, function | `reference/tool-use.md` |
| Prompt or system-prompt authoring | prompt, system prompt, instructions | `reference/prompting.md`, `reference/subagent-prompts.md` |
| Evaluation | eval, test, measure, benchmark, accuracy | `reference/evaluation.md` |
| Safety / hardening | injection, jailbreak, hallucination, leak, guardrail | `reference/guardrails.md` |

If the request spans buckets, pick the dominant one and note the secondary. Do not read all seven reference files — that defeats progressive disclosure.

### Step 2: Define success criteria

Never start work against an undefined target. State two to four criteria that are **specific, measurable, achievable, and relevant**.

Bad: "good performance."
Good: "accurate sentiment classification, 95% agreement with the labeled set across 100 cases."

Even subjective goals quantify. Bad: "safe outputs." Good: "fewer than 0.1% of outputs across 10,000 trials flagged for toxicity by the content filter."

If the task is a one-off piece of writing rather than a system, criteria are still required, just lighter: audience, length, tone, and the single thing the deliverable must accomplish.

State the criteria to the user in one or two lines before proceeding. Do not ask permission to continue — proceed unless a criterion is genuinely undecidable without input.

### Step 3: Choose model and effort

Read `reference/model-selection.md`. State the choice and the reason in one sentence.

Defaults when the user has not specified:

- Classification, moderation, screening, routing, high-volume, latency-critical → **Haiku 4.5**
- Structured extraction, predictable pipelines, frontend, computer use → **Sonnet 5**
- Complex agentic coding, enterprise work, support chat, high-accuracy document work → **Opus 5**
- Tasks measured in hours or days, hardest unsolved problems → **Fable 5**

Effort defaults to `high`. Lower it to `medium` or `low` for routine work; raise it to `xhigh` for the hardest coding and agentic work. Effort is the primary cost and latency control.

### Step 4: Construct the prompt

Read `reference/prompting.md`. Apply, in this order:

1. **Be explicit.** State the desired output format and constraints. If you want above-and-beyond work, ask for it — do not expect it to be inferred from a vague brief.
2. **Give the reason, not only the request.** Explain why the instruction matters; the model generalizes from the explanation.
3. **Put the bulk of the prompt in the first user turn.** The system prompt is for the role, and little else.
4. **Structure with XML tags.** `<instructions>`, `<context>`, `<examples>`, `<input>` — consistent, descriptive, nested where there is a natural hierarchy.
5. **Add three to five examples**, wrapped in `<example>` tags inside `<examples>`. Make them relevant, diverse, and covering edge cases.
6. **Long inputs go at the top**, above the query and instructions. For documents over 20k tokens, ask for relevant quotes to be extracted first, then reasoned over.
7. **Say what to do, not what to avoid.** "Write in flowing prose paragraphs" beats "do not use markdown."

Then run the removal pass — these patterns actively hurt on current models:

- Delete emphatic pressure: "CRITICAL:", "You MUST", "If in doubt, use [tool]". These cause overtriggering. Plain "Use [tool] when…" is correct.
- Delete verification instructions ("double-check your answer", "add a final verification step", "use a subagent to verify") when targeting a model that self-verifies. Remove them; do not rewrite them.
- Delete forced progress-update scaffolding ("after every 3 tool calls, summarize progress").
- Delete any instruction to reproduce, transcribe, or explain internal reasoning in the response text.
- Delete prefilled assistant turns and `budget_tokens`; both are removed from current models.

Do not do this pass by eye. Run the linter, which checks every pattern above mechanically:

```bash
python scripts/lint_prompt.py path/to/prompt.txt
```

It reads a file, a directory, or stdin, and reports each finding with the reason and the fix. Errors correspond to patterns that return an HTTP 400; warnings still run but produce worse output. Structural findings are advice — a long prompt with no XML tags, no examples, or a document placeholder with no quote grounding.

Run it on the user's prompt before rewriting, and on your own draft afterwards. If a finding is a deliberate exception, suppress it on that line with `lint-prompt: ignore=W001` rather than leaving it unexplained.

### Step 5: Decide on delegation

Read `reference/subagent-prompts.md` **before writing any subagent prompt**.

Delegate when the work is genuinely independent, parallelizable, or needs isolated context. Do not delegate work finishable in a handful of tool calls, sequential work needing shared state, or single-file edits.

When you do delegate, the subagent prompt must be self-contained — the subagent cannot see this conversation. It must carry the objective, the context it needs, the constraints, the exact return format, and what not to do. Launch independent subagents in a single message so they run in parallel.

### Step 6: Execute

Do the work. While working:

- Make independent tool calls in parallel; make dependent ones sequentially. Never guess a parameter to force parallelism.
- Keep untrusted third-party content inside `tool_result` blocks, never concatenated into instructions.
- Ground factual claims in what you actually read. Do not describe a file you have not opened.
- Report status against evidence. If something is unverified, say so.

### Step 7: Verify

Check the deliverable against the criteria stated in Step 2 — the specific ones, not a generic once-over.

Match the verification method to the stakes:

| Method | Use when |
|---|---|
| Code / exact match | The answer is checkable mechanically. Fastest and most reliable. |
| Self-review against the rubric | Judgment calls, moderate stakes. |
| Fresh-context subagent verifier | High stakes, long runs, or work you produced yourself. Outperforms self-critique because it has no attachment to the approach. |

Prefer more test cases with automated grading over few cases with hand grading.

### Step 8: Report

End with a short block naming what was applied. Not a lecture — four to six lines:

```
Applied: <model> at <effort> — <one-line reason>
Structure: <which prompting techniques>
Delegation: <none | N subagents, why>
Verified: <method> against <criteria>
Removed: <any anti-patterns stripped from the user's existing prompt, if applicable>
```

If the user handed you an existing prompt, list what you removed and why. That list is usually the most valuable part of the response.

---

## Bundled script

**`scripts/lint_prompt.py`** — flags patterns in a prompt that break or degrade on current models. Run it, do not read it.

```bash
python scripts/lint_prompt.py prompt.txt          # human-readable
python scripts/lint_prompt.py prompts/ --json     # machine-readable, whole directory
cat prompt.md | python scripts/lint_prompt.py -   # stdin
```

Exit code is 1 when an error-severity finding is present, 0 otherwise. Use `--fail-on warning` in a CI gate, `--min-severity warning` to hide structural advice.

## Reference files

Read only what Step 1 selected.

- **[reference/model-selection.md](reference/model-selection.md)** — model comparison, pricing, effort levels, thinking defaults, per-model behavioral quirks and migration traps.
- **[reference/prompting.md](reference/prompting.md)** — core techniques, output and format control, verbosity, life after prefill, the removal list.
- **[reference/subagent-prompts.md](reference/subagent-prompts.md)** — when to delegate, the subagent prompt template, parallel dispatch, verifier subagents.
- **[reference/tool-use.md](reference/tool-use.md)** — tool definition design, `tool_choice`, parallel calls, tool context management, tool search, programmatic tool calling.
- **[reference/evaluation.md](reference/evaluation.md)** — success criteria, eval design principles, grading methods, LLM-grader rubrics.
- **[reference/guardrails.md](reference/guardrails.md)** — hallucinations, output consistency, jailbreaks, indirect prompt injection, prompt leak.
- **[reference/agent-architecture.md](reference/agent-architecture.md)** — long-horizon runs, memory, context editing and compaction, stop reasons, budgets, state tracking.

---

## Rules that apply to every task

These hold regardless of bucket. They are the short version of the whole skill.

**Choose before you optimize.** Selecting a different model often fixes latency and cost faster than any prompt change. Not every failing evaluation is a prompting problem.

**Get it working before you make it fast.** Engineer a prompt that performs well without model or token constraints first, then apply latency reduction. Optimizing early hides what peak quality looks like.

**Examples beat instructions.** Showing three to five well-chosen examples steers format, tone, and structure more reliably than describing them.

**Positive examples beat prohibitions.** Demonstrating the style you want works better than listing what to avoid.

**Less pressure, not more.** Current models are highly responsive to the system prompt. Emphatic instructions written to fight older models' reluctance now cause the opposite failure.

**Structure beats filtering for untrusted content.** Placing third-party content in `tool_result` blocks and JSON-encoding it does more than any regex filter.

**Uncertainty is a feature.** Explicitly permitting "I don't have enough information to confidently assess this" measurably reduces fabrication.

**One task, one prompt.** When a task has distinct subtasks, prompt and evaluate each separately rather than writing one prompt that does everything adequately.

---

## Old patterns

<details>
<summary>Techniques that no longer apply on current models</summary>

**Prefilled assistant responses.** Providing a partial assistant message on the last turn to control format or skip preamble. Removed; requests containing it return a 400. Use structured outputs, or an explicit instruction such as "Do not preamble." Assistant messages elsewhere in the conversation are unaffected.

**`budget_tokens` / manual extended thinking.** Setting a fixed thinking budget. Removed on current models; returns a 400. Use adaptive thinking and control depth with `effort`; use `max_tokens` as the hard ceiling.

**Sampling parameters for stylistic variety.** `temperature`, `top_p`, and `top_k` at non-default values return a 400 on current Sonnet-class models. Guide tone through the system prompt, or ask the model to propose several distinct directions and pick one.

**Anti-laziness prompting.** Instructions written to force tool use or thoroughness on older models — "CRITICAL: You MUST use this tool", "If in doubt, use [tool]", "Default to using [tool]". These now cause overtriggering. Dial them back to neutral phrasing or remove them.

</details>

Files in this skill

  • SKILL.md13.5 KB
  • reference/agent-architecture.md13.9 KB
  • reference/evaluation.md5.8 KB
  • reference/guardrails.md8.9 KB
  • reference/model-selection.md8.6 KB
  • reference/prompting.md12.8 KB
  • reference/subagent-prompts.md11.5 KB
  • reference/tool-use.md9.8 KB
  • scripts/lint_prompt.py15.3 KB

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…