Applies Anthropic's documented best practices for building with Claude to the task at hand — model and effort selection, prompt construction, subagent delegation, tool design, evaluation, and guardrails. Use when the user invokes /best-practice, or asks for work done "by best practices", "the way Anthropic recommends", "as the official documentation recommends", or uses the Russian equivalents "по лучшим практикам", "сделай как рекомендует Anthropic", "как в официальной документации". Also us...
Installs into .claude/skills of the current project.
Are you the author of Best Practice?
Add the live security badge to your README. It updates with every re-scan.
[](https://www.skillsdirectory.com/skills/starsinc1708-best-practice)
---
name: best-practice
description: Applies Anthropic's documented best practices for building with Claude to the task at hand — model and effort selection, prompt construction, subagent delegation, tool design, evaluation, and guardrails. Use when the user invokes /best-practice, or asks for work done "by best practices", "the way Anthropic recommends", "as the official documentation recommends", or uses the Russian equivalents "по лучшим практикам", "сделай как рекомендует Anthropic", "как в официальной документации". Also use when the user asks how to prompt Claude, choose a model, set effort, write a subagent prompt, design evaluations, or harden an agent against prompt injection.
---
# Best Practice
Do the requested work the way Anthropic's documentation says to do it — and make the reasoning visible, so the user learns the rule rather than just receiving the output.
This skill has two modes. Detect which one applies before doing anything else.
**Execution mode** — the user wants a task done, and wants it done properly.
Trigger: `/best-practice <task>`, or a task request carrying a quality qualifier ("по лучшим практикам", "the way Anthropic recommends", "properly", "как в документации").
Behavior: run the workflow below, then produce the deliverable.
**Advisory mode** — the user is asking what the best practice *is*.
Trigger: "how should I prompt for X", "which model for Y", "how do I write a subagent prompt", "is this prompt good".
Behavior: skip to the relevant reference file, answer with the rule and its source, skip the deliverable.
---
## Workflow (execution mode)
Copy this checklist into your working notes and check items off as you go.
```
Best Practice Progress:
- [ ] Step 1: Classify the task
- [ ] Step 2: Define success criteria
- [ ] Step 3: Choose model and effort
- [ ] Step 4: Construct the prompt
- [ ] Step 5: Decide on delegation
- [ ] Step 6: Execute
- [ ] Step 7: Verify against the criteria from Step 2
- [ ] Step 8: Report what was applied
```
### Step 1: Classify the task
Put the request in exactly one bucket. The bucket determines which reference files you read.
| Bucket | Signals | Read |
|---|---|---|
| Generation | write, draft, build, create, design | `reference/prompting.md` |
| Classification | route, categorize, label, moderate, triage | `reference/evaluation.md`, `reference/prompting.md` |
| Extraction / summarization | summarize, extract, pull out, condense | `reference/prompting.md` |
| Agentic / multi-step | automate, orchestrate, agent, pipeline, long-running | `reference/agent-architecture.md`, `reference/tool-use.md` |
| Tool or integration design | tool, API, MCP, connector, function | `reference/tool-use.md` |
| Prompt or system-prompt authoring | prompt, system prompt, instructions | `reference/prompting.md`, `reference/subagent-prompts.md` |
| Evaluation | eval, test, measure, benchmark, accuracy | `reference/evaluation.md` |
| Safety / hardening | injection, jailbreak, hallucination, leak, guardrail | `reference/guardrails.md` |
If the request spans buckets, pick the dominant one and note the secondary. Do not read all seven reference files — that defeats progressive disclosure.
### Step 2: Define success criteria
Never start work against an undefined target. State two to four criteria that are **specific, measurable, achievable, and relevant**.
Bad: "good performance."
Good: "accurate sentiment classification, 95% agreement with the labeled set across 100 cases."
Even subjective goals quantify. Bad: "safe outputs." Good: "fewer than 0.1% of outputs across 10,000 trials flagged for toxicity by the content filter."
If the task is a one-off piece of writing rather than a system, criteria are still required, just lighter: audience, length, tone, and the single thing the deliverable must accomplish.
State the criteria to the user in one or two lines before proceeding. Do not ask permission to continue — proceed unless a criterion is genuinely undecidable without input.
### Step 3: Choose model and effort
Read `reference/model-selection.md`. State the choice and the reason in one sentence.
Defaults when the user has not specified:
- Classification, moderation, screening, routing, high-volume, latency-critical → **Haiku 4.5**
- Structured extraction, predictable pipelines, frontend, computer use → **Sonnet 5**
- Complex agentic coding, enterprise work, support chat, high-accuracy document work → **Opus 5**
- Tasks measured in hours or days, hardest unsolved problems → **Fable 5**
Effort defaults to `high`. Lower it to `medium` or `low` for routine work; raise it to `xhigh` for the hardest coding and agentic work. Effort is the primary cost and latency control.
### Step 4: Construct the prompt
Read `reference/prompting.md`. Apply, in this order:
1. **Be explicit.** State the desired output format and constraints. If you want above-and-beyond work, ask for it — do not expect it to be inferred from a vague brief.
2. **Give the reason, not only the request.** Explain why the instruction matters; the model generalizes from the explanation.
3. **Put the bulk of the prompt in the first user turn.** The system prompt is for the role, and little else.
4. **Structure with XML tags.** `<instructions>`, `<context>`, `<examples>`, `<input>` — consistent, descriptive, nested where there is a natural hierarchy.
5. **Add three to five examples**, wrapped in `<example>` tags inside `<examples>`. Make them relevant, diverse, and covering edge cases.
6. **Long inputs go at the top**, above the query and instructions. For documents over 20k tokens, ask for relevant quotes to be extracted first, then reasoned over.
7. **Say what to do, not what to avoid.** "Write in flowing prose paragraphs" beats "do not use markdown."
Then run the removal pass — these patterns actively hurt on current models:
- Delete emphatic pressure: "CRITICAL:", "You MUST", "If in doubt, use [tool]". These cause overtriggering. Plain "Use [tool] when…" is correct.
- Delete verification instructions ("double-check your answer", "add a final verification step", "use a subagent to verify") when targeting a model that self-verifies. Remove them; do not rewrite them.
- Delete forced progress-update scaffolding ("after every 3 tool calls, summarize progress").
- Delete any instruction to reproduce, transcribe, or explain internal reasoning in the response text.
- Delete prefilled assistant turns and `budget_tokens`; both are removed from current models.
Do not do this pass by eye. Run the linter, which checks every pattern above mechanically:
```bash
python scripts/lint_prompt.py path/to/prompt.txt
```
It reads a file, a directory, or stdin, and reports each finding with the reason and the fix. Errors correspond to patterns that return an HTTP 400; warnings still run but produce worse output. Structural findings are advice — a long prompt with no XML tags, no examples, or a document placeholder with no quote grounding.
Run it on the user's prompt before rewriting, and on your own draft afterwards. If a finding is a deliberate exception, suppress it on that line with `lint-prompt: ignore=W001` rather than leaving it unexplained.
### Step 5: Decide on delegation
Read `reference/subagent-prompts.md` **before writing any subagent prompt**.
Delegate when the work is genuinely independent, parallelizable, or needs isolated context. Do not delegate work finishable in a handful of tool calls, sequential work needing shared state, or single-file edits.
When you do delegate, the subagent prompt must be self-contained — the subagent cannot see this conversation. It must carry the objective, the context it needs, the constraints, the exact return format, and what not to do. Launch independent subagents in a single message so they run in parallel.
### Step 6: Execute
Do the work. While working:
- Make independent tool calls in parallel; make dependent ones sequentially. Never guess a parameter to force parallelism.
- Keep untrusted third-party content inside `tool_result` blocks, never concatenated into instructions.
- Ground factual claims in what you actually read. Do not describe a file you have not opened.
- Report status against evidence. If something is unverified, say so.
### Step 7: Verify
Check the deliverable against the criteria stated in Step 2 — the specific ones, not a generic once-over.
Match the verification method to the stakes:
| Method | Use when |
|---|---|
| Code / exact match | The answer is checkable mechanically. Fastest and most reliable. |
| Self-review against the rubric | Judgment calls, moderate stakes. |
| Fresh-context subagent verifier | High stakes, long runs, or work you produced yourself. Outperforms self-critique because it has no attachment to the approach. |
Prefer more test cases with automated grading over few cases with hand grading.
### Step 8: Report
End with a short block naming what was applied. Not a lecture — four to six lines:
```
Applied: <model> at <effort> — <one-line reason>
Structure: <which prompting techniques>
Delegation: <none | N subagents, why>
Verified: <method> against <criteria>
Removed: <any anti-patterns stripped from the user's existing prompt, if applicable>
```
If the user handed you an existing prompt, list what you removed and why. That list is usually the most valuable part of the response.
---
## Bundled script
**`scripts/lint_prompt.py`** — flags patterns in a prompt that break or degrade on current models. Run it, do not read it.
```bash
python scripts/lint_prompt.py prompt.txt # human-readable
python scripts/lint_prompt.py prompts/ --json # machine-readable, whole directory
cat prompt.md | python scripts/lint_prompt.py - # stdin
```
Exit code is 1 when an error-severity finding is present, 0 otherwise. Use `--fail-on warning` in a CI gate, `--min-severity warning` to hide structural advice.
## Reference files
Read only what Step 1 selected.
- **[reference/model-selection.md](reference/model-selection.md)** — model comparison, pricing, effort levels, thinking defaults, per-model behavioral quirks and migration traps.
- **[reference/prompting.md](reference/prompting.md)** — core techniques, output and format control, verbosity, life after prefill, the removal list.
- **[reference/subagent-prompts.md](reference/subagent-prompts.md)** — when to delegate, the subagent prompt template, parallel dispatch, verifier subagents.
- **[reference/tool-use.md](reference/tool-use.md)** — tool definition design, `tool_choice`, parallel calls, tool context management, tool search, programmatic tool calling.
- **[reference/evaluation.md](reference/evaluation.md)** — success criteria, eval design principles, grading methods, LLM-grader rubrics.
- **[reference/guardrails.md](reference/guardrails.md)** — hallucinations, output consistency, jailbreaks, indirect prompt injection, prompt leak.
- **[reference/agent-architecture.md](reference/agent-architecture.md)** — long-horizon runs, memory, context editing and compaction, stop reasons, budgets, state tracking.
---
## Rules that apply to every task
These hold regardless of bucket. They are the short version of the whole skill.
**Choose before you optimize.** Selecting a different model often fixes latency and cost faster than any prompt change. Not every failing evaluation is a prompting problem.
**Get it working before you make it fast.** Engineer a prompt that performs well without model or token constraints first, then apply latency reduction. Optimizing early hides what peak quality looks like.
**Examples beat instructions.** Showing three to five well-chosen examples steers format, tone, and structure more reliably than describing them.
**Positive examples beat prohibitions.** Demonstrating the style you want works better than listing what to avoid.
**Less pressure, not more.** Current models are highly responsive to the system prompt. Emphatic instructions written to fight older models' reluctance now cause the opposite failure.
**Structure beats filtering for untrusted content.** Placing third-party content in `tool_result` blocks and JSON-encoding it does more than any regex filter.
**Uncertainty is a feature.** Explicitly permitting "I don't have enough information to confidently assess this" measurably reduces fabrication.
**One task, one prompt.** When a task has distinct subtasks, prompt and evaluate each separately rather than writing one prompt that does everything adequately.
---
## Old patterns
<details>
<summary>Techniques that no longer apply on current models</summary>
**Prefilled assistant responses.** Providing a partial assistant message on the last turn to control format or skip preamble. Removed; requests containing it return a 400. Use structured outputs, or an explicit instruction such as "Do not preamble." Assistant messages elsewhere in the conversation are unaffected.
**`budget_tokens` / manual extended thinking.** Setting a fixed thinking budget. Removed on current models; returns a 400. Use adaptive thinking and control depth with `effort`; use `max_tokens` as the hard ceiling.
**Sampling parameters for stylistic variety.** `temperature`, `top_p`, and `top_k` at non-default values return a 400 on current Sonnet-class models. Guide tone through the system prompt, or ask the model to propose several distinct directions and pick one.
**Anti-laziness prompting.** Instructions written to force tool use or thoroughness on older models — "CRITICAL: You MUST use this tool", "If in doubt, use [tool]", "Default to using [tool]". These now cause overtriggering. Dial them back to neutral phrasing or remove them.
</details>