Design rulebook and diagnostic layer for web UI, distilled from Refactoring UI by Adam Wathan and Steve Schoger. Use when building a new page or component, refining UI that looks off or amateurish, running a design review, or establishing design tokens (color shades, type scale, spacing scale, elevation, radius). Covers visual hierarchy, layout and spacing, typography, color and contrast, depth and shadows, images, and finishing touches, with Tailwind mappings throughout. Also triggers on non...
Installs into .claude/skills of the current project.
Are you the author of Refactoring Ui?
Add the live security badge to your README. It updates with every re-scan.
[](https://www.skillsdirectory.com/skills/edisonmbli-refactoring-ui)
---
name: refactoring-ui
description: Design rulebook and diagnostic layer for web UI, distilled from Refactoring UI by Adam Wathan and Steve Schoger. Use when building a new page or component, refining UI that looks off or amateurish, running a design review, or establishing design tokens (color shades, type scale, spacing scale, elevation, radius). Covers visual hierarchy, layout and spacing, typography, color and contrast, depth and shadows, images, and finishing touches, with Tailwind mappings throughout. Also triggers on non-English requests for the same things - for example Chinese: 页面/界面设计, 样式调整, UI 太丑, 设计评审, 视觉层级, 配色, 对比度, 间距, 排版, 阴影, 空状态, 设计规范, 设计系统. Not for backend logic, copy editing, or changes with no visual surface.
---
# Refactoring UI
## What this skill is
This skill is the **constraint and diagnosis layer** for interface work. It supplies the rules a design must satisfy and finds where a design breaks them.
It is deliberately **not a generator of aesthetics**. When a design-generation capability is present (a Design mode, `frontend-design`, a canvas tool), that capability decides what the interface *looks like*; this skill decides what it must *obey*, and afterwards checks whether it did. See [Composing with design generators](#composing-with-design-generators).
The rules come from a book whose authors went on to build Tailwind CSS. Its scales are Tailwind's defaults, so the path from principle to code is unusually short.
## Step 0 — Detect before you decide
**Never make a design judgment before running this.** It takes one pass and prevents the two worst failure modes: fighting an existing design system, and emitting code the project can't use.
Detect and state briefly:
1. **Framework** — React/Vue/Svelte/plain HTML; SSR or SPA.
2. **Tailwind, and which major version.** v4 declares theme in CSS via `@theme`; v3 uses `tailwind.config.js`. Check `package.json`, then look for `@import "tailwindcss"` (v4) vs `@tailwind base` (v3).
3. **Existing design tokens** — CSS custom properties, a theme object, a `tokens.json`, a `theme/` directory.
4. **Component library** — shadcn/ui, MUI, Ant Design, Chakra, Radix. These carry their own token contracts — for shadcn/ui, load `16-component-libraries.md` before writing or editing any component.
5. **Dark mode** — present, planned, or absent.
**The prime directive that follows from this: an existing system always wins.** Where the project has a system, translate the book's principles into that system's vocabulary. Where it has none, propose one. Never overwrite, never silently migrate, never add a build-time dependency the user didn't ask for.
## Reading the reference files
The reference files mix two kinds of content, and confusing them is the main way this skill can do damage. The notation keeps them apart.
**Invariants — the book's actual rules.** In every `Values` block, **bold numbers are invariants**: font weights 400/500 and 600/700, a maximum of three text colors, one primary action per page, 45–75 characters per line, 4.5:1 and 3:1 contrast, no two scale steps closer than 25%, hue rotation capped at 20–30°, five elevation levels. These do not vary by project. Apply them as written.
**Illustrative choices — everything else.** Unbolded values in `Values`, and every size, spacing and color in a code example, are one valid instance of the rule, not the required answer.
**Colors in examples are always placeholders**, written in braces so they cannot be pasted verbatim:
| Placeholder | Resolves to |
|---|---|
| `{primary}` | the project's primary/brand ramp |
| `{neutral}` | the project's grey ramp — `gray`, `slate`, `zinc`, `stone`, or a custom one |
| `{danger}` | the project's destructive ramp |
| `{accent}` | a semantic accent ramp (warning, success, info) |
`bg-{primary}-600` means "the 600 stop of whatever this project's primary is." Never emit a brace into real code, and never substitute Tailwind's own default palette names unless Step 0 confirmed the project actually uses them. Defaulting to indigo and gray because the examples showed them is the single most recognizable tell of machine-generated UI.
**Resolution order for any concrete value:**
1. A token that already exists in the project (found in Step 0).
2. A token generated by Workflow A in this session.
3. The book's default scale from the relevant reference file — and say that you fell back to it.
If none of the three yields a value, that gap is itself a finding: the project needs a token before this decision can be made consistently.
## Route by intent
| The user is asking to… | Workflow | Load |
|---|---|---|
| set up a theme, palette, design system, "we have no design standards" | **A · Tokens** | `11-design-tokens.md` |
| build a new page, screen, component, feature UI | **B · Build** | `12-component-recipes.md` + lens refs as needed |
| improve / clean up / "make this look better" / "this looks off" | **C · Refine** | `13-audit-rubric.md`, then lens refs |
| review, critique, "what's wrong with this", pre-launch check | **D · Audit** | `13-audit-rubric.md` + `14-antipatterns.md` |
| understand a specific principle, "why does this look bad" | **Explain** | the single relevant lens ref |
Intent is often mixed. "Build me a settings page in this app that has no design system" is A then B. Run them in order and say so.
## The Twelve Laws
These apply to every task. They are the highest-cost rules to violate and need no reference lookup.
1. **Hierarchy is the job.** Making something look designed is mostly about making importance visible — not about styling. A flat wall of equally-weighted content reads as noise no matter how good the colors are. `§2.1`
2. **Emphasize by de-emphasizing.** When the primary element won't stand out, soften what competes with it rather than shouting louder. `§2.4`
3. **Every value comes from a system.** Font sizes, spacing, colors, shadows, radii, border widths — all chosen from a predefined scale. No hand-tuned one-offs, no arbitrary values. `§1.5 §3.2 §4.1`
4. **Start with too much white space and remove it.** Space added until something stops looking bad always undershoots. `§3.1`
5. **Space around a group must exceed space within it.** This one rule fixes most "why does this look confusing" problems. `§3.6`
6. **Size is not the only lever.** Use weight and color. Two weights (400/500 and 600/700) and three text colors (dark / grey / lighter grey) carry almost all UI hierarchy. Never go below 400 for UI text. `§2.2`
7. **Never put grey text on a colored background.** The effect you want is reduced contrast, not greyness. Hand-pick a color at the background's hue with adjusted saturation and lightness. Not white-at-lower-opacity. `§2.3`
8. **Semantics are secondary to hierarchy.** One primary action per page, a couple of secondary, a few tertiary. A destructive action is not automatically a loud one. `§2.8`
9. **Choose elements semantically, style them hierarchically.** An `h1` is a document-structure decision, not a font-size decision. Section titles usually behave as labels and should be small. `§2.6`
10. **Light comes from above.** Raised = lighter top edge plus a small, sharp shadow below. Inset = shadow at the top, lighter bottom edge. Everything about depth follows from this. `§6.1`
11. **Text measures 45–75 characters (20–35em).** Line-height rises with line length (1.5 narrow → 2 wide) and falls as font size grows (large headlines can sit at 1). `§4.3 §4.5`
12. **Never rely on color alone.** Color supports a signal the design already carries some other way. `§5.7`
## Workflow A · Tokens
Establishing a design system where none exists, or filling gaps in a partial one.
1. **Interview.** Load `11-design-tokens.md` and run its question script — at most four rounds, batched, every question skippable. If the user says "just pick for me," take the documented defaults and go straight to step 2.
2. **Generate.** Run `scripts/generate_palette.py` for each color ramp. The book's algorithm is base(500) → edges(900/100) → midpoints(700/300) → fill(800/600/400/200), with saturation raised as lightness moves away from 50% and hue rotated no more than 20–30° toward 60/180/300 to lighten or 0/120/240 to darken. If Python is unavailable, follow the manual procedure in `11-design-tokens.md` — same algorithm, done by hand.
3. **Verify.** Run `scripts/check_contrast.py` over every foreground/background pairing the system implies. 4.5:1 for normal text, 3:1 for large.
4. **Emit** the five deliverables defined in `11-design-tokens.md`: `design-tokens.json`, the Tailwind theme block for the detected version, `tokens.css`, a human-readable `DESIGN.md`, and `preview.html`.
Do not invent hex values by hand. The palette algorithm exists precisely because eyeballed shades drift.
## Workflow B · Build
The order is the method. Following it out of order is how interfaces end up over-decorated and under-structured.
1. **Start with a feature, not a shell.** Design the actual functionality before the navigation, container, or logo. `§1.1`
2. **Work in grayscale first.** Force spacing, size and contrast to carry the hierarchy before any color exists. `§1.2.1`
3. **Rank the content** into primary / secondary / tertiary, and the actions into one primary / few secondary / several tertiary. `§2.1 §2.8`
4. **Lay out from the scale.** Spacing, sizing and type all drawn from the system; group spacing obeys Law 5. `§3.2 §3.6`
5. **Set the measure and rhythm.** 45–75 characters, line-height matched to width and size. `§4.3 §4.5`
6. **Then color**, then **depth**, then **finishing touches** — in that order. `§5 §6 §8`
7. **Design the empty state as part of the feature**, not afterward. Hide supporting UI (tabs, filters) that does nothing until content exists. `§8.4`
8. **Verify visually** — see [Visual verification](#visual-verification).
Consult `12-component-recipes.md` for buttons, cards, form groups, tables, dropdowns, modals, alerts and empty states rather than re-deriving them.
**Steps 1–3 must leave a trace, not just happen in your head.** A single generation pass tends to compress "grayscale first, then rank, then color" into "write the colored markup directly" — the order gets lost because nothing forces it to surface. Before the first line of markup, write a short constraint brief as a leading comment in the file you're about to produce: the content ranking from step 3 (primary / secondary / tertiary, one primary action), and the token values steps 4–6 will draw from. It's the process made checkable, not a formality — a reviewer (human or a later Workflow D pass) can see the ranking was actually decided instead of trusting that it was.
## Workflow C · Refine
Audit first, then change. Never rewrite a component wholesale because it "feels off."
1. Run Workflow D to produce findings.
2. Order the fixes by the sweep order below — structural fixes first, because they frequently dissolve the cosmetic complaints. Adding white space often removes the question of whether a border was needed.
3. Apply changes as a reviewable diff. Each change cites its rule ID.
4. Re-verify visually.
## Workflow D · Audit
Diagnose only. Do not edit files in this workflow unless the user asks.
Load `13-audit-rubric.md` for the severity model and the finding schema. The report is a table:
| # | Lens | Severity | Location | Rule | Now | Proposed | Effort |
Severity:
- **P0** — hierarchy has collapsed, text is unreadable, or contrast fails WCAG. Usability is affected.
- **P1** — off-system values, ambiguous group spacing, grey on color, unranked actions. Consistency is affected.
- **P2** — missing polish, unexploited depth, default-looking components.
Report P0s even when the user asked about something else.
## Depth modes
| Mode | When | Behavior |
|---|---|---|
| **Focused** | user names a direction ("look at the colors") | Load 1–2 lens refs, single pass |
| **Sweep** *(default for "review this")* | broad request, one screen or component | Run the self-contained checklist in `13-audit-rubric.md`, lens by lens, in order — no chapter file needed. Load one only to explain a rule or settle a disputed finding |
| **Parallel** | large surface, many files, **and the user agrees** | One subagent per lens, each returning the finding schema; then merge |
**Sweep order is fixed and not negotiable:**
```
hierarchy → layout & spacing → typography → color → depth → images → finishing touches
```
This is the book's own ordering and it is causal: hierarchy problems change what counts as a spacing problem, spacing changes what counts as a border problem, and finishing touches are meaningless on a broken structure.
**Merging parallel results:** deduplicate on `(file, line, rule)`. When two lenses propose conflicting changes to the same element, the earlier lens in the sweep order wins. Never present the raw concatenation of subagent output — merge, deduplicate, re-rank by severity, and cut anything that survived only as a restatement.
**Parallel is opt-in.** A cold subagent has to rebuild context that this session already holds, so it only pays off across a genuinely large surface. Ask before spawning.
## Visual verification
Reading source code is not a design review. Whenever the project can be rendered:
1. Start the preview server and open the page.
2. Screenshot at desktop (1280) **and** mobile (375).
3. Judge from the pixels — actual measure, actual contrast, actual group spacing, whether breakpoints collapse.
Static analysis cannot see a line that runs to 110 characters, a heading whose weight vanishes against its background, or a card grid that stacks wrong at 400px. Findings that came only from reading code should be marked as unverified.
## Composing with design generators
When a Design mode or design-generation skill is available, do not compete with it. Compose in two positions:
- **Pre-flight — supply the constraints.** Before generation, emit a brief containing the token values it must use and the prohibitions it must respect (the Twelve Laws, plus anything from `14-antipatterns.md` relevant to the surface). A generator inside good constraints produces work that needs far less correction.
- **Post-flight — audit the output.** After generation, run Workflow D against the result. This is the structurally missing step in every generate-only flow: a generator does not come back and critique its own work.
When no such capability exists, Workflow B stands alone.
## Scope guard
- Change only what the user asked about. Report other findings; don't act on them.
- Never install a dependency, never migrate a styling approach, never introduce Tailwind into a project that doesn't use it. Recommend, and let the user decide.
- Never replace an existing design system with the book's defaults. Map onto it instead.
- Don't manufacture a diff to demonstrate that a rule was applied. If the current implementation already satisfies the rule, say so and move on.
- When a rule and a deliberate project convention conflict, surface the conflict; the convention wins until the user says otherwise.
## Output language
Reference files and this skill are written in English. **Deliverables follow the user's language** — audit reports, `DESIGN.md`, interview questions and explanations are written in whatever language the user is writing in.
## Reference index
| File | Load when |
|---|---|
| `00-coverage-matrix.md` | maintaining this skill; verifying no rule was dropped |
| `01-starting-from-scratch.md` | process, fidelity, personality, choosing constraints |
| `02-hierarchy.md` | anything about emphasis, importance, labels, buttons, weight |
| `03-layout-spacing.md` | spacing, white space, widths, grids, responsive sizing |
| `04-typography.md` | type scale, fonts, measure, line-height, alignment, letter-spacing |
| `05-color.md` | palettes, shades, HSL, greys, contrast, accessibility |
| `06-depth.md` | shadows, elevation, raised/inset, layering |
| `07-images.md` | photos, text over images, icon and screenshot scaling, uploads |
| `08-finishing-touches.md` | polish, accents, backgrounds, empty states, borders, component reinvention |
| `10-tailwind-mapping.md` | translating any rule into Tailwind v3 or v4 |
| `11-design-tokens.md` | Workflow A, always |
| `12-component-recipes.md` | building or fixing a specific component |
| `13-audit-rubric.md` | Workflows C and D, always |
| `14-antipatterns.md` | diagnosing from a symptom; reviewing AI-generated UI |
| `15-beyond-the-book.md` | dark mode, focus states, target sizes, motion, breakpoints, z-index |
| `16-component-libraries.md` | Step 0 detected shadcn/ui (or another component library with its own token contract) |
Everything in `15-beyond-the-book.md` is an extension, not the book's position, and is labeled as such wherever it appears.
## Continuous improvement
Two habits worth passing to the user when the moment fits (`§9`):
- On any design you admire, ask what the designer did that **you would never have thought to do** — an inverted datepicker background, a button placed inside a text input, two colors in a single headline. The unintuitive choices are where new technique comes from.
- Rebuild interfaces you like from scratch **without opening devtools**. Working out why your version looks different is how details like tightened heading line-height, letter-spaced uppercase, and layered shadows get learned rather than memorized.
---
*Rules derived from* Refactoring UI *by Adam Wathan & Steve Schoger — [refactoringui.com](https://www.refactoringui.com). This skill is an independent reimplementation of the book's principles as executable guidance; it is not a substitute for reading it.*