Review and rewrite copy, web UI, slides, READMEs, commits, PRs, marketing pages, educational chapters, class materials, notebooks, and generated media to strip AI-isms — Claudeisms, GPT-isms, Codexisms, Geminisms, engineering-artifact slop, and the v0/Lovable design look — producing a line-item fix plan as static HTML. Use before shipping any outward-facing text or design, when something "reads like AI", or when auditing a property for machine tells. NOT for grammar/spell checking, SEO optimi...
Installs into .claude/skills of the current project.
Are you the author of Make Copy And Media Human?
Add the live security badge to your README. It updates with every re-scan.
[](https://www.skillsdirectory.com/skills/curiositech-make-copy-and-media-human)
---
license: Apache-2.0
name: make_copy_and_media_human
description: Review and rewrite copy, web UI, slides, READMEs, commits, PRs, marketing pages, educational chapters, class materials, notebooks, and generated media to strip AI-isms — Claudeisms, GPT-isms, Codexisms, Geminisms, engineering-artifact slop, and the v0/Lovable design look — producing a line-item fix plan as static HTML. Use before shipping any outward-facing text or design, when something "reads like AI", or when auditing a property for machine tells. NOT for grammar/spell checking, SEO optimization, plagiarism detection, or judging whether a third party used AI (this is an editing skill, and the evidence says authorship detection does not work).
allowed-tools: Read,Write,Edit,Bash,Grep,Glob,WebFetch
argument-hint: '[file-or-directory] [--out report.html] [--findings findings.json] [--json findings.json]'
metadata:
provenance:
kind: first-party
owners: [port-daddy]
scope: public
authorship:
maintainers: [port-daddy]
category: Writing & Editing
tags:
- humanize
- copy-editing
- ai-isms
- design-review
- engineering-artifacts
- voice
pairs-with:
- skill: agent-pr-authoring
reason: that skill writes the PR; this skill catches the generated scaffold, the unverified test plan, and the commit-style drift before a maintainer does
- skill: port-daddy-marketing-copy
reason: that skill drafts portdaddy.dev copy in the house voice; this skill audits the draft for machine tells before publish
- skill: port-daddy-expository-writer
reason: that skill drafts long-form concept/tutorial prose; this skill catches the cadence and structure tells that voice guide alone won't
- skill: web-design-expert
reason: that skill produces the visual design; this skill hunts the v0/Lovable generated-look tells (defaults, glassmorphism, indigo) in the result
io-contract:
kind: deliverable
consumes:
- kind: draft-copy-or-media
format: markdown
- kind: humanization-judge-findings
format: json
produces:
- kind: humanized-copy-or-media
format: markdown
- kind: humanization-audit-findings
format: json
- kind: humanization-fix-plan-report
format: html
---
# Make Copy and Media Human
Strip the machine accent from anything outward-facing. This skill catalogs the
tells per model dialect and per medium, detects them in two layers, strikes them,
and hands you a line-item fix plan as one self-contained HTML file.
## Read this first
You are going to find things. Before you act on them, understand what a finding
means, because the evidence here is uncomfortable and the skill is built around
it.
A study of seven detectors found substantial false-positive rates on non-native
English essays. That result concerns those detectors and that corpus; it does
not establish the probability that any particular author used AI. Fonts, colors,
punctuation, and teaching mistakes cannot settle authorship either.
That is why this skill edits and never accuses. Each proposed fix must improve the artifact for its intended reader, which means you never
have to answer the authorship question to do the work. If someone asks you to
determine whether a colleague or a student used AI, decline and offer to edit the
text instead. `references/fairness-and-false-positives.md` has the full argument
and the citations, and you should read it before your first review.
## Three laws
**No topical keyword lists.** Phrase-level tropes get judged by you against a
rubric, never by substring matching over free text. Three narrow exceptions are
allowed because they are not about topic or taste. The first is a closed set of
<!-- humanize:ignore-start -->closed syntactic residues from metadata or transport, such
as `oaicite` tokens or a `utm_source=chatgpt.com` parameter.<!-- humanize:ignore-end --> The second is closed sets of grammatical
forms measured as a rate, like sentence-final participles or nominalization
suffixes, which carry published effect sizes. The third is the era-versioned
excess-vocabulary marker list, reported as density against a baseline and always
labelled with its era. Everything else belongs to the judge pass.
**Judge the delta, not the absolute.** This is the law that makes the skill fair
and also makes it work. Rhythm signals such as em-dash rate, comma rate,
contraction rate and sentence-length variance are fluency proxies, and they
overlap badly between people and models. Melville runs more em dashes than
GPT-4-class output does. Llama runs none at all. One anti-slop linter measured
its em-dash rule warning on roughly 64% of legitimate technical blog posts before
recalibration. So pass `--baseline` with a few pieces of the author's own prior
writing whenever you can get it. Without a baseline the script caps that whole
family at low severity, on purpose.
**The output has to pass its own review.** The report uses Georgia and Menlo, an
oxide-red accent, nothing under 14px, no emoji, no indigo. Review the bundle with its own scripts and report residual findings explicitly. If this skill's own artifacts looked generated, nothing it says would
land.
## The six families
Findings carry a `family`, and the family tells you how much to trust the
finding. This matters more than severity.
`residue` covers unconverted citation tokens, tracking parameters, and leaked
assistant boilerplate. Verify their context before removing them: examples and
quotations can contain these deliberately. Unicode spacing belongs to typography,
not provenance. U+202F is used in French locale formatting; preserve meaningful
spacing and script joiners unless a concrete rendering defect requires repair.
`form` covers grammatical patterns with measured effect sizes, like participial
tails at 5.3 times the human rate or nominalizations at roughly twice. Humans do
these too, so read the `false_positive_when` line before you cut.
`rhythm` covers punctuation and sentence-length habits. Weak on their own, real
against a baseline. This is where false accusations come from.
`shape` covers document assembly: heading density, bullet colonization, rules
between every section, the obligatory future-directions block. Usually safe to
act on, because the fix improves the document regardless.
Weight `form` and `shape` over the phrase-level items when the two disagree. The
largest community analysis of what makes writing sound like AI concluded against
its own premise on this, finding cosmetic tells mostly noise. The quantitative
version is stronger still: a study of 61,608 stories separated human from
generated fiction at 93.2% macro-F1 using discourse-level narrative features
alone, with every stylistic cue stripped out. Structure survives a model that has
learned not to say "delve". Phrases do not.
`code` covers engineering artifacts. The highest-precision checks here are all
relative, comparing a change against the repo's own log, idiom, and PR norms.
A contributor who read the surrounding code passes them automatically.
### Declared but not wired
A useful interaction-risk pattern is a visible affordance whose failure behavior
was never tested. Static source can contain an Escape handler without proving
that Escape closes the modal, restores focus, or works in the relevant state.
The catalog includes a browser-trial study of this gap; its numerical claims
were not live-verified in this research pass, so they are not a calibrated
severity threshold here.
So when auditing anything interactive, ask of each affordance whether its label
is present without its behaviour. A sort caret with no `aria-sort`. A disabled
button with no reason. A `<dialog>` written but opened with `.show()` rather than
`.showModal()`. An `onMutate` with no `onError`. A "Last 30 days" label with no
range control. Two related shapes travel with it: the interface that is a
one-to-one rendering of the data model rather than a designed view (every column
a column, every config key a switch), and the content whose cardinality exactly
equals a layout constant (four cards because the grid has four columns).
These patterns locate missing behavior. They do not distinguish a model
author from a human author who did not test the interface.
`defect` is the odd one out, and the most useful. These are not inferences about
who built something; they are things that are broken. The page scrolls sideways
at 390px. The button is not a button. The grey text fails contrast. You reproduce
them by opening the page, so they carry no fairness caveat at all and you act on
them with full confidence. Work this family first: a site that does not work on a
phone has a bigger problem than a site that reads a bit generated.
`defect` findings share their standing with one more class that sits outside the
families and outranks everything. Citation pathology is about truth rather than
authorship: a dead DOI, a reference that
does not support its sentence, a statistic with no study behind it. Check those
with complete confidence, because you are verifying a claim rather than inferring
an author. It is the highest-yield check in the skill and it has no fairness cost.
## The one rubric that matters most
If you only ever run one judge question, run this one: **what detail in this
passage could only have come from this author, this reader, or this thing?**
Five communities arrived at it independently, in their own vocabularies. People
flagging LinkedIn slop call it generic abstraction. Researchers studying fake
reviews call it emphasizing generic product merits over idiosyncratic
experience. Reddit moderators, cold-email testers and dating-app users all
describe the same failure. It's the rubric behind `specificity-starvation`,
`textureless-anecdote`, `substitutable-reply`, `synthetic-review-shape` and
`mirror-back-research-opener`, and it survives every model improvement, because
the thing it asks for is knowledge the model doesn't have.
## Placeholder copy produces placeholder layout
For web pages the taste findings and the content findings are one family, and
the arrow runs in a direction worth knowing. Underspecified input gets answered
with the statistical average of every page on the internet, and the average page
is mediocre. Ask for a landing page without giving it real content and it invents
placeholder copy, and placeholder copy forces placeholder layout.
That explains the tells together rather than separately. The eyebrow is empty
because the template has an eyebrow slot and nothing true was available to fill
it. The hero has no image because there is no product to photograph. The grid has
three cards because three is what you pick when the number of real things is
unknown. The pull quote quotes nothing because a quote is a provenance relation
and there was no document to pull from.
So the fix usually runs upstream of the design. Supply the real content, and most
of the layout problems stop being layout problems. When you cannot supply it,
that absence is the finding, and a page that admits it beats a page that fills
the gap with a gradient.
## Ask what the reader already knows
Expository writing fails in ways sentence-level review never catches. A blog post
touts improvements over a v2 nobody saw. Internal codenames arrive unglossed. A
term of art is used three times and defined never. Everything is explained at the
same depth, so nothing is signposted as hard.
It is tempting to call these one failure, and I did until the evidence said
otherwise. There are three mechanisms, and they need different fixes.
**Context leakage.** The model writes from its context window, which holds the
previous versions, the internal thread and the repository, and nothing in the
loop marks which parts the reader was present for. It treats availability as
sharedness. This is the curse of knowledge with a machine behind it, and it is
the family `unearned-prior-reference`, `repo-context-leak` and
`definition-after-use` catch. **Fix it with a referent audit:** for every back-
reference, name, and term, check that the thing it points at exists on the page.
**Altitude lock**, and this is the one that breaks the tidy story. When the
audience is stated explicitly in the prompt — so the model has been told what the
reader knows and the context confusion is gone — it still fails. Explanations
generated for different stated audiences come out indistinguishable in reading
level, and match their intended level about half the time against roughly
four-fifths for human-written ones. **Prompting the audience does not fix this.**
The author has to supply a depth plan: which two ideas are hard, which get a
worked example, which get one sentence.
**Template completion.** The structure is announced rather than built, the
promised idea never arrives, the release notes wear an essay's clothes. Form
before content, which is the `shape` family. **Fix it by deleting the scaffold.**
The diagnostic, when you want one question: **could a competent stranger who
arrived from a search result follow this, and where is the first place they would
have to stop and look something up?** That spot is the fix. But note it only
surfaces the first family reliably — altitude lock produces prose a stranger can
follow and learn nothing from, so for that one ask instead which two things here
are hard, and whether the piece spends more time on them than on the easy parts.
## Set severity by reader harm
Do not convert estimated AI prevalence on a platform into editing severity.
A broken exercise, fabricated reference, or inaccessible control matters because
of its effect on the reader. A fashionable font or an unusually tidy paragraph
is at most a review cue. Findings need a location, an observable consequence,
a counterexample, and a repair; uncertain authorship adds none of these.
## Teach a concept before spending it
For lessons, academic books, tutorials, research explanations, and study notes,
load `references/educational-exposition.md` and
`references/education-and-chrome-research.md`. Record the intended reader's prior
knowledge and the next task they should be able to do. Then trace each hard
concept from its definition through a worked instance to an independent attempt.
A definition alone does not make the concept available for later reasoning.
Use the **close-the-page test**: can a learner explain the key step, distinguish
an example from a near miss, and solve a changed case without the answer visible?
Treat that as a proposed assessment until an actual learner attempts it. Readable
prose, a completed annotation map, and a model's self-review do not prove learning.
The route depends on purpose. A reference manual may point to a prerequisite
lesson. An expert chapter may omit elementary practice. A novice lesson normally
needs a complete worked example, reduced guidance, and a later retrieval task.
Keep the hard step visible; remove repetition of what the learner already knows.
## Give each title layer a job
An eyebrow, title, subtitle, and introductory line are four opportunities to say
something useful, not four slots that must be filled. For each, name the distinct
job: locating the reader, naming the subject, narrowing the scope, or supplying a
constraint. Temporarily remove it. If the reader loses neither information nor
orientation, cut it. If useful scope is lost, move that scope into the surviving
heading or opening sentence.
This applies to web heroes, note templates, slide titles, LaTeX chapter openings,
and repeated figure/title/caption stacks. Three layers is a **local review cue**,
not a research-derived limit. Keep useful section labels, venue-required titles,
accessible headings, theorem labels, and captions that explain a figure.
`references/content-chrome.md` contains the decision tree and counterexamples.
```bash
python3 scripts/review_learning_structure.py chapter.md page.html chapter.tex \
--out structure-findings.json
python3 scripts/humanize_review.py chapter.md --findings structure-findings.json \
--out report.html
```
The first script finds structural title-stack candidates, not redundant meaning.
Its optional learning-map mode checks explicit reviewer annotations; it cannot
infer mastery from chapter text. See its module documentation and
`templates/learning-map.json` for the input contract. Replace its illustrative
source path and line numbers with reviewed locations; the validator checks the
annotation structure, not whether those locations establish learning. A model reviewer must still
check teaching quality and the rendered page or PDF.
## Start here: what am I looking at?
Answer one question — what kind of artifact is in front of you — and this
says which files to open. Every reference file then opens with its own decision
diagram for when it applies, and an index of its contents.
```mermaid
flowchart TD
START(["What am I reviewing?"]) --> PROSE["Prose, essay, README,<br/>blog, email, reply"]
START --> WEB["Web page, component,<br/>CSS, JSX"]
START --> CODE["Commit, PR, review,<br/>test, source file"]
START --> DOC["Long document<br/>or slide deck"]
START --> MKT["Landing copy, social post,<br/>listing, resume"]
START --> PAPER["Paper, preprint,<br/>referee report, .tex"]
START --> EDU["Lesson, notes, textbook,<br/>notebook, tutorial"]
EDU --> E1["educational-exposition<br/>education-and-chrome-research"]
E1 --> RUN
START --> FIG["Figure, chart,<br/>plotting script"]
START --> MEDIA["Generated image,<br/>video or audio"]
PROSE --> P1["claudeisms · gptisms-codexisms<br/>other-model-dialects"]
WEB --> W1["web-build-defects FIRST"]
W1 --> W2["then: visual-design · typographic-craft<br/>interaction-and-motion · forms-and-input<br/>accessibility · performance · product-ux-writing<br/>unopened-surfaces · dark-patterns"]
CODE --> C1["engineering-artifact-tells"]
DOC --> D1["document-and-deck-structure<br/>content-chrome"]
MKT --> M1["marketing-and-platform-tells"]
PAPER --> R1["research-papers · latex-source"]
FIG --> F1["scientific-figures"]
MEDIA --> G1["generated-media-tells<br/>provenance first"]
P1 --> RUN["Run the pipeline"]
W2 --> RUN
C1 --> RUN
D1 --> RUN
M1 --> RUN
R1 --> RUN
F1 --> RUN
G1 --> RUN
```
**Open `references/web-build-defects.md` before anything else on a web page.**
Everything in it is reproducible by opening the page, so it carries no fairness
caveat and no claim about authorship — and a page that scrolls sideways at
390px has a bigger problem than reading a bit generated.
## The pipeline
Three layers. The script is layer one; you are layer three.
```mermaid
flowchart TD
A(["Artifact in hand"]) --> B{"Author's prior writing available?"}
B -->|yes| C["Collect 2-3 samples"]
B -->|no| D["Proceed. Every rhythm finding caps at LOW"]
C --> E["LAYER 1 - STRUCTURAL<br/>humanize_review.py FILE --baseline GLOB"]
D --> E
E --> F{"Any finding marked rendered?"}
F -->|yes| G["LAYER 2 - RENDERED<br/>render_check.py FILE --viewports 390,768,1280<br/>--probe-modals --probe-a11y"]
F -->|no| H["LAYER 3 - JUDGE"]
G --> H
H --> I["Read catalog.json. Ask every llm-judge item,<br/>plus every structural item the scripts<br/>do NOT implement - each is marked<br/>'Automated here: no' in its reference file"]
I --> J{"Item marked assistive?"}
J -->|yes| K["A person must hear it. Do NOT write<br/>an assertion and call a pass evidence"]
J -->|no| L["Write findings.json"]
K --> L
L --> M["Merge: humanize_review.py FILE --findings findings.json"]
M --> N{"Fix now?"}
N -->|yes| O["Edit file by file, then re-run to verify"]
N -->|no| P["Deliver the report and stop"]
O --> Q{"Clean?"}
Q -->|no| E
Q -->|yes| P
```
## Process
### Scope the input and find a baseline
Identify the medium and the stakes. A tweet gets the judge pass only; a marketing
site gets both layers plus a render check. Then go looking for the author's own
earlier work, because it turns the weakest half of the catalog into the strongest.
Two or three published pieces are enough. For a repo, the baseline is the repo:
read `git log --oneline -30` and three recently merged PRs before you judge a
commit or a description.
For decks, dump the text and speaker notes first with python-pptx, review that as
prose, then review the visual idiom separately.
### Run the structural layer
```bash
python3 scripts/humanize_review.py FILE... --baseline 'posts/*.md' \
--out report.html --json structural.json
```
It measures countable things only: densities, variances, ratios, codepoints, hex
values, font names, identifier overlap. Thresholds come from
`references/catalog.json`, so the rubric and the code cannot drift apart. Run
`--validate` if you ever doubt that they agree.
### Render the page before you judge it
For anything that ships as a web page, the static pass is half the story. It can
tell you there are no responsive variants; it cannot tell you the pricing table
is 1180px wide inside a 390px screen, or which element is at fault.
```bash
python3 scripts/render_check.py URL_OR_FILE --json render.json
python3 scripts/humanize_review.py page.html --findings render.json --out report.html
```
That script is the one thing in this bundle with a dependency, on Playwright, and
it says so when it is missing rather than failing obscurely. It opens the page at
390, 768 and 1280, and reports horizontal overflow with the offending elements
named, tap targets under 44px, text under 12px, and contrast below AA. Those
findings merge into the same report through the normal `--findings` path.
If you cannot run it, open the page yourself at phone width. Half the findings in
`references/web-build-defects.md` are visible in ten seconds that way.
Two of its flags turn a source-level guess into a verdict, and both cover states
the author's own browser is never in:
```bash
python3 scripts/render_check.py URL --probe-modals --probe-a11y
```
`--probe-modals` opens each plausible trigger and tests the four behaviours that
matter — focus moves in, Escape closes, Tab stays contained, focus returns.
`--probe-a11y` applies the WCAG 1.4.12 text-spacing values and measures what
stops fitting, then switches to forced colors and confirms which elements lost
their only visible boundary. Both matter because a static pass reports success on
exactly these: the Escape handler is present in source whether or not it runs,
and a gradient looks like a boundary right up until the OS reverts it.
### The mechanism of each lane, one line each
Each lane file opens with its own mechanism and a decision diagram for when it
applies, then an index of its contents with severity, family, and whether the
scripts already automate the check. This table is the index; the file is the
argument.
| lane | the mechanism in one sentence |
| --- | --- |
| [web build defects](references/web-build-defects.md) | Reproducible by opening the page. No authorship claim, no fairness caveat — act on this file first. |
| [visual design](references/visual-design-tells.md) | One default is a coincidence; the cluster is the finding. |
| [generated media](references/generated-media-tells.md) | Provenance first, pixels last — the artifact heuristics age in months. |
| [document and deck structure](references/document-and-deck-structure.md) | Assembly survives a full rewrite of the prose, which is why a reader feels it without being able to name it. |
| [marketing and platform](references/marketing-and-platform-tells.md) | The venue has a house form, and the tell is the form arriving complete. |
| [engineering artifacts](references/engineering-artifact-tells.md) | The diff is already visible; a commit that restates it says nothing. |
| [unopened surfaces](references/unopened-surfaces.md) | Not that the output is strange — that a whole class of output was never looked at. |
| [dark patterns](references/dark-patterns-the-model-inherits.md) | The violation is usually not in either component but in the gap between them. |
| [accessibility](references/accessibility-beyond-the-checklist.md) | A machine fully verifies roughly 30% of the criteria; a clean scan is not a clean page. |
| [product UX writing](references/product-ux-writing.md) | The less the system knows about the failure, the warmer the copy gets. |
| [typographic craft](references/typographic-craft-and-tokens.md) | Never a wrong value — one value where a function of context belonged. |
| [performance](references/performance-as-a-design-tell.md) | The generator can see the markup and cannot see the waterfall. |
| [interaction and motion](references/interaction-and-motion.md) | Motion's quality lives in time, and it is emitted as a string nobody watches run. |
| [forms and input](references/forms-and-input.md) | The one surface where the output is the START of the user's work. |
| [research papers](references/research-papers.md) | Prose and checkable commitments arrive at the same confidence, and nothing marks which is which. |
| [LaTeX source](references/latex-source.md) | Never that it failed to compile — it compiled, and nobody opened the PDF. |
| [scientific figures](references/scientific-figures.md) | The render is on the far side of a step the author never watched, so the defect is in the code. |
| [tool fingerprints](references/tool-fingerprints.md) | What made the page. Never who wrote it, never whether it is good. |
### Provenance is a separate question, and a separate mode
```bash
python3 scripts/humanize_review.py page.html --provenance
```
This reports which tool published the page and stops. It writes no report and
carries no severity, and that separation is the point: **a fingerprint tells you
what made the page, never who wrote it and never whether it is any good.**
Three collapses to refuse. Provenance is not authorship — a scaffold
dependency proves where a project started, not who has committed since.
Provenance is not defect — a CDN host is infrastructure, and stripping it
gains nothing. And absence proves nothing, because every badge in that table is
removable by paying, so "no badge" is not evidence of hand-building. Wix and
Squarespace are the sharpest case: their AI flows now seed the ordinary editor,
so AI-built and hand-built output are identical in markup, and "this is
AI-generated because it is on Wix" is false.
What does earn a finding is the overlap set, where a provenance string and a
craft problem are the same bytes: a meta description reading `Generated by v0`,
a share card that is the builder's logo or an auto-screenshot of a preview
build, a favicon that is a build tool's logo, a palette still at a library's
exact default values, a generated headshot on an author byline. The structural
pass reports those as ordinary findings. For every one of them the fix is to do
the work the string points at, never to hide the string — after which the
string disappears on its own. An audit that says "strip the generator meta" is
teaching concealment.
`references/dark-patterns-the-model-inherits.md` is about shapes the commercial
web supplied. A model producing a fake countdown or an accept-only cookie banner
is not choosing to deceive; Princeton found 1,818 dark-pattern instances across
1,254 of 11,000 shopping sites, so these ARE the majority pattern. The output is
still deceptive and sometimes unlawful, which makes it the author's decision, not
yours. The useful thing to notice is that the violation usually lives in the gap
BETWEEN two correct components rather than inside either one: the analytics
snippet is right, the consent banner is right, nothing gates one on the other.
So the checks that matter most there are relational and rendered — load the
page in a fresh profile and watch the network, rather than reading either half.
### Run the judge pass
Read `references/catalog.json`. For every item whose `detection_type` is
`llm-judge`, ask its rubric question of the text. For every item whose
`detection_type` is `structural` but which the script does not implement, ask it
yourself; the script covers the countable subset, not all of them.
Each catalog item carries a `detection_type` — the kind of evidence that
settles it. Three of the four are mechanical: `structural` (decidable from the
source text), `rendered` (needs a real browser at a real viewport), and
`llm-judge` (needs you to read it and answer the item's rubric question).
The fourth is `assistive`, and it means what it says: a headless browser cannot
settle it, because the finding is what a person *hears*.
Do not write an assertion for one and treat a pass as evidence — that is how a
manual check gets laundered into a green tick, which is the failure this whole
skill exists to name. Report those as "not checked" unless someone actually put a
screen reader on, and point them at `references/five-minute-manual-pass.md`.
Read as a hostile editor with taste, not as a checklist executor. Then do one
free-form pass asking what else smells generated, because the catalog is a floor
rather than a ceiling.
Write findings to JSON matching `templates/output-template.md`. Each one needs a
file, a line, an excerpt, an ism that exists in the catalog, a dialect, a
severity, an explanation, and a rewrite. A high-severity finding without a
rewrite is a complaint, not a fix.
### Merge, render, fix, re-run
```bash
python3 scripts/humanize_review.py FILE... --findings judge.json --out report.html
```
Work `templates/rewrite-checklist.md` top-down, highest severity first. Then run
both layers again. Use `--fail-on high` if you want this in CI. Never call
something clean without the re-run.
### Validate an update
[README.md](README.md) contains the runnable entry points;
[CHANGELOG.md](CHANGELOG.md) records revision scope. Read
[the teaching and title examples](examples/before-after-learning-and-chrome.md)
when a rubric needs a concrete repair. Run
`scripts/test_learning_structure.py` for parser and annotation counterexamples,
and `scripts/test_review_regressions.py` for the Unicode, citation-URL, and
eyebrow judgment boundaries. These are offline checks; learner outcomes and
rendered teaching quality still require direct review.
### Maintaining the catalog
`references/catalog.json` is the source of truth for rubrics, thresholds, false
positives, and currency. Edit it, then run
`python3 scripts/regenerate_references.py`, which rebuilds every per-dialect
markdown view and refuses to write if any item would land in no file. The
markdown files are generated, so do not hand-edit them.
Every item carries a `currency` field, because tells decay. Vendors patch the
famous ones: one provider's em-dash rate fell from 10.62 per thousand words to
0.29 across model generations. Items marked `obsolete` stay in the catalog as
cautions rather than tests. The entry on mangled hands in generated images is
already one of them.
## The references
Every generated file opens with a decision diagram for when it applies and an
index of its contents. Load the one the router pointed at; you do not need the
rest.
| Reference | Load when |
|---|---|
| `references/fairness-and-false-positives.md` | Before your first review, and any time someone asks you to judge authorship |
| `references/educational-exposition.md` | Teaching progression, worked examples, retrieval, transfer, research reasoning, and instructional media |
| `references/content-chrome.md` | Repeated title layers and decorative navigation in HTML, Markdown, slides, and LaTeX |
| `references/education-and-chrome-research.md` | Evidence strength, study boundaries, source audit, and evaluation plan |
| `references/review-decisions.md` | Detailed progression and title-layer decision trees, genre exceptions, and medium coverage |
| `references/instructional-media-and-notes.md` | Video, audio, diagrams, notebooks and notes whose task fidelity needs direct inspection |
| `references/source-audit-2026-09-19.json` | Dated inventory of source-note review and explicitly bounded live verification |
| `references/catalog.json` | Always, at judge-pass time; the machine-readable rubric with thresholds and false positives |
| `references/claudeisms.md` | Text suspected from Claude: staccato, dashes, negation frames, escalating compliments · **39 items** |
| `references/gptisms-codexisms.md` | READMEs, code comments, service-voice copy, emoji headers · **35 items** |
| `references/other-model-dialects.md` | Gemini caveat stacks, DeepSeek and Qwen register, Grok voice, register leveling · **14 items** |
| `references/visual-design-tells.md` | Web UI and landing pages: the defaults nobody chose, clustering together · **33 items** |
| `references/document-and-deck-structure.md` | Long documents, READMEs and slide decks — how the thing was assembled, independent of any sentence in it · **28 items** |
| `references/marketing-and-platform-tells.md` | Landing copy, social posts, cold email, listings, résumés — venues with a house form · **30 items** |
| `references/generated-media-tells.md` | A generated image, video or audio clip — provenance first, pixels last · **11 items** |
| `references/fiction-and-narrative-tells.md` | Fiction, narrative, and anything told as a story; the strongest tells in the catalog live here · **6 items** |
| `references/web-build-defects.md` | Any web page, before anything else: responsive, semantics, contrast, scaffold residue · **36 items** |
| `references/engineering-artifact-tells.md` | Commits, PRs, code review, tests, docs, source files · **27 items** |
| `references/unopened-surfaces.md` | Navigation, i18n, docs sites, commerce, email templates, print and PDF — the surfaces a model has no way to open · **36 items** |
| `references/dark-patterns-the-model-inherits.md` | Consent banners, countdowns, activity counters, cancellation and checkout: shapes the commercial web supplied, several of them unlawful · **9 items** |
| `references/tool-fingerprints.md` | Asked which tool built a page, or about to report a generator string as a finding — read the first two items before the other nineteen · **21 items** |
| `references/five-minute-manual-pass.md` | Reviewing any web page for accessibility — eleven keyboard and screen-reader steps, written for someone who has never used one |
| `references/accessibility-beyond-the-checklist.md` | After the manual pass, or when an automated scan came back clean and you need the other 70% · **40 items** |
| `references/product-ux-writing.md` | Any string a logged-in user reads mid-task: errors, empty states, button labels, confirmations, notifications, field hints · **35 items** |
| `references/typographic-craft-and-tokens.md` | Reviewing a stylesheet or a design system: leading, tracking, scale, weight, numerals, colour roles, tokens · **38 items** |
| `references/performance-budget-and-folklore.md` | Before any performance finding — eleven budget numbers a designer can hold, and ten pieces of folklore to drop |
| `references/performance-as-a-design-tell.md` | Images, fonts, video, libraries and the head stack: where the source is right and the delivery is not · **29 items** |
| `references/interaction-and-motion.md` | Any page with animation, scroll effects or hover styling: duration, easing, reduced motion, and what a touchscreen does with a hover state · **26 items** |
| `references/forms-and-input.md` | Any form: the disabled submit, autocomplete tokens, mobile keyboards, error handling and what survives a failed submit · **25 items** |
| `references/research-papers.md` | A manuscript, preprint, abstract or referee report — checks that point outward at the world rather than inward at style · **28 items** |
| `references/latex-source.md` | A `.tex` or `.bib` file — it compiled, and nobody opened the PDF · **30 items** |
| `references/scientific-figures.md` | A plot, chart or panel, and the `.py`/`.R` that drew it — most of it is checkable in the code · **29 items** |
| `references/sources.md` | When you need citations |
| `templates/output-template.md` | Drafting a judge-pass finding or the delivery summary |
| `agents/openai.yaml` | Delegating a review to a subagent |
## Shibboleths
- **A concept mentioned is not a concept taught.** Name the learner action that
makes it safe to use the concept in the next argument.
- **An example is worked when the difficult decision is visible.** Showing the
answer after "simplifying" is not enough when simplification is the lesson.
- **Comfort is not mastery.** Ask for an unaided explanation or changed problem;
do not use a fluent recap or a confidence rating as the exit criterion.
- **A heading names content; it does not justify the page's existence.** Keep
scope and orientation, cut promotional restatement.
- **Count title layers, then judge their jobs.** Never infer redundancy from
capitalization, a missing digit, or a font choice.
- **A caption is part of the argument.** Preserve assumptions, scales, and
uncertainty when trimming decorative text.
- **Evidence about learning is not evidence about AI prevalence.** A classic
pedagogy experiment supports a repair, not a claim that models uniquely fail.
A rate is a cue, not a verdict. Compare punctuation and rhythm against the
author and genre; change them only when the passage benefits.
A pull quote should excerpt or accurately attribute something. Preserve a
clearly labeled original aphorism; repair a falsely attributed quotation.
Defaults can cluster in deliberate design systems. Inspect whether the layout,
type and color serve the content before proposing a change.
A comment that repeats an obvious operation adds little. Explain the constraint
or remove the repetition, regardless of who wrote the code.
Parallel lists can be easier to compare. Break their symmetry only when it
hides real differences in importance or content.
A sentence that announces significance should state the mechanism or evidence.
A number, proper noun or participial clause is neither necessary nor sufficient
for a useful sentence.
## Dos and don'ts
Do quantify before you flag rhythm, and get a baseline whenever you can. Do
preserve meaning exactly when you rewrite, since you are removing an accent
rather than content. Do read two or three adjacent published pieces to learn the
property's voice first. Do run this pipeline over your own draft before you
deliver it. Do say "clean" when it is clean, because a zero-finding report is a
real result.
Don't paraphrase-launder. Running text through one more model adds a second
accent on top of the first, and the fixes have to be surgical instead: edit the
flagged span and leave the rest byte-identical.
Don't build keyword lists to catch phrases. Don't strip personality along with
tropes, since contractions, opinions and irregularity are the goal rather than
the casualties. Don't treat low-severity items as a to-do list; they are taste
calls the author gets to keep. Don't add findings as comments in the source file,
because the HTML report is the deliverable.
Don't strip an agent attribution trailer without checking the repo first. Roughly
forty projects in the public contribution-policy survey require disclosure, so
removing an honest one can violate policy. Detect it and ask.
## Failure modes
**Paraphrase laundering.** You spot it when the fixed text has new tropes the
original lacked. Rewrites should touch only the flagged span.
**Voice flattening.** The post-fix text is clean but beige, because the author's
tics went out with the machine's. Diff against their known-human writing and put
the fingerprints back. A baseline run makes this visible instead of a matter of
taste.
**Checklist myopia.** The judge pass returns only catalog items and no novel
observations. The free-form pass is not optional.
**Flagging quoted evidence.** A document that demonstrates AI-isms will flag the
AI-isms it demonstrates. Wrap those regions in `<!-- humanize:ignore-start -->`
and `<!-- humanize:ignore-end -->`, which is what this bundle's own examples do.
**Stale advice.** Acting on an `obsolete` catalog item produces confident wrong
answers. Check the currency line, especially on image forensics, where hands and
garbled text stopped working as tests around 2025. Lead with provenance instead.
**Forensics drift.** Someone asks whether a student or employee used AI, and the
skill answers. It is out of scope. Detection-for-accusation has miserable
false-positive economics, and the fairness reference exists so you can say why.
Worked examples live alongside those: `examples/before-after-prose.md` is a
launch announcement in machine accent and then edited,
`examples/before-after-landing-page.md` walks the v0 look token by token, and
`examples/sample-report.html` shows what a finished fix plan looks like.
<!-- humanize:ignore-start -->
<!-- BEGIN BUNDLE INDEX (auto: index_references.py) -->
## Skill Bundle Index
*Every file in this skill, and when to open it. Auto-generated; from the repository root run `python3 skills/skill-architect/scripts/index_references.py skills/make_copy_and_media_human --fix`.*
**root**
- [`CHANGELOG.md`](CHANGELOG.md) — Changelog — <!-- humanize:ignore-start A changelog is heading-and-bullet dense by design, which is the documented false positive for heading-spam and bu
- [`README.md`](README.md) — Make Copy and Media Human — Strip the machine accent from copy, web UI, slides, READMEs, commits, PRs, marketing pages, lessons, academic books, notebooks, and generate
**`agents/`**
- [`agents/openai.yaml`](agents/openai.yaml) — openai (data/schema)
**`examples/`**
- [`examples/before-after-landing-page.md`](examples/before-after-landing-page.md) — Before / After — Landing Page (the v0 look, token by token) — Verified by running `scripts/humanize_review.py` against the Before block above.
- [`examples/before-after-learning-and-chrome.md`](examples/before-after-learning-and-chrome.md) — Before / After — Learning, Research, and Title Chrome — These specimens are synthetic teaching data and editorial demonstrations.
- [`examples/before-after-prose.md`](examples/before-after-prose.md) — Before / After — Launch Announcement — The same announcement, machine accent vs.
- [`examples/sample-report.html`](examples/sample-report.html)
**`references/`**
- [`references/accessibility-beyond-the-checklist.md`](references/accessibility-beyond-the-checklist.md) — Accessibility beyond the checklist — Everything in this file is in the part of accessibility that automation cannot reach.
- [`references/catalog.json`](references/catalog.json) — catalog (data/schema)
- [`references/claudeisms.md`](references/claudeisms.md) — Claudeisms — and the generic prose tells Claude amplifies — Tells most associated with Claude-family output, plus the cross-model prose tells that show up strongest in Claude registers.
- [`references/content-chrome.md`](references/content-chrome.md) — Title layers and content chrome — Count visible layers, then ask what each contributes.
- [`references/dark-patterns-the-model-inherits.md`](references/dark-patterns-the-model-inherits.md) — Dark patterns the model inherits — A model that produces a fake countdown or an asymmetric cookie banner is not choosing to deceive.
- [`references/document-and-deck-structure.md`](references/document-and-deck-structure.md) — Document and deck structure — how the thing was assembled — Document-shape tells: how generated long-form docs, READMEs and slide decks are put together, independent of any sentence in them.
- [`references/education-and-chrome-research.md`](references/education-and-chrome-research.md) — Education, semantic progression, and chrome — This memo adds an education and interface-structure lens to the humanization research.
- [`references/educational-exposition.md`](references/educational-exposition.md) — Educational exposition and concept progression — Audit what a reader can do with an idea before later material depends on it.
- [`references/engineering-artifact-tells.md`](references/engineering-artifact-tells.md) — Engineering-artifact tells — commits, PRs, reviews, code, tests, docs — What generated engineering work looks like in the artifacts maintainers actually read.
- [`references/fairness-and-false-positives.md`](references/fairness-and-false-positives.md) — Fairness and false positives — read this before you act on any finding — Hand-written, not generated from the catalog.
- [`references/fiction-and-narrative-tells.md`](references/fiction-and-narrative-tells.md) — Fiction and narrative tells — What generated fiction does at the level of story rather than sentence.
- [`references/five-minute-manual-pass.md`](references/five-minute-manual-pass.md) — The five-minute manual pass — Written for someone who has never used a screen reader.
- [`references/forms-and-input.md`](references/forms-and-input.md) — Forms and input — where the output is the start of the user's work — A form is the one surface where the model's output is the BEGINNING of the user's work rather than the end of it.
- [`references/generated-media-tells.md`](references/generated-media-tells.md) — Generated images, video and audio — provenance first — A separate file because the REVIEW ORDER is different.
- [`references/gptisms-codexisms.md`](references/gptisms-codexisms.md) — GPT-isms and Codexisms — ChatGPT's service voice and README register, and the code-comment tells of Codex/Copilot-shaped generation.
- [`references/instructional-media-and-notes.md`](references/instructional-media-and-notes.md) — Instructional media and source-bound notes — Review generated diagrams, video, audio, notebooks and notes against the task they claim to serve.
- [`references/interaction-and-motion.md`](references/interaction-and-motion.md) — Interaction and motion — the property a generator cannot watch — Motion is the one design property whose entire quality lives in TIME, and a generator emits it as a static string it can never watch run.
- [`references/latex-source.md`](references/latex-source.md) — LaTeX source — it compiled, and nobody opened the PDF — LaTeX source is a program nobody in the loop has run — and the tell is never that it failed to compile.
- [`references/marketing-and-platform-tells.md`](references/marketing-and-platform-tells.md) — Marketing copy and platform posts — written to a template — Landing-page copy, social posts, cold email, listings and résumés.
- [`references/other-model-dialects.md`](references/other-model-dialects.md) — Other model dialects — Gemini, Kimi, DeepSeek, Qwen, Llama, Grok — and cross-model translationese — Distinctive tics per model family, plus the affect and register tells that mark any machine output regardless of vendor.
- [`references/performance-as-a-design-tell.md`](references/performance-as-a-design-tell.md) — Performance and payload as a design tell — **The organising idea.** A generator can see the markup it is writing.
- [`references/performance-budget-and-folklore.md`](references/performance-budget-and-folklore.md) — A performance budget a designer can hold, and the folklore to drop — Two things in one file, because they are the same argument.
- [`references/product-ux-writing.md`](references/product-ux-writing.md) — UX writing inside the product — The strings a logged-in user reads mid-task: errors, empty states, button labels, confirmations, notifications, field hints.
- [`references/research-papers.md`](references/research-papers.md) — Research papers — checks that point outward, at the world — A paper's sentences and its checkable commitments come out of the same machinery at the same confidence, and nothing in the finished artifac
- [`references/review-decisions.md`](references/review-decisions.md) — Review decisions — learning sequence and title chrome — This reference turns two high-confusion findings into editorial decisions.
- [`references/scientific-figures.md`](references/scientific-figures.md) — Scientific figures — checked in the code that drew them — A figure is produced by code, and the render sits on the other side of a step the author never watched.
- [`references/source-audit-2026-09-19.json`](references/source-audit-2026-09-19.json) — source audit 2026 09 19 (data/schema)
- [`references/sources.md`](references/sources.md) — Sources — Published catalogs, stylometry research, and essays the catalog draws on.
- [`references/tool-fingerprints.md`](references/tool-fingerprints.md) — Tool fingerprints — provenance, and the few that are also defects — Read the first two items in this file before the other nineteen, because they are the rules the rest depends on.
- [`references/typographic-craft-and-tokens.md`](references/typographic-craft-and-tokens.md) — Typographic craft and design-system structure — **The governing mechanism, and the thing to say when you report any item here.** A generator emits a stylesheet that is internally consisten
- [`references/unopened-surfaces.md`](references/unopened-surfaces.md) — Unopened surfaces — navigation, i18n, docs, commerce, email, print — Everything in this file is about a surface that was never opened.
- [`references/visual-design-tells.md`](references/visual-design-tells.md) — Visual design tells — the v0/Lovable look — What makes a UI read as generated: the defaults nobody chose, clustering together.
- [`references/web-build-defects.md`](references/web-build-defects.md) — Web build defects — the half you can reproduce — A different KIND of finding from the rest of this catalog.
**`scripts/`**
- [`scripts/humanize_review.py`](scripts/humanize_review.py) — humanize_review.py — flag AI-isms in copy/media and emit a static HTML fix plan.
- [`scripts/regenerate_references.py`](scripts/regenerate_references.py) — Regenerate references/*.md from references/catalog.json. Stdlib only.
- [`scripts/render_check.py`](scripts/render_check.py) — render_check.py — open a page at real viewports and report what breaks.
- [`scripts/review_learning_structure.py`](scripts/review_learning_structure.py) — Offline structure and learning-map review for the humanization skill.
- [`scripts/test_learning_structure.py`](scripts/test_learning_structure.py) — In-memory tests for review_learning_structure.py; no temporary files needed.
- [`scripts/test_review_regressions.py`](scripts/test_review_regressions.py) — Regressions for semantic overreach; no network, model, or temporary files.
**`templates/`**
- [`templates/learning-map.json`](templates/learning-map.json) — learning map (data/schema)
- [`templates/output-template.md`](templates/output-template.md) — Judge-Pass Finding + Delivery Template — Fill this in during step 3 (judge pass) and step 4 (delivery) of the process in `SKILL.md`.
- [`templates/rewrite-checklist.md`](templates/rewrite-checklist.md) — Rewrite Checklist — run after every humanizing pass — Work the report top-down, highest severity first, then verify each line below.
<!-- END BUNDLE INDEX -->
<!-- humanize:ignore-end -->