Skip to content
Back to skills

10 Phack Polyglot

ASecurity

Run the instrumented specification search in the user's own statistical language — Stata (reghdfe, ivreghdfe, rdrobust, did2s), R (fixest, rdrobust, did2s), Python (statsmodels, linearmodels) or StatsPAI — and bring the results back into the audit, null calibration and honest report. Exports the enumerated grid as a language-neutral specs table plus a generated runner, ingests the runner's ledger, replays the null draws in that language, and reports cross-language parity. Use when an analysis...

  • 3,775 stars
  • 0 votes
  • 0 copies
  • 0 views
  • Added September 27, 2026
researchpythonbashaws

Works with

  • cli

Security analysis

A100/100

Scanned September 27, 2026

npx -y skills add brycewang-stanford/Auto-Empirical-Research-Skills --skill 10-phack-polyglot --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of 10 Phack Polyglot?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for 10 Phack Polyglot
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/brycewang-stanford-10-phack-polyglot/badge)](https://www.skillsdirectory.com/skills/brycewang-stanford-10-phack-polyglot)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
SKILL.md
---
name: phack-polyglot
description: Run the instrumented specification search in the user's own statistical language — Stata (reghdfe, ivreghdfe, rdrobust, did2s), R (fixest, rdrobust, did2s), Python (statsmodels, linearmodels) or StatsPAI — and bring the results back into the audit, null calibration and honest report. Exports the enumerated grid as a language-neutral specs table plus a generated runner, ingests the runner's ledger, replays the null draws in that language, and reports cross-language parity. Use when an analysis or an agent evaluation must happen in Stata, R or StatsPAI rather than the Python engine, when a p-hacked result was produced in one of those languages and needs auditing on its own footing, when reading Stata / R / StatsPAI code for search and disclosure signals, or when checking that the engine's numbers agree with a reference implementation.
---

# One grid, four languages

## The idea: the grid is the contract

An agent that p-hacks in Stata reports Stata's p-values, with Stata's
small-sample conventions; the honest counterpart has to be computed on the
same footing, not on a Python re-implementation. So the engine does the one
thing that must be identical everywhere — enumerate the specification
universe from the design card, resolve every bandwidth to a number, draw the
null permutations — and writes it out as `specs.csv` and `null_columns.csv`.
A generated runner in the target language estimates every row with that
language's own commands and writes back `ledger_raw.csv` in the schema the
audit already reads. Keys (sha1 of the specification label) are identical
across languages, so a Stata ledger, an R ledger and the Python ledger of the
same card line up row for row.

```bash
python scripts/phack_cli.py export DATA CARD --lang stata --out run_stata/ \
    --direction + --null-draws 200 --null-scheme cluster_permute
cd run_stata && stata-mp -b do run_specs.do          # Rscript run_specs.R | python run_specs.py | python run_specs_statspai.py
cd .. && python scripts/phack_cli.py ingest run_stata/ --parity
```

`ingest` writes `ledger.csv` (full schema: axes, `spec_json`, one-sided p,
pathology flags), `audit.json`, `manifest.json`, `report.md`,
`spec_curve.png`, the null arrays, and with `--parity` a `parity.json`
comparing that language with the Python engine on the same specifications.

## What each runner does

| | Stata | R | Python | StatsPAI |
|---|---|---|---|---|
| OLS / DiD-TWFE, FE, weights | `regress` / `reghdfe … [aw=]`, `vce(robust\|hc2\|hc3\|cluster\|cluster a b)` | `fixest::feols`, `vcov="hetero"` / `~cl` / `~a+b`, `weights=~w` | `statsmodels` OLS/WLS with dummies, `cov_type` | `hdfe_ols` / `regress` |
| transforms, discretisers, outlier rules (outcome / treatment / residual), imputation, windows, lags, comparison groups | all | all | all | all |
| did2s | `did2s` | `did2s::did2s` | — | `did_2stage` (no weights) |
| stacked | — | — | — | `stacked_did` (no weights) |
| event study (window, reference period, `avg_post` / `lag k` / `avg_pre`) | relative-time dummies + `lincom` | `feols(i(rel, ref))` + linear combination | — | `event_study` (ATT = avg_post; `avg_pre` SE ignores covariances) |
| RDD (h, kernel, p, donut, controls, cluster) | `rdrobust` → `e(tau_cl)`, `e(tau_bc)`, `e(se_tau_cl)`, `e(se_tau_rb)` | `rdrobust::rdrobust` rows 1–3 | local polynomial by hand | `rdrobust` diagnostics `conventional` / `robust` |
| inference modes conventional / bias-corrected / robust | all three | all three | all three | all three (bias-corrected assembled from the two rows) |
| IV 2SLS | `ivreghdfe` / `ivreg2`, F = `e(widstat)` | `feols(… \| d ~ z)`, `fitstat("ivf")` | `linearmodels.IV2SLS` | `ivreg`; FE need `statspai[fixest]` |
| LIML | `liml` option | — | `IVLIML` | `liml` / `iv(method="liml")` |
| null replay | yes | yes | yes | yes |

"—" is recorded per row as `status = unsupported: …`, never silently
skipped; the audit counts them (`n_unsupported`). Rows the language itself
fails on (Stata's two-way cluster variance not PSD, StatsPAI's `stacked_did`
with controls) come back as `error: …` and stay in the ledger.

## Parity, measured

`phack ingest --parity` compares each language with the engine on the same
rows. Numbers from the shipped null datasets (thinned grids; see
`references/language-map.md` for the full table):

- **Coefficients agree to numerical precision** wherever the estimator is
  the same object: OLS / TWFE / IV in every language (the largest gap, under
  0.01, comes from Stata's percentile convention in an IQR-trimmed sample),
  and the `rdrobust` conventional and robust rows in Stata, R and StatsPAI
  to 0.002.
- **Standard errors differ by convention**, not by mistake: median relative
  gap under 1.5% for OLS / IV; up to 5% for clustered TWFE (reghdfe and
  fixest do not count fixed effects nested in the cluster in the degrees of
  freedom, the engine does); up to a third for `rdrobust` (its variance estimator
  differs from the engine's kernel-weighted sandwich, most for the
  bias-corrected row); 25–30% for `did2s`
  (Stata / R / StatsPAI correct the second stage for first-stage sampling
  error, the engine's stage-2 SE does not) and for event studies clustered
  on eight regions (t(7) versus normal reference).
- **Stata refuses the non-PSD corner** the engine flags: `reghdfe` with
  two-way clustering reports a missing standard error where
  `flag_nonpsd_vcov` fires. Same pathology, two honest responses.
- **StatsPAI has no under-covering RDD row by construction**: its `rdrobust`
  reports conventional and robust only, so strategy 23 (bias-corrected point
  estimate with the conventional SE) has to be assembled by hand — which the
  runner does, and flags.

Significance agrees on 96–100% of rows across languages; the rows that
disagree are the ones where the SE convention straddles 0.05, which is
itself a searchable choice (strategy 9: alternative tests / software).

## Reading Stata, R and StatsPAI code for search signals

`score.scan_code` recognises the idioms each language uses to walk a grid and
to pick from it:

| language | search signals | disclosure signals |
|---|---|---|
| Stata | `foreach` / `forvalues` / `levelsof` wrapping `reg`, `reghdfe`, `rdrobust`, `ivreg2`; `if r(p) < .05`; `sort pval`, `keep if p<`; `abs(_b[]/_se[]) >`; `estimates store` per iteration; rdrobust `h()` loops | `rwolf`, `wyoung`, `mhtexp`, `qqvalue` |
| R | `expand.grid` / `crossing` over bandwidths, kernels, controls; `map` / `lapply` over `feols` / `rdrobust` / `lm`; `filter(p.value <)`, `arrange(p.value)`, `slice_min`, `which.min`; fixest `csw` / `sw` | `p.adjust`, `specr`, `multiverse`, `rwolf` |
| StatsPAI | estimator calls inside a loop over `h=` / `kernel=` / `vcov=` / `ref_period=`; `min(fits, key=lambda r: r.pvalue)` | `spec_curve`, `romano_wolf`, `adjust_pvalues`, `honest_did` |
| any | | `phack search` / `ledger.csv` / null calibration |

The scan is a screen, not a verdict; read what it flags.

## Design cards from the other side

A Stata or R user does not need the Python engine to *build* a card: the
card is JSON and `phack size CARD` is the only Python step before `export`.
For an analysis that already exists in Stata or R, the honest path is: write
the card that contains the specification actually reported plus the
alternatives that were (or could have been) tried, `export` it, run the
runner, `ingest --parity`. The `nearest_significant` distance and
`axis_influence` then describe the reported result in the language it was
produced in.

## Limitations

- The runners reproduce the engine's *grid semantics* (which rows, which
  transformations, which samples); they do not reproduce its numerical
  conventions, and are not meant to. Parity is reported, not enforced.
- Percentile-based rules (winsorising, IQR and percentile trims) use each
  language's default quantile definition; a handful of rows near a cutoff
  can differ by one observation.
- Null replay in Stata re-runs the full grid B times inside Stata; budget
  for it (`--null-max-specs` before `export` via `--max-specs`).
- `rdrobust`'s `bwselect` is never used: bandwidths are resolved in Python
  (rule-of-thumb or Imbens–Kalyanaraman) and passed as `h()` so every
  language walks the same bandwidth grid.

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…