Installs into .claude/skills of the current project.
Are you the author of Brain Alignment Causal Dissociation Attention Heads?
Add the live security badge to your README. It updates with every re-scan.
[](https://www.skillsdirectory.com/skills/hiyenwong-brain-alignment-causal-dissociation-attention-head)
---
name: brain-alignment-causal-dissociation-attention-heads
description: "检验脑对齐头是否因果重要时用。对齐与计算解离。"
metadata:
arxiv_id: "2609.37991"
published: "2026-09-29"
authors: "Christopher Pinier, Gustaw Opiełka, Hannes Rosenbusch, Taylor Webb, Michael D. Nunez, Claire E. Stevenson"
source: "arXiv cs.AI, q-bio.NC (UvA + Princeton)"
tags: [brain-ai-alignment, interpretability, attention-heads, causal-ablation, eeg-frp, function-vectors, concept-vectors]
---
# Brain Alignment vs. Causal Importance in LLM Attention Heads
**arXiv 2609.37991** (2026-09-29) — Pinier et al. (University of Amsterdam + Princeton). Direct causal test of the brain-AI alignment inference: does representational similarity to human EEG identify the units a model actually uses to solve the task?
**Activation keywords**: neuroscience, brain network, neural dynamics, computational neuroscience, brain-AI alignment, EEG, attention heads, causal ablation, function vectors, concept vectors, interpretability, RSA
## Core Finding
**Alignment and causation dissociate.** Across 17 LLMs (3B–72B; Llama/Qwen/Phi-4/DeepSeek-R1-Distill, base + instruct variants):
- Brain-aligned heads contribute *something* to performance, but removing them is substantially less disruptive than removing heads selected by attribution patching (FV heads): peak excess damage 12.8 pp @ 24.5% removed vs. FV's 42.6 pp @ 12.5% removed — less than 1/3 the damage despite removing ~2× more heads.
- Brain score ↔ patching score: near-zero Spearman ρ in every model (mean .017, range [-.017, .077])
- Brain score ↔ concept score (CV): positive in Llama (.100–.325) and DeepSeek (.269), mixed in Qwen base, **negative in all instruction-tuned Qwen** (-.205 to -.037), ~zero in Phi-4. Not a universal abstraction signal.
**Verdict**: "Brain alignment thus captures how the model reads the stimulus, and only faintly captures how it represents the pattern and solves the task."
## Method (direct causal test protocol)
1. **Task**: abstract pattern completion AAABAAA→B (8 patterns; icons for humans, words/symbols for models; 3-shot raw prompts)
2. **Three head scores** (computed without/with human data):
- **Brain score**: Spearman correlation between head's pattern-level RDM (8×8, head output averaged within each pattern) and human frontal fixation-related potential (FRP) RDM
- **Patching score (FV)**: average indirect effect — patch head's clean-prompt output into corrupted prompt (first block of every demo + query re-drawn at random); gain in clean-answer probability. Attribution patching ≈ full patching (r=.983, 95% top-20 overlap)
- **Concept score (CV)**: RSA of head outputs over 1,200 items varying pattern×alphabet (German/English/symbols)×response format, against same-pattern design; same-alphabet as control
3. **Cumulative zero-ablation** in each ranking, vs. size-matched random controls (5×)
4. **Attention profile clustering**: 8 patterns × 7 positions attention matrix per head; cluster top-20 brain heads; build frozen templates; carrier = loading ≥ .50; independent clustering as check
5. **Gaze comparison**: human fixation-duration map on same 8×7 grid, cell-wise correlation with model attention
## Key Results
1. **Two recurring attention families among brain-aligned heads**:
- **Novelty heads**: attend to the *distinctive* element (the lone B). 44% of all heads; correlates with where humans look in all 17 models (mean ρ=.373). **But removal is *less damaging than random ablation*** (never exceeded random by >0.3 pp; model keeps 63.5% accuracy after removing 15.6% of heads vs. random median <50%).
- **Repetition heads**: attend to *repeated* positions (second A). 17.5% of all heads, overrepresented in top-20 brain heads (22.6%); correlate with concept scores (ρ=.45) and co-occur in layers 16–35; removal costs ~5.5–11.2 pp (more than random) **but leaves the pattern representation intact** (concept RSA preserved; 100% pattern decodability from later concept heads).
2. **Novelty heads are potentially a "spandrel"**: softmax forces every head to distribute attention somewhere; a recognizable profile can coexist with negligible task contribution. Additional 8-task battery + Pythia-1.4B 19-checkpoint training sweep confirm: attention profile strengthens over training while causal contribution stays below threshold.
3. **Gaze confound honesty**: in 5/8 patterns the only unrepeated symbol IS the answer — gaze/attention convergence concentrated there (.47 vs .12 in other patterns); cannot separate "attend because answer" from "attend because rare". But ablations show the model can attend to it without using it. FV heads also attend to same symbol (.28) — attention location alone does not separate idle from working heads.
## Pitfalls for Brain-AI Alignment Research
- **Never infer computation from alignment alone**: similarity to human neural data does not reliably identify performance-critical components — always pair alignment maps with causal interventions (patching/ablation) on the same task
- **RDM correlation ≠ causal role**: representational correspondence can reflect shared *stimulus reading* (salience/rarity detection) rather than shared *solution computation*
- Instruction-tuned models can invert the alignment-concept relationship (Qwen instruct: negative) — alignment results may not transfer across training variants of the same family
- Group-average RDM with unequal participant data contribution + only 28 dependent pattern comparisons = weak statistical target; use held-out participants and multiple RDM constructions
- Zero-ablation effects can be masked by redundant heads — layer-matched controls and patching preferred for weak claims
- Ablation must be measured against size-matched random baselines, not absolute accuracy drop
## Applications
- **Audit protocol for alignment claims**: (1) compute alignment scores, (2) compute causal scores independently, (3) cross-correlate, (4) ablate in both rankings vs. random. Report dissociation if present — an honest negative result is publishable and informative
- Brain-aligned-but-idle heads (novelty family) are a candidate "neural listening interface": they read stimuli the way humans do without perturbing computation — potentially safe targets for brain-model comparison or steering that don't damage capability
- Cross-model universality check: do attention profiles recur across 17 models? Template matching + independent clustering as consistency check
- Training-dynamics angle: track fixed head set across Pythia checkpoints to separate "profile emergence" from "causal role emergence"
## Related Skills
- `heterogeneous-neural-predictivity-lm` (evaluating LM neural predictivity)
- `llm-brain-alignment-training-data` (alignment drivers)
- `mllm-brain-alignment-task-probing` (task-conditioned probing)
- `representation-steering` (activation interventions)
## Source
- arXiv: 2609.37991 — EEG + eye-tracking data from Pinier et al. (2025) 400-trial experiment; code/data availability per paper appendix