Skip to content
Back to skills

Brain Alignment Causal Dissociation Attention Heads

ASecurity

检验脑对齐头是否因果重要时用。对齐与计算解离。

  • 3 stars
  • 0 votes
  • 0 copies
  • 0 views
  • Added October 3, 2026
researchperformance

Security analysis

A100/100

Scanned October 3, 2026

npx -y skills add hiyenwong/ai_collection --skill brain-alignment-causal-dissociation-attention-heads --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Brain Alignment Causal Dissociation Attention Heads?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for Brain Alignment Causal Dissociation Attention Heads
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/hiyenwong-brain-alignment-causal-dissociation-attention-head/badge)](https://www.skillsdirectory.com/skills/hiyenwong-brain-alignment-causal-dissociation-attention-head)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
SKILL.md
---
name: brain-alignment-causal-dissociation-attention-heads
description: "检验脑对齐头是否因果重要时用。对齐与计算解离。"
metadata:
  arxiv_id: "2609.37991"
  published: "2026-09-29"
  authors: "Christopher Pinier, Gustaw Opiełka, Hannes Rosenbusch, Taylor Webb, Michael D. Nunez, Claire E. Stevenson"
  source: "arXiv cs.AI, q-bio.NC (UvA + Princeton)"
  tags: [brain-ai-alignment, interpretability, attention-heads, causal-ablation, eeg-frp, function-vectors, concept-vectors]
---

# Brain Alignment vs. Causal Importance in LLM Attention Heads

**arXiv 2609.37991** (2026-09-29) — Pinier et al. (University of Amsterdam + Princeton). Direct causal test of the brain-AI alignment inference: does representational similarity to human EEG identify the units a model actually uses to solve the task?

**Activation keywords**: neuroscience, brain network, neural dynamics, computational neuroscience, brain-AI alignment, EEG, attention heads, causal ablation, function vectors, concept vectors, interpretability, RSA

## Core Finding

**Alignment and causation dissociate.** Across 17 LLMs (3B–72B; Llama/Qwen/Phi-4/DeepSeek-R1-Distill, base + instruct variants):
- Brain-aligned heads contribute *something* to performance, but removing them is substantially less disruptive than removing heads selected by attribution patching (FV heads): peak excess damage 12.8 pp @ 24.5% removed vs. FV's 42.6 pp @ 12.5% removed — less than 1/3 the damage despite removing ~2× more heads.
- Brain score ↔ patching score: near-zero Spearman ρ in every model (mean .017, range [-.017, .077])
- Brain score ↔ concept score (CV): positive in Llama (.100–.325) and DeepSeek (.269), mixed in Qwen base, **negative in all instruction-tuned Qwen** (-.205 to -.037), ~zero in Phi-4. Not a universal abstraction signal.

**Verdict**: "Brain alignment thus captures how the model reads the stimulus, and only faintly captures how it represents the pattern and solves the task."

## Method (direct causal test protocol)

1. **Task**: abstract pattern completion AAABAAA→B (8 patterns; icons for humans, words/symbols for models; 3-shot raw prompts)
2. **Three head scores** (computed without/with human data):
   - **Brain score**: Spearman correlation between head's pattern-level RDM (8×8, head output averaged within each pattern) and human frontal fixation-related potential (FRP) RDM
   - **Patching score (FV)**: average indirect effect — patch head's clean-prompt output into corrupted prompt (first block of every demo + query re-drawn at random); gain in clean-answer probability. Attribution patching ≈ full patching (r=.983, 95% top-20 overlap)
   - **Concept score (CV)**: RSA of head outputs over 1,200 items varying pattern×alphabet (German/English/symbols)×response format, against same-pattern design; same-alphabet as control
3. **Cumulative zero-ablation** in each ranking, vs. size-matched random controls (5×)
4. **Attention profile clustering**: 8 patterns × 7 positions attention matrix per head; cluster top-20 brain heads; build frozen templates; carrier = loading ≥ .50; independent clustering as check
5. **Gaze comparison**: human fixation-duration map on same 8×7 grid, cell-wise correlation with model attention

## Key Results

1. **Two recurring attention families among brain-aligned heads**:
   - **Novelty heads**: attend to the *distinctive* element (the lone B). 44% of all heads; correlates with where humans look in all 17 models (mean ρ=.373). **But removal is *less damaging than random ablation*** (never exceeded random by >0.3 pp; model keeps 63.5% accuracy after removing 15.6% of heads vs. random median <50%).
   - **Repetition heads**: attend to *repeated* positions (second A). 17.5% of all heads, overrepresented in top-20 brain heads (22.6%); correlate with concept scores (ρ=.45) and co-occur in layers 16–35; removal costs ~5.5–11.2 pp (more than random) **but leaves the pattern representation intact** (concept RSA preserved; 100% pattern decodability from later concept heads).
2. **Novelty heads are potentially a "spandrel"**: softmax forces every head to distribute attention somewhere; a recognizable profile can coexist with negligible task contribution. Additional 8-task battery + Pythia-1.4B 19-checkpoint training sweep confirm: attention profile strengthens over training while causal contribution stays below threshold.
3. **Gaze confound honesty**: in 5/8 patterns the only unrepeated symbol IS the answer — gaze/attention convergence concentrated there (.47 vs .12 in other patterns); cannot separate "attend because answer" from "attend because rare". But ablations show the model can attend to it without using it. FV heads also attend to same symbol (.28) — attention location alone does not separate idle from working heads.

## Pitfalls for Brain-AI Alignment Research

- **Never infer computation from alignment alone**: similarity to human neural data does not reliably identify performance-critical components — always pair alignment maps with causal interventions (patching/ablation) on the same task
- **RDM correlation ≠ causal role**: representational correspondence can reflect shared *stimulus reading* (salience/rarity detection) rather than shared *solution computation*
- Instruction-tuned models can invert the alignment-concept relationship (Qwen instruct: negative) — alignment results may not transfer across training variants of the same family
- Group-average RDM with unequal participant data contribution + only 28 dependent pattern comparisons = weak statistical target; use held-out participants and multiple RDM constructions
- Zero-ablation effects can be masked by redundant heads — layer-matched controls and patching preferred for weak claims
- Ablation must be measured against size-matched random baselines, not absolute accuracy drop

## Applications

- **Audit protocol for alignment claims**: (1) compute alignment scores, (2) compute causal scores independently, (3) cross-correlate, (4) ablate in both rankings vs. random. Report dissociation if present — an honest negative result is publishable and informative
- Brain-aligned-but-idle heads (novelty family) are a candidate "neural listening interface": they read stimuli the way humans do without perturbing computation — potentially safe targets for brain-model comparison or steering that don't damage capability
- Cross-model universality check: do attention profiles recur across 17 models? Template matching + independent clustering as consistency check
- Training-dynamics angle: track fixed head set across Pythia checkpoints to separate "profile emergence" from "causal role emergence"

## Related Skills
- `heterogeneous-neural-predictivity-lm` (evaluating LM neural predictivity)
- `llm-brain-alignment-training-data` (alignment drivers)
- `mllm-brain-alignment-task-probing` (task-conditioned probing)
- `representation-steering` (activation interventions)

## Source
- arXiv: 2609.37991 — EEG + eye-tracking data from Pinier et al. (2025) 400-trial experiment; code/data availability per paper appendix

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…