Installs into .claude/skills of the current project.
Are you the author of Context Modulated Emrnn Memory?
Add the live security badge to your README. It updates with every re-scan.
[](https://www.skillsdirectory.com/skills/hiyenwong-context-modulated-emrnn-memory)
---
name: context-modulated-emrnn-memory
description: "建模PFC上下文调制记忆时用。低秩门控RNN+键值EM情境检索。"
metadata:
arxiv_id: "2609.37791"
published: "2026-09-29"
authors: "Hayoung Song, JeongJun Park, Qihong Lu, Giacomo Vedovati, Monica D. Rosenberg, Zachariah M. Reagh, ShiNung Ching"
source: "arXiv cs.AI, cs.NE"
tags: [computational-neuroscience, working-memory, episodic-memory, context-modulation, low-rank-rnn, key-value-memory, pfc, hippocampus, bayesian-inference, naturalistic-fmri]
---
# Context-Modulated EM-RNN: PFC Context Gating of Working & Episodic Memory
**arXiv 2609.37791** (2026-09-29) — Song, Park, Lu, Vedovati, Rosenberg, Reagh & Ching (UT Austin / WashU in St. Louis / U Chicago). An RNN + key-value episodic-memory buffer that *infers* situational context at test time (sticky Bayesian, no labels) and uses it to modulate both working memory (low-rank gating of shared recurrent connectivity) and episodic retrieval (content × context similarity), validated against 33-subject naturalistic fMRI and "aha"-press human retrieval data from the same movie. Code: github.com/hyssong/contextMod; fMRI: OpenNeuro ds005658.
**Activation keywords**: neuroscience, brain network, neural dynamics, computational neuroscience, working memory, episodic memory, context modulation, low-rank RNN, PFC, hippocampus, memory-augmented network, key-value memory, Bayesian context inference, naturalistic fMRI, RSM similarity
## Architecture (one screen)
Inputs: CLIP (50d video) + CLAP (50d audio) per ~4s scene → leaky RNN (h ∈ R^100, λ=0.7) predicts next-scene semantics (10d: characters/place/time). Context π ∈ R^4 (storyline posterior, sums to 1) is a *modulatory* signal, not an input.
1. **Low-rank WM gating** (Eq. 4): shared recurrence W0 element-wise gated by context-weighted unit-rank outer products (m_k, n_k trainable):
`h_t = (1−λ)·h_{t−1} + λ·tanh( (W0 ⊙ Σ_k π_t^(k)·m_k n_k^T)·h_{t−1} + W_v v_t + W_a a_t + b )`
2. **Sticky Bayesian context inference** (test time, labels absent): prior = posterior_{t−1}·T with sticky transition T (γ=0.9 self); likelihood of context k ≈ −β/2·MSE(y_{t+1}, ŷ_{t+1}^(k)), β=10 — simulate a next-scene prediction under each k, better prediction ⇒ higher likelihood; posterior via log-space Bayes + log-sum-exp. π_0 = uniform. Succeeds only because context-specific parameters were trained with π provided.
3. **Context-modulated EM retrieval** (Eqs. 16–18): buffer stores keys `k_t = W_k[v_t,a_t]`, values `v_t = h_t`, AND contexts `π_t` (M=1000 slots). Retrieval weights = `softmax( content_sim ⊙ relu(context_sim), τ=0.1 )` where content_sim = q_t·K^T/√(τD) (D=100), context_sim = π_t·C^T/√(τD) (D=4). Retrieved m_t = V·attn^T; output `ŷ = W_y(h_t·α + m_t·(1−α))`, α=0.5. EM flushed only at test start.
## Key results (numbers to cite)
| Condition | test pred r | ctx infer % | brain-RSM r | human-retrieval r |
|---|---|---|---|---|
| no modulation | 0.414 | – | 0.0020 | 0.265 @200 iter |
| input mod | 0.514 | 58.7 | 0.0039 | – |
| output mod | 0.597 | 75.4 | 0.0053 | – |
| full WM (4 indep. W_h) | 0.555 | 53.2 | 0.0080 | – |
| **low-rank WM** | 0.583 | 60.6 | **0.0087** | 0.207 (worse than none!) |
| **WM+EM (low-rank + ctx retrieval)** | 0.565 | 62.1 | highest (WM+EM > WM > none, t(32)=9.5) | **0.283; human-like 4× faster (r=0.2 @ iter 22 vs 94)** |
brain-RSM = model hidden-state scene-RSM vs fMRI parcel RSMs, avg over 200 Schaefer parcels × 33 subjects, 598 scenes. All modulation-vs-baseline diffs FDR-corrected p<.0001 (paired t(32)>9). Human retrieval = 48×48 event matrix from participants' "aha" insight explanations (>40% reference past events).
## Transferable design patterns
1. **Modulate the middle, not the ends.** Context gating of recurrent dynamics (WM) beats input or output gating for brain alignment. WHERE modulation sits is a first-class design decision — test it as an explicit factor.
2. **Low-rank > full switching.** Gating a shared W0 with context-weighted unit-rank corrections beats per-context independent matrices: straighter state-space trajectories (tortuosity t(19)=5.5, p<.0001), lower-dimensional dynamics (PC1 40.8% vs 34.4%), more orthogonal context endpoints and connectivity (t(19)=5.2). Fewer parameters, better separation.
3. **Explicit context in the memory buffer = division of labor.** Storing π alongside keys/values frees the key-value (attention) system from learning context structure implicitly — human-like retrieval emerges 4× faster. Trade-off: the KV content representation becomes less human-aligned on its own (content-sim ↔ human retrieval: 0.331 no-mod vs 0.171 WM+EM) — context sim carries the structure instead.
4. **Selective retrieval is load-bearing.** Multiplicative content×context gating helps only with small softmax τ (0.1); τ=0.5, τ=1.0, no-softmax progressively degrade to baseline. Human retrieval is sparse/selective by nature.
5. **Context modulation can hurt in isolation.** WM-only modulation drops human-retrieval similarity below baseline (0.207 vs 0.265). Mechanism benefits are conditional on the full system — ablate combinations, not just add-ons.
6. **Modulate readout scores, not addressing.** Low-rank-modulating the key/query *transforms* instead of retrieval scores: r=0.139 vs 0.283 (t(19)=7.5, p<.0001). Gate the similarity computation, not the input pathway to it.
7. **Test-time context inference from prediction quality.** Posterior over discrete contexts = per-context simulated prediction + Gaussian-MSE likelihood + sticky HMM prior (γ=0.9). No context labels needed at deployment, but π must be provided during training (chicken-and-egg).
8. **Validation stack for naturalistic modeling:** (a) model-brain RSM correlation (hidden-state scene RSM vs parcel RSMs), (b) model-human behavior (retrieval-weight matrix vs behavioral retrieval graph), (c) dynamical geometry (forward-simulate π from 0→1: tortuosity / PCA dimensionality / endpoint orthogonality). Reusable triplet for any model-vs-fMRI paper.
## When to use
- Memory-augmented RNNs/agents facing *changing situations* (regimes, tasks, users, sessions): treat inferred context as an explicit low-rank gate on recurrence plus a multiplicative prior on memory retrieval, rather than hoping attention learns it implicitly.
- Naturalistic neuroscience modeling: CLIP/CLAP scene embeddings → scene-by-scene RSM vs Schaefer-parcel BOLD comparison pipeline.
- Any system where context is latent at deployment but observable during training.
## Honest limits
- Brain-RSM correlations are small (r ≈ 0.009); alignment dominated by visual/auditory areas — PFC is implemented algorithmically (Bayes), so no PFC module activity is compared.
- Only 4 hand-annotated contexts (TV storylines); inference ceiling ~62%.
- Human retrieval matrix is sparse and aggregated; events 46–48 excluded in the original dataset.
- Task accuracy alone does not separate conditions much — the case rests on brain/behavioral alignment, not prediction performance.