Skip to content
Back to skills

Nqs Mechanistic Interpretability

ASecurity

Apply sparse autoencoders to analyze internal representations of neural quantum states and steer quantum properties.

  • 3 stars
  • 0 votes
  • 0 copies
  • 1 view
  • Added September 11, 2026
code-qualitydebugging

Security analysis

A100/100

Scanned September 11, 2026

npx -y skills add hiyenwong/ai_collection --skill nqs-mechanistic-interpretability --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Nqs Mechanistic Interpretability?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for Nqs Mechanistic Interpretability
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/hiyenwong-nqs-mechanistic-interpretability/badge)](https://www.skillsdirectory.com/skills/hiyenwong-nqs-mechanistic-interpretability)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
SKILL.md
---
name: nqs-mechanistic-interpretability
description: Apply sparse autoencoders to analyze internal representations of neural quantum states and steer quantum properties.
trigger_keywords: ["nqs interpretability", "neural quantum states analysis", "sparse autoencoder quantum", "feature steering NQS", "mechanistic interpretability quantum"]
---

# NQS Mechanistic Interpretability

## Description

Methodology from arXiv:2607.01336 that applies sparse autoencoders (SAEs) to analyze the internal activations of Neural Quantum States. Despite being trained only on variational objectives, NQS learn interpretable physical concepts including spin correlations, symmetries, and topological order. Causal feature steering enables controlled manipulation of learned quantum properties.

## Core Methodology

1. **Feature Extraction**: Train sparse autoencoders on NQS residual stream activations to extract interpretable features
2. **Concept Identification**: Map extracted features to physical concepts (spin correlations, symmetries, topological order)
3. **Causal Steering**: Intervene on specific feature dimensions to manipulate quantum properties controllably
4. **Physical Grounding**: Verify that learned features correspond to actual physical observables

## Key Patterns

- **Residual Stream Analysis**: Extract features from intermediate layers of the NQS architecture
- **Sparse Decomposition**: Use SAEs to decompose dense activations into sparse, interpretable feature vectors
- **Feature-Concept Mapping**: Correlate individual SAE features with known physical quantities (magnetization, correlation functions, topological invariants)
- **Causal Intervention**: Zero-out, amplify, or swap specific features to test their causal role in quantum property prediction

## Applications

- Understanding what NQS actually learn about quantum systems
- Debugging and improving variational ansätze design
- Discovering emergent physical concepts from neural representations
- Guiding architecture design based on interpretability insights

## Activation

Use when: analyzing what neural quantum states learn, applying mechanistic interpretability to physics models, extracting physical concepts from neural activations, steering quantum model behavior.

**Keywords**: sparse autoencoders, mechanistic interpretability, neural quantum states, feature steering, physical concepts, residual stream, topological order, spin correlations

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…