HySTAR - fix the value-decomposition basis (overlapping sparse hypergraph scaffold) while learning adaptive spatiotemporal representations, eliminating structural target drift in cooperative multi-agent RL. Use for credit assignment under partial observability, stable high-order coalition value decomposition, or when dynamic grouping keeps reshuffling your learning targets.
Installs into .claude/skills of the current project.
Are you the author of Anchored Hypergraph Credit Assignment?
Add the live security badge to your README. It updates with every re-scan.
[](https://www.skillsdirectory.com/skills/hiyenwong-anchored-hypergraph-credit-assignment)
---
name: anchored-hypergraph-credit-assignment
description: HySTAR - fix the value-decomposition basis (overlapping sparse hypergraph scaffold) while learning adaptive spatiotemporal representations, eliminating structural target drift in cooperative multi-agent RL. Use for credit assignment under partial observability, stable high-order coalition value decomposition, or when dynamic grouping keeps reshuffling your learning targets.
category: ai_collection
trigger_words: multi-agent RL, credit assignment, structural target drift, hypergraph, MAPPO, value decomposition, cooperative MARL, coalition value, SMAC, Google Research Football, agent death, grouping topology
---
# HySTAR: Anchored Hypergraphs for Stable Credit Assignment
**Source**: arXiv:2609.31531v1 (2026-09-25) — Luo, Zhang, Kuang, Yuan, Zeng et al. (8 authors), cs.LG.
## Problem: Structural Target Drift
Cooperative MARL under partial observability with shared rewards must assign team outcomes to individual agents **and high-order coalitions**. Two failure modes:
- **MAPPO-style critics**: compress joint behavior into one global value → no fine-grained credit.
- **Dynamic grouping critics**: adaptively reconstruct the interaction topology at every step → the *mapping from agents/coalitions to value components keeps changing* → the value-decomposition learning target is a moving object. The paper names this **structural target drift**: representation learning and target assignment are entangled, so neither converges cleanly.
## Core Mechanism: Anchor the Basis, Adapt the Features
HySTAR separates the two roles:
1. **Anchored decomposition scaffold** (fixed): an **overlapping sparse hypergraph** with uniform coverage — the *basis* mapping agents↔coalitions to value components never changes during training.
2. **Adaptive representation** (learned): a **spatiotemporal encoder** represents physical and task-dependent interactions, feeding *features* into the fixed scaffold.
3. **Agent-specific advantages**: combine **temporal relevance** (TD advantages) with **structural relevance** (hypergraph-derived coalition membership) to construct per-agent advantage estimates.
The scaffold being fixed kills structural target drift; the encoder being adaptive preserves expressiveness.
## Recipe
1. Start from MAPPO (shared reward, centralized critic).
2. Define a hypergraph over agents: hyperedges = candidate coalitions (overlapping allowed); ensure uniform coverage of all agents. Freeze this topology for the whole run.
3. Learn a spatiotemporal encoder (e.g., attention over agent observations + positions) that produces interaction-aware embeddings.
4. Value head: decompose global value over hypergraph nodes/edges using the encoder features — the *decomposition structure is fixed, only feature-dependent weights are learned*.
5. Advantage per agent = f(temporal advantage, structural advantage from its hyperedge memberships).
6. Train actors with per-agent advantages (PPO-style clipping as in MAPPO).
## Validation Results
- **SMAC (hardest settings)**: +16.7% relative gain over MAPPO, +15.6% over HYGMA (prior hypergraph method with dynamic topology).
- **Google Research Football**: ranks first on all 6 scenarios.
- **Traffic Junction**: convergence epochs reduced by up to 40.2% vs. MAGIC.
- **MPE**: highest episode rewards.
- Controlled experiments (topology variants, agent-death, neighborhood sizes, parameter sweeps) all support **anchoring the scaffold while adapting representations**.
## Reusable Design Principle
**When a learning target depends on a structure that the learner itself adapts, split the structure into (fixed scaffold) + (adaptive features).** Fixed scaffolds stabilize optimization targets; adaptive features carry the expressiveness. The failure pattern — dynamic grouping → target drift → unstable credit assignment — recurs in: continual learning (task-boundary drift), neural architecture search (supernet reparameterization), and any self-referential learning system.
**Activation**: credit assignment, MARL, hypergraph value decomposition, structural target drift, MAPPO variants, coalition credit, anchored scaffold.