Skip to content
Back to skills

Maximum Entropy Objective

ASecurity

Compute SAC maximum entropy objective diagnostics for rewards, log probabilities, discounts, and temperature-scaled entropy bonuses.

  • 247 stars
  • 0 votes
  • 0 copies
  • 2 views
  • Added September 9, 2026
testingpythonperformance

Security analysis

A100/100

Scanned September 9, 2026

npx -y skills add VectorSpaceLab/AREX-Skill --skill maximum_entropy_objective --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Maximum Entropy Objective?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for Maximum Entropy Objective
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/vectorspacelab-maximum-entropy-objective/badge)](https://www.skillsdirectory.com/skills/vectorspacelab-maximum-entropy-objective)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

SKILL.md
---
name: maximum_entropy_objective
description: Compute SAC maximum entropy objective diagnostics for rewards, log probabilities, discounts, and temperature-scaled entropy bonuses.
---

# Maximum Entropy Objective

Use this skill when a recovery or implementation must verify the SAC maximum-entropy objective from Haarnoja et al. The skill is appropriate for bounded checks, synthetic replay batches, and cross-checks of entropy bonus signs. Do not use it to claim full MuJoCo performance.

## Inputs
- A reward sequence.
- Matching policy log probabilities.
- Discount factor `gamma`.
- Temperature or reward-scale coefficient `alpha`.

## Outputs
- Discounted reward-only return.
- Discounted soft return using `reward - alpha * log_prob`.
- Per-step entropy bonuses.

## Workflow
1. Validate reward and log-probability lengths.
2. Convert entropy to the sign convention `-log_prob`.
3. Add `alpha * entropy` to reward at each step.
4. Compute discounted totals for reward-only and soft objectives.
5. Use the difference as a mechanism diagnostic, not as a task score.

## Validation
Run `python tests/test_entropy_objective.py` or validate the whole tree with the Distiller skill validator.

## Limitations
This module computes objective diagnostics only. It does not implement replay sampling, neural critics, or environment interaction.

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…