Skip to content
Back to skills

Td Agent

ASecurity

Pick between Q-learning, SARSA, Expected SARSA for a tabular or small-feature RL task. Use when you need help with td agent.

  • 8 stars
  • 0 votes
  • 0 copies
  • 0 views
  • Added September 8, 2026
developmentgo

Security analysis

A100/100

Scanned September 8, 2026

npx -y skills add anubhavg-icpl/vibe --skill td-agent --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Td Agent?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for Td Agent
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/anubhavg-icpl-td-agent/badge)](https://www.skillsdirectory.com/skills/anubhavg-icpl-td-agent)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
SKILL.md
---
name: td-agent
description: Pick between Q-learning, SARSA, Expected SARSA for a tabular or small-feature RL task. Use when you need help with td agent.
license: CC-BY-NC-SA-4.0
phase: 9
lesson: 4
metadata:
  version: 1.0.0
  tags: [rl, td-learning, q-learning, sarsa]
---

Given a tabular or small-feature environment, output:

1. Algorithm. Q-learning / SARSA / Expected SARSA / n-step variant. One-sentence reason tied to on-policy vs off-policy and variance.
2. Hyperparameters. α, γ, ε, decay schedule.
3. Initialization. Q_0 value (optimistic vs zero) and justification.
4. Convergence diagnostic. Target learning curve, `|Q - Q*|` check if DP is possible.
5. Deployment caveat. How will exploration behave at inference? Is SARSA's conservatism needed?

Refuse to apply tabular TD to state spaces > 10⁶. Refuse to ship a Q-learning agent without a max-bias caveat. Flag any agent trained with ε held at 1.0 throughout (no exploitation phase).

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…