Skip to content
Back to skills

Ppo Gym Wrapper Kaggle Env

ASecurity

Wrap a Kaggle competitive game environment as an OpenAI Gym env with continuous action space for training PPO agents via stable-baselines3

  • 61 stars
  • 0 votes
  • 0 copies
  • 1 view
  • Added September 12, 2026
developmentpythongoapi

Works with

  • api

Security analysis

A100/100

Scanned September 12, 2026

npx -y skills add wenmin-wu/ds-skills --skill ppo-gym-wrapper-kaggle-env --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Ppo Gym Wrapper Kaggle Env?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for Ppo Gym Wrapper Kaggle Env
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/wenmin-wu-ppo-gym-wrapper-kaggle-env/badge)](https://www.skillsdirectory.com/skills/wenmin-wu-ppo-gym-wrapper-kaggle-env)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
SKILL.md
---
name: tabular-ppo-gym-wrapper-kaggle-env
description: Wrap a Kaggle competitive game environment as an OpenAI Gym env with continuous action space for training PPO agents via stable-baselines3
---

# PPO Gym Wrapper for Kaggle Environments

## Overview

Kaggle game AI competitions (Halite, Kore, Lux) use `kaggle_environments` which don't conform to the Gym API. Wrapping them in a `gym.Env` subclass with defined observation and action spaces enables training with stable-baselines3 PPO (or SAC, A2C). The wrapper handles state encoding, action translation, opponent management, and episode termination.

## Quick Start

```python
import gym
from gym import spaces
import numpy as np
from kaggle_environments import make
from stable_baselines3 import PPO
from stable_baselines3.common.monitor import Monitor

class KaggleGymEnv(gym.Env):
    def __init__(self, opponent="random"):
        super().__init__()
        self.env = make("kore_fleets", debug=True)
        self.opponent = opponent
        self.observation_space = spaces.Box(-1, 1, shape=(21*21*4+3,), dtype=np.float32)
        self.action_space = spaces.Box(-1, 1, shape=(3,), dtype=np.float32)

    def reset(self):
        self.trainer = self.env.train([None, self.opponent])
        obs = self.trainer.reset()
        return self._encode(obs)

    def step(self, action):
        game_action = self._decode(action)
        obs, reward, done, info = self.trainer.step(game_action)
        return self._encode(obs), reward, done, info

env = Monitor(KaggleGymEnv())
model = PPO("MlpPolicy", env, verbose=1)
model.learn(total_timesteps=100_000)
```

## Workflow

1. Subclass `gym.Env`, define observation and action spaces
2. In `reset()`, create a trainer via `env.train([None, opponent])` — `None` marks the learning agent
3. In `step()`, decode continuous actions to game actions, call `trainer.step()`
4. Encode raw observations into the observation space format
5. Wrap with `Monitor` for logging, train with PPO or similar algorithm
6. Export the trained policy as a Kaggle submission agent

## Key Decisions

- **Action space**: continuous Box(-1, 1) is simpler than MultiDiscrete; decode to game actions in `step()`
- **Opponent**: start with "random", then self-play or a heuristic agent for curriculum
- **Observation shape**: flatten the grid tensor + append scalar features (turn number, total kore, ship count)
- **Reward**: use shaped rewards (delta board value) rather than sparse win/loss for faster learning
- **Self-play**: alternate the trained agent as opponent every N episodes for robustness

## References

- [Reinforcement Learning baseline in Python](https://www.kaggle.com/code/lesamu/reinforcement-learning-baseline-in-python)

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…