Skip to content
Back to skills

oh-my-coding-maas-gateway

FSecurity

Operational companion for the oh-my-coding-maas-gateway LiteLLM proxy stack. Provides context and commands for health checks, validation, upgrades, key/model management, debug routing, metrics, and recovery.

  • 3 stars
  • 0 votes
  • 0 copies
  • 0 views
  • Added September 19, 2026
ai-agentspythongobashsqldockergitapidatabase

Works with

  • claude code
  • cli
  • api

Security analysis

F17/100
  • criticalPipes output to a shell interpreter
  • mediumUses curl or wget to download content
  • criticalModifies startup scripts or system services for persistence
  • criticalExfiltrates credentials via HTTP — exact pattern from Snyk ToxicSkills study
  • criticalSends environment variables or credentials to an external URL
  • criticalDownloads and executes remote scripts — classic supply chain attack
  • mediumInstalls packages at runtime which could introduce malicious dependencies

Pro scans all 20 files and shows the line behind each finding

Scanned October 4, 2026

npx -y skills add wallacelw/oh-my-coding-maas-gateway --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of oh-my-coding-maas-gateway?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for oh-my-coding-maas-gateway
[![Security: F — Skills Directory](https://www.skillsdirectory.com/api/skills/wallacelw-oh-my-coding-maas-gateway/badge)](https://www.skillsdirectory.com/skills/wallacelw-oh-my-coding-maas-gateway)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
SKILL.md
---
name: oh-my-coding-maas-gateway
description: Operational companion for the oh-my-coding-maas-gateway LiteLLM proxy stack. Provides context and commands for health checks, validation, upgrades, key/model management, debug routing, metrics, and recovery.
---

# oh-my-coding-maas-gateway — Operational Companion

Operational companion for a self-hosted LiteLLM proxy routing Huawei MaaS
models to opencode, Codex CLI, Claude Code CLI, and Pi agent with virtual
keys, multi-key load balancing, and Prometheus + Grafana observability.

## When Invoked

Present this menu. Default (just press Enter) loads context without action:

```
What would you like to do?

  1) Health check      — quick status of all services
  2) Run validation    — full end-to-end validation
  3) Upgrade           — check for and apply updates
  4) Uninstall         — remove all or part of the gateway
  5) Install skill     — install a new skill into all agents
  6) Mint new keys     — add MaaS or virtual keys

  Choice [1-6] or Enter for context only:
```

After completing an action, ask if they need anything else. For anything
not in the menu (model management, debug routing, metrics), respond using
the reference sections below.

## Project Location

The gateway is at `/home/oh-my-coding-maas-gateway` (the default install
location). `cd` there first:

```bash
cd /home/oh-my-coding-maas-gateway
```

If installed elsewhere, locate the repo by finding the directory that
contains `scripts/04_validate.sh`.

## Available Scripts

| Script | Purpose | Key flags |
|--------|---------|-----------|
| `scripts/bootstrap.sh` | Install or upgrade the entire stack | `--tool=`, `--virtual-key=`, `--api-key=`, `-y`/`--yes`, `--dry-run`, `--no-skill` |
| `scripts/update.sh` | Check and update individual components (tools + infrastructure); shows project + component versions | `--check`, `--all`, `--dry-run` |
| `scripts/04_validate.sh` | End-to-end validation (run anytime) | `--litellm-only`, `--opencode-only`, `--codex-only`, `--claude-code-only`, `--pi-only`, `--skip-opencode`, `--skip-codex`, `--skip-claude-code`, `--skip-pi`, `--dry-run` |
| `scripts/05_skill.sh` | Install THIS companion skill into agents | `--yes`, `--dry-run`, `--no-skill` |
| `scripts/06_backup.sh` | Dump/restore the LiteLLM PostgreSQL DB (spend history, virtual keys, budgets) | `--restore FILE`, `--keep N`, `--dry-run`, `--yes` |
| `scripts/07_dashboard_shots.sh` | Capture dashboard screenshots for visual verification | `--out=`, `--dry-run` |
| `scripts/install-skill.sh` | Install ANY skill into all detected agents | `--name=`, `--source=`, `--dry-run` |
| `scripts/uninstall.sh` | Remove all or part of the gateway | `--tool=`, `--docker`, `--repo`, `--all`, `--dry-run`, `--yes` |
| `scripts/02_litellm.sh` | Regenerate LiteLLM config + restart (after editing `.env` or `models.sh`) | `--routing-strategy=`, `--dry-run` |
| `scripts/01_env.sh` | Regenerate `.env` (after key changes) | `--force` |

## Services

| Service | URL | Auth |
|---------|-----|------|
| LiteLLM Proxy | `http://127.0.0.1:4000` | Virtual key |
| LiteLLM Admin UI | `http://127.0.0.1:4000/ui` | Master key (from `.env`) |
| Grafana Dashboard | `http://127.0.0.1:3000` | admin password (from .env) |
| Prometheus | `http://127.0.0.1:9090` | None |

---

## Option 1: Health Check

```bash
docker compose ps
curl -sf http://127.0.0.1:4000/health/liveliness && echo "LiteLLM: healthy" || echo "LiteLLM: unhealthy"
curl -sf http://127.0.0.1:3000/api/health && echo "Grafana: healthy" || echo "Grafana: unhealthy"
```

Report: how many containers are running, which are healthy, any issues.
If problems found, suggest fixes from the Recovery table below.

## Option 2: Run Validation

```bash
./scripts/04_validate.sh
```

If failures occur, match them against the Recovery table, suggest the fix,
and offer to run it. WARN messages are advisory only.

## Option 3: Upgrade

```bash
curl -fsSL https://raw.githubusercontent.com/wallacelw/oh-my-coding-maas-gateway/main/scripts/bootstrap.sh | bash
```

After upgrade, remind user to restart any running coding tools.
If Grafana looks stale: `docker compose restart grafana`.

### Update coding tools only

To check and update individual components without re-running the full
pipeline. Components are grouped into two categories:

- **Coding Tools** — opencode, oh-my-opencode-slim, Codex CLI, Claude
  Code, Pi agent
- **Infrastructure** — LiteLLM, Grafana, Prometheus, PostgreSQL (pinned,
  display-only — never auto-updated)

```bash
./scripts/update.sh              # interactive: show grouped table, select which to update
./scripts/update.sh --check      # show version table only
./scripts/update.sh --all        # update all components with updates available
./scripts/update.sh --dry-run    # show what would be updated
```

The script detects installed components, checks current vs latest
versions, and offers selective updates. It does NOT touch passwords,
API keys, or virtual keys — only updates binaries, npm packages, and
Docker images. After updating Docker images, the affected service is
automatically pulled and restarted.

## Option 4: Uninstall

Ask what to remove:

```bash
./scripts/uninstall.sh --all --dry-run   # preview
./scripts/uninstall.sh --tool=opencode   # one agent
./scripts/uninstall.sh --all             # everything
```

## Option 5: Install Skill

Install a **new** skill (not this companion) into all detected coding agents.

**Step 1**: Ask the user for:
- **Skill name** — a short directory name (e.g. `my-deploy-skill`)
- **Source** — a local file path or URL to a SKILL.md file

**Step 2**: Preview what would be installed:
```bash
./scripts/install-skill.sh --name=<name> --source=<source> --dry-run
```

**Step 3**: If the user confirms, install:
```bash
./scripts/install-skill.sh --name=<name> --source=<source>
```

This installs into all detected agents:
- opencode: `~/.config/opencode/skills/<name>/SKILL.md`
- codex: `~/.codex/skills/<name>/SKILL.md`
- pi: `~/.pi/agent/skills/<name>/SKILL.md`
- claude: `~/.claude/skills/<name>/SKILL.md`

**Step 4**: Remind the user to restart their coding agents for the new
skill to be discovered.

## Option 6: Mint New Keys

Ask the user which type of key:

**a) Add a MaaS load-balancing key**

This adds another Huawei MaaS API key for load balancing across multiple
keys, increasing throughput.

**Step 1**: Ask the user for the new MaaS API key (from Huawei cloud
console, region ap-southeast-1, starts with `sk-`).

**Step 2**: Read the current key count from `.env`:
```bash
CURRENT_COUNT=$(grep '^HUAWEI_MAAS_API_KEY_COUNT=' .env | cut -d= -f2 | tr -d '"')
NEW_INDEX=$CURRENT_COUNT
NEW_COUNT=$((CURRENT_COUNT + 1))
```

**Step 3**: Append the new key to `.env` and update the count:
```bash
echo "HUAWEI_MAAS_API_KEY_${NEW_INDEX}=\"sk-the-new-key\"" >> .env
sed -i "s/^HUAWEI_MAAS_API_KEY_COUNT=.*/HUAWEI_MAAS_API_KEY_COUNT=${NEW_COUNT}/" .env
```

**Step 4**: Regenerate LiteLLM config and restart (this creates
deployments for all models across all keys including the new one):
```bash
./scripts/02_litellm.sh
```

**Step 5**: Verify:
```bash
./scripts/04_validate.sh
```

**b) Mint a LiteLLM virtual key**

This creates a new virtual key for an additional coding tool or custom
integration. Each virtual key has its own budget and access control.

**Step 1**: Ask the user for:
- **Key alias** — a name for the key (e.g. `my-tool`)
- **Budget** — max spend in USD, or 0 for unlimited

**Step 2**: Read the master key from `.env`:
```bash
MASTER_KEY=$(grep '^LITELLM_MASTER_KEY=' .env | cut -d= -f2 | tr -d '"')
```

**Step 3**: Mint the key (omitting `models` grants access to all models,
matching `scripts/helpers/keys.sh`):
```bash
curl -X POST http://127.0.0.1:4000/key/generate \
  -H "Authorization: Bearer $MASTER_KEY" \
  -H "Content-Type: application/json" \
  -d '{"key_alias": "<alias>", "max_budget": <budget>}'
```

**Step 4**: Show the returned key to the user and tell them to use it as
the API key in their tool's config.

**Step 5**: Remind the user they can view all keys at
`http://127.0.0.1:4000/ui` (login: `admin` / master key).

---

## Model Management

Models are in `scripts/helpers/models.sh`. Format:
```
model_name:tpm:rpm:max_tokens:max_input:max_output:input_cost:output_cost:cache_read_cost:cache_creation_cost
```
`cache_read_cost` and `cache_creation_cost` are 0 for models without
cache support.
Off-peak pricing is configured separately in the `OFF_PEAK_PRICING` array
(format: `model_name|hours_utc|input_cost|output_cost|cache_read_cost`).
Only models with off-peak pricing are listed; rates are absolute values.

Current models: `glm-5.3`, `glm-5.2`, `glm-5.1`, `deepseek-v4.1-flash`.

**List models**:
```bash
sed -n '/^MODELS=(/,/^)/p' scripts/helpers/models.sh | grep -E '^[[:space:]]*"' | sed 's/^[[:space:]]*"//; s/:.*//' | sort
```

**Add a model**: add a line to the `MODELS` array in `scripts/helpers/models.sh`,
then update `configs/litellm/config.yaml.template`, `configs/opencode/opencode.json.template`,
and `configs/codex/model_catalog.json` with the new model. If the model surfaces
reasoning (`reasoning_effort` pass-through or thinking mode), add it to the
`REASONING_MODELS` array. If it has off-peak
pricing, add it to the `OFF_PEAK_PRICING` array. If it accepts image input,
add it to the `VISION_MODELS` array and set `modalities` to include
`image` in `input` on its `opencode.json.template` entries. Update
`configs/opencode/oh-my-opencode-slim.json.template` only if agents should be
assigned the new model. Then regenerate (this creates 2N deployments per model,
two per API key — one per format):
```bash
./scripts/02_litellm.sh
./scripts/04_validate.sh
```

**Remove a model**: delete the line from `MODELS`, then same regenerate +
validate.

## Debug Routing

**401 errors**:
```bash
docker compose logs litellm --tail 100 | grep 401
```

**Slow/no response**:
```bash
curl -sf http://127.0.0.1:4000/health/liveliness
docker compose logs litellm --tail 100 | grep -i error
curl -sf 'http://127.0.0.1:9090/api/v1/query?query=litellm_request_total_latency_metric_sum' | jq .
```

**Inference smoke test**:
```bash
MASTER_KEY=$(grep '^LITELLM_MASTER_KEY=' .env | cut -d= -f2 | tr -d '"')
curl -X POST http://127.0.0.1:4000/v1/chat/completions \
  -H "Authorization: Bearer $MASTER_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model": "glm-5.1", "messages": [{"role": "user", "content": "hi"}], "max_tokens": 5}'
```

## View Metrics

```bash
curl -sf 'http://127.0.0.1:9090/api/v1/query?query=litellm_proxy_total_requests_metric_total' | jq .
curl -sf 'http://127.0.0.1:9090/api/v1/query?query=litellm_spend_metric_total' | jq .
curl -sf 'http://127.0.0.1:9090/api/v1/query?query=rate(litellm_deployment_failure_responses_total[5m])' | jq .
```

Grafana: `http://127.0.0.1:3000` — 44-panel dashboard (7 row headers + 37 visualization panels).

### Visual dashboard verification

After any change to `configs/grafana/dashboards/main.json`, capture the
rendered dashboard and review it visually — JSON validity says nothing
about layout, colors, or broken queries.

```bash
./scripts/07_dashboard_shots.sh   # writes full.png + band-NN.png to /tmp/dashboard-shots
```

Feed the bands to a vision-capable agent and have it check them against
the design intent: row order, panel alignment, color semantics, units in
titles, and empty or broken panels. Fix defects, re-run the script, and
re-verify until clean. `04_validate.sh` reports whether the capture
capability is installed (pass/skip — optional, no hard dependency).
One-time setup:

```bash
pip3 install --user --break-system-packages playwright && python3 -m playwright install chromium
```

## Verify Spend and Off-Peak Discount

LiteLLM tracks per-request spend in its database. Fetch recent requests with
the master key:

```bash
MASTER_KEY=$(grep '^LITELLM_MASTER_KEY=' .env | cut -d= -f2 | tr -d '"')
curl -s "http://127.0.0.1:4000/spend/logs" -H "Authorization: Bearer $MASTER_KEY" \
  | jq '[.[] | {model: .model_group, start: .startTime, spend: .spend,
                tokens_in: .prompt_tokens, tokens_out: .completion_tokens}] | .[0:10]'
```

The response is a bare JSON array, newest first, up to 10000 entries. A
`limit` query param is ignored — slice with jq as above. Each entry's
`startTime` is an ISO 8601 UTC string; `spend` is USD; `cache_hit` is a
string (`"True"`/`"False"`/`"None"`), not a boolean.

**Off-peak window**: glm-5.2 and glm-5.1 bill at 70% of peak rates and
deepseek-v4.1-flash at 50%, from 13:00 to 00:00 UTC (21:00–07:59 Beijing).
glm-5.3 has flat pricing (no off-peak discount). LiteLLM checks the window
when the request completes, so a request started at 23:59 UTC bills at peak
if it finishes after 00:00 UTC.

**Verify the discount**: recompute a request's cost from the peak rates and
compare — off-peak spend is exactly 70% (glm-5.2/glm-5.1) or 50%
(deepseek-v4.1-flash) of the peak cost for the same tokens.
Observed examples (small requests, no cache):

```text
peak:      01:39 UTC  glm-5.2  13 in / 38 out   → $0.0001854 = 13×$1.4/M  + 38×$4.4/M
off-peak:  23:58 UTC  glm-5.2  13 in / 252 out  → $0.0007889 = 13×$0.98/M + 252×$3.08/M
```

Large requests rarely match the simple formula — cached input tokens are
billed at the (also discount-scaled) cache-hit rate. To verify the discount,
use small requests or compare spend-per-token across the window boundary.

If a request inside the off-peak window bills at 100% of peak, the running
container may have loaded a config without `off_peak_pricing` blocks — check
`configs/litellm/config.yaml` and `docker compose restart litellm`.

---

## Recovery

| Symptom | Fix |
|---------|-----|
| `.env not found` / `placeholder value` | `./scripts/01_env.sh` |
| Fewer than 4 containers running | `docker compose up -d`, wait 30s |
| LiteLLM liveness probe fails | `docker compose logs litellm --tail 50` |
| Inference smoke test fails | Check MaaS key in `.env`; `docker compose logs litellm --tail 100` |
| `opencode not found` / config issues | `./scripts/03a_opencode.sh` |
| `codex not found` / config issues | `./scripts/03b_codex.sh` |
| `claude not found` / config issues | `./scripts/03c_claude_code.sh` |
| `pi not found` / config issues | `./scripts/03d_pi.sh` |
| Prometheus not reachable | `docker compose up -d prometheus`, wait 10s |
| `/metrics` endpoint not responding | `docker compose restart litellm`, wait 15s |
| Grafana not reachable | `docker compose up -d grafana`, wait 20s |
| Docker daemon not running | `systemctl start docker` |
| Port 4000/3000/9090 in use | `lsof -i :<port>`, stop conflicting process |
| Need to preserve spend history before a reset | `./scripts/06_backup.sh` first — then `docker compose down -v` is safe |
| `git pull` conflicts on upgrade | `git stash && git pull && git stash pop` |
| Coding tool outdated version | `./scripts/update.sh --check` to see available updates, then `./scripts/update.sh` to update |
| Stale models in config (deepseek-v4-pro, deepseek-v4-flash) | `./scripts/03a_opencode.sh && ./scripts/03d_pi.sh` to regenerate configs from current catalog |

## Remote Access

**SSH forwarding (recommended):**
```bash
ssh -L 4000:127.0.0.1:4000 -L 3000:127.0.0.1:3000 -L 9090:127.0.0.1:9090 user@vm
```

**Bind to all interfaces:** set `BIND_ADDRESS="0.0.0.0"` in `.env`, then
`docker compose up -d`.

Files in this skill

  • AGENTS.md13.9 KB
  • CHANGELOG.md42 KB
  • INSTALLATION.md23.3 KB
  • REFERENCE.md33.4 KB
  • SKILL.md11.2 KB
  • VERSION7 B
  • configs/.env.template2.2 KB
  • docker-compose.yml4.3 KB
  • scripts/01_env.sh16.4 KB
  • scripts/02_litellm.sh13.6 KB
  • scripts/03a_opencode.sh9.5 KB
  • scripts/03b_codex.sh6 KB
  • scripts/03c_claude_code.sh7.6 KB
  • scripts/03d_pi.sh8.3 KB
  • scripts/04_validate.sh42.5 KB
  • scripts/05_skill.sh4.3 KB
  • scripts/bootstrap.sh27.8 KB
  • scripts/install-skill.sh3.3 KB
  • scripts/uninstall.sh16.2 KB
  • scripts/update.sh16.4 KB

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…