Push Grafana Cloud dashboards and alert rules (dashboard- and alerts-as-code). Runs playbooks/deploy_dashboard.yml and playbooks/deploy_alerts.yml. TRIGGER when: the user wants to push, deploy, or update the Grafana dashboard or alert rules, apply dashboard-as-code, set up the cluster health dashboard, add or change a Grafana alert, or asks why no alerts exist. SKIP: if the user only wants to query Grafana Cloud (that is the MCP server from /sales-demos-mcp) or deploy Alloy (that is /sales-de...
Installs into .claude/skills of the current project.
Are you the author of Sales Demos Dashboard?
Add the live security badge to your README. It updates with every re-scan.
[](https://www.skillsdirectory.com/skills/ericcames-sales-demos-dashboard)
---
name: sales-demos-dashboard
description: "Push Grafana Cloud dashboards and alert rules (dashboard- and alerts-as-code). Runs playbooks/deploy_dashboard.yml and playbooks/deploy_alerts.yml. TRIGGER when: the user wants to push, deploy, or update the Grafana dashboard or alert rules, apply dashboard-as-code, set up the cluster health dashboard, add or change a Grafana alert, or asks why no alerts exist. SKIP: if the user only wants to query Grafana Cloud (that is the MCP server from /sales-demos-mcp) or deploy Alloy (that is /sales-demos-alloy)."
---
# sales-demos-dashboard
Push Grafana Cloud dashboards and alert rules defined as committed JSON. Issues
[#275](https://github.com/ericcames/sales.demos/issues/275) (dashboard) and
[#629](https://github.com/ericcames/sales.demos/issues/629) (alerts).
## There is an AAP path now too (#318)
`AAP Observability - 2 Deploy Dashboards` runs the same playbook from AAP, so
this no longer has to come off a laptop. Use whichever suits; the skill is still
the quicker loop while iterating on dashboard JSON.
**It is not per-environment, and that surprises people.** One Grafana Cloud
serves both environments, so the template exists in both controllers and pushes
to the *same* folder — running it from demo also updates what sandbox sees.
This skill contains **no logic**. All the work is in
[`playbooks/deploy_dashboard.yml`](../../../playbooks/deploy_dashboard.yml). See
`CLAUDE.md` → *Skills and playbooks*.
## What it does
1. Creates a "Sales Demos" folder in Grafana Cloud (idempotent)
2. Reads `playbooks/files/grafana/cluster-health.json`
3. Pushes the dashboard via the Grafana HTTP API with `overwrite: true`
The dashboard covers cluster nodes, KubeVirt VMs, AAP platform health, and
logs. A `cluster` template variable makes it work for both sandbox and demo.
`deploy_alerts.yml` does the same for alert rules
(`playbooks/files/grafana/alert-rules.json`): it PUTs one rule group,
`sales-demos-health`, into the same folder. The PUT replaces the whole group, so
a rule removed from the JSON is removed from Grafana. Seven rules, each labelled
by `cluster`:
| Rule | Fires when |
|---|---|
| Alloy federation down | a cluster reported metrics in the last hour but not now (5m) |
| AAP controller metrics down | the AAP metrics scrape answered in the last hour but not now (5m) |
| Running VM count dropped | fewer VMs running than 10 minutes ago (1m) — expected after a teardown |
| AAP jobs stuck pending | any pending job for 15m |
| Free-tier series budget above 80% | over 8,000 active series stack-wide (15m) |
| Node under disk pressure | kubelet reports DiskPressure on any node (immediate) — it is already evicting, and AAP job pods request no ephemeral-storage so they rank among the first taken (#782) |
| Node disk approaching the eviction threshold | a node's `/var` is above 82% used for 15m — eviction begins at 85% (#782) |
**No contact point is configured.** Firing alerts follow the stack's default
notification policy; adding a receiver would put an address in a public repo.
Rules stay editable in the UI (`X-Disable-Provenance`), and the next run puts the
committed version back.
AAP path: `AAP Observability - 3 Deploy Alerts`, not per-environment, like template 2.
## Preflight Check
Run these before doing anything else. Every one must pass.
```bash
VAULT_ID="sales.demos@$HOME/secrets/.vault_pass_sales_demos"
# 1. Vault password file exists
test -s "$HOME/secrets/.vault_pass_sales_demos" \
&& echo "pass: vault password file" \
|| echo "FAIL: ~/secrets/.vault_pass_sales_demos missing"
# 2. secrets.yml exists and is vault-encrypted
head -c 15 playbooks/group_vars/all/secrets.yml 2>/dev/null | grep -q '^\$ANSIBLE_VAULT' \
&& echo "pass: secrets.yml is vault-encrypted" \
|| echo "FAIL: secrets.yml missing or NOT encrypted — see /sales-demos-first-time"
# 3. Grafana Cloud Editor SA token is filled in
ansible-vault view playbooks/group_vars/all/secrets.yml --vault-id "$VAULT_ID" 2>/dev/null \
| python3 -c "
import sys, yaml
d = yaml.safe_load(sys.stdin) or {}
keys = ['grafana_cloud_url', 'grafana_cloud_editor_sa_token']
bad = [k for k in keys if d.get(k, 'CHANGEME') == 'CHANGEME' or k not in d]
print(('FAIL: missing or CHANGEME: ' + ', '.join(bad)) if bad
else 'pass: Grafana Cloud dashboard credentials filled in')
"
# 4. Dashboard JSON exists
test -f playbooks/files/grafana/cluster-health.json \
&& echo "pass: dashboard JSON exists" \
|| echo "FAIL: playbooks/files/grafana/cluster-health.json missing"
```
If any check fails, stop and tell the user exactly which one and the fix shown
beside it. Do not attempt the run with a failing prerequisite.
### If the Editor SA token is missing
The user must create it manually in Grafana Cloud:
1. Administration > Service Accounts > Add
2. Name: `sales-demos-editor`, Role: **Editor**
3. Add token > copy the `glsa_...` value
4. Add to vault as `grafana_cloud_editor_sa_token`
This is separate from the Viewer SA token used by the MCP server.
## Run
```bash
mkdir -p ~/ansible-logs
export ANSIBLE_LOG_PATH=~/ansible-logs/deploy-dashboard-$(date +%F-%H%M).log
./utilities/run-ansible.sh playbooks/deploy_dashboard.yml -i inventory \
--vault-id sales.demos@~/secrets/.vault_pass_sales_demos
```
**Always set `ANSIBLE_LOG_PATH`** — logs live outside the repo, in
`~/ansible-logs/`. Tell the user the path.
**No `--limit` needed.** This playbook targets localhost because Grafana Cloud
is a single external service. Do NOT pass `--limit sandbox` or `--limit demo`
— localhost is not in those groups and the play will skip with "no hosts
matched".
**`-i inventory` is required** even though the play targets localhost, because
Ansible needs the inventory path to resolve `group_vars/all/` for vault
variable loading.
This takes under 30 seconds.
### Alert rules
```bash
export ANSIBLE_LOG_PATH=~/ansible-logs/deploy-alerts-$(date +%F-%H%M).log
./utilities/run-ansible.sh playbooks/deploy_alerts.yml -i inventory \
--vault-id sales.demos@~/secrets/.vault_pass_sales_demos
# reversal — deletes the rule group, then asserts it is gone
./utilities/run-ansible.sh playbooks/deploy_alerts.yml -i inventory \
-e alerts_state=absent \
--vault-id sales.demos@~/secrets/.vault_pass_sales_demos
```
The playbook reads the group back and asserts every committed rule uid is there,
so a green run already means Grafana holds the rules. Confirm from the agent's
side anyway: `alerting_manage_rules` with `operation: list` should show the five
rules in folder "Sales Demos", each `normal` unless something is genuinely wrong.
## Verify via Grafana MCP
Use the Grafana MCP server to confirm the dashboard was pushed:
1. **Queries match:** `get_dashboard_panel_queries` with uid
`sales-demos-cluster-health` — returns every panel's query expression. This
is the fastest way to confirm a specific panel change landed.
2. **Dashboard exists:** `search_dashboards` with query `cluster health` —
should return "Sales Demos - Cluster Health" in the "Sales Demos" folder.
3. **Full model:** `get_dashboard_by_uid` with uid
`sales-demos-cluster-health` — returns the complete dashboard. Use
`get_dashboard_property` with a JSONPath to check a specific field without
pulling the whole model (e.g., `$.panels[*].options.textMode`).
Then open the dashboard URL printed by the playbook and confirm panels render
with live data.
## v1/v2 schema note
The playbook pushes via the legacy v1 API (`POST /api/dashboards/db`). This
Grafana Cloud stack stores dashboards in `v0alpha1` format
(`status.conversion.storedVersion`) and converts to v2 on read. The v1 write
path has been reliable for all fields so far, but if a future panel option
appears wrong in the live dashboard despite the v1 API returning the correct
value, check the v2 apiserver directly:
```
GET /apis/dashboard.grafana.app/v2/namespaces/stacks-<stack-id>/dashboards/sales-demos-cluster-health
```
`<stack-id>` is your stack's numeric ID: Grafana Cloud portal › your stack ›
Details, or the `stack_id` label on `grafanacloud_instance_info` in the
`grafanacloud-usage` data source. Like the stack URL, it identifies the
account, so it is never written into a tracked file (#635).
Compare `resourceVersion` and `generation` between the v2 response and what
the Grafana UI is rendering — a mismatch indicates read replica lag or a
stale client session, not a write failure.
## When it finishes
Report the dashboard URL and tell the user the dashboard is live in Grafana
Cloud. Remind them to select the `cluster` variable (e.g., `sandbox`) to see
data.
## If it fails
| Symptom | Cause | Fix |
|---|---|---|
| Assertion fails on `grafana_cloud_editor_sa_token` | Token not in vault | Create the Editor SA in Grafana Cloud UI, add to vault |
| `401 Unauthorized` | Token expired or revoked | Regenerate in Grafana Cloud > Service Accounts |
| `403 Forbidden` | Token has Viewer role, not Editor | Create a new SA with Editor role |
| `412 Precondition Failed` on folder creation | Folder already exists and was modified | Already handled by the playbook (accepts 200, 409, 412) |
| `Attempting to decrypt but no vault secrets found` | `--vault-id` missing | Add `--vault-id sales.demos@~/secrets/.vault_pass_sales_demos` |
| `no hosts matched` / skipping | `--limit` was passed | Remove `--limit` — this play targets localhost |
Never paste a Grafana Cloud URL or token into a commit message, issue, or PR.
This repo is public — see `CLAUDE.md`.