Phase 3 — build demo VMs on OpenShift Virtualization with Terraform and register them as managed hosts in AAP, ready for the daily-demo content to run against. Runs playbooks/provision_vm.yml. TRIGGER when: the user asks to provision, create, build or spin up demo VMs, wants a Linux or Windows VM for a demo, or asks for a specific size tier. SKIP: if the environment has never been set up — that is sales-demos-setup — or if the user wants to destroy VMs, which is sales-demos-teardown.
Installs into .claude/skills of the current project.
Are you the author of Sales Demos Provision?
Add the live security badge to your README. It updates with every re-scan.
[](https://www.skillsdirectory.com/skills/ericcames-sales-demos-provision)
---
name: sales-demos-provision
description: "Phase 3 — build demo VMs on OpenShift Virtualization with Terraform and register them as managed hosts in AAP, ready for the daily-demo content to run against. Runs playbooks/provision_vm.yml. TRIGGER when: the user asks to provision, create, build or spin up demo VMs, wants a Linux or Windows VM for a demo, or asks for a specific size tier. SKIP: if the environment has never been set up — that is sales-demos-setup — or if the user wants to destroy VMs, which is sales-demos-teardown."
---
# sales-demos-provision
Phase 3. Runs `terraform/ocpvirt/` and registers the resulting VMs in AAP so the
demo content has hosts to run against.
This skill contains **no logic**. All the work is in
[`playbooks/provision_vm.yml`](../../../playbooks/provision_vm.yml), which is the
same playbook the `Linux Day 1 - 1 Provision` job template runs, with survey
answers mapped to the same variable names. See `CLAUDE.md` →
*Skills and playbooks*.
**Prefer the job template when AAP is available** — it is the demo-able path and
the one a customer sees. This skill is for when you are working ahead of AAP, or
debugging a run without the controller in the way.
## The inputs are the contract
| Variable | Values | Default |
|---|---|---|
| `vm_size_tier` | `small`, `medium`, `large` | `small` |
| `os_type` | `linux`, `windows`, `both` | `linux` |
| `vm_role` | `web`, `db`, `app` (1-8 lowercase alphanumeric) | `web` |
| `vm_count` | `1` or `2` | `1` |
These names are shared verbatim with the AAP survey and
`terraform/ocpvirt/variables.tf`. Changing one means changing all three.
**`vm_role` and `vm_count` build a farm (#389).** VMs are named
`{role}-{os}-{index}` — `web-win-1`, `db-lnx-2` — and `vm_count=1` behaves
exactly as a single-VM run always did. Two things follow that are easy to get
wrong:
- **Each role has its OWN Terraform state** (`secret_suffix=<env>-<os>-<role>`),
the same way each OS has since #301. So `vm_role=db` cannot disturb a running
`web` farm — and **a teardown must be given the role it was built with**, or it
inits an empty state, destroys nothing, and still reports success.
- **`vm_role` is capped at 8 characters** because a Windows computer name is
capped at 15 and the name is used verbatim as the NetBIOS hostname. The longest
name this builds is `{8}-win-{2 digits}` = 15 exactly. A non-empty
`name_suffix` spends from the same budget and is checked at plan time.
**`os_type=windows` or `both` requires that the environment is linked to the
published CIS L1 hardened Windows golden image.** CNV ships `win2k22` as an empty
DataSource placeholder; on a new environment, run `sales-demos-windows-image` first
to fill it. The playbook preflights that DataSource and **warns rather than
refuses**, because `os_type=both` still gets a working Linux guest and linking a
minute later fixes the Windows half without re-provisioning.
## Preflight Check
```bash
./utilities/preflight.sh "${ENV:-sandbox}" --terraform
```
**Never pipe the run through `tee`.** In a pipeline the exit status comes from
`tee`, not `ansible-playbook`, so a failed run reports success.
## Run
```bash
./utilities/run-ansible.sh playbooks/provision_vm.yml -i inventory --limit sandbox \
-e target_env=sandbox -e os_type=linux -e vm_size_tier=small \
--vault-id sales.demos@~/secrets/.vault_pass_sales_demos
```
Idempotent — re-running converges rather than rebuilding. A second run reports
`changed=0`, and that is the check that it is behaving.
## Verify it in the EE before merging a change
See `/sales-demos-verify-ee` for why and how. The one command:
```bash
utilities/run-in-ee.sh playbooks/provision_vm.yml \
-i inventory --limit sandbox -e target_env=sandbox \
--vault-id sales.demos@~/secrets/.vault_pass_sales_demos
```
## What it does
1. Ensures the **state namespace** exists — the Terraform kubernetes backend
needs it before there is any state to describe it with. It does **not** create
the VM namespace; Terraform owns that.
2. `terraform init` against the kubernetes backend, then `apply`.
3. Reads the outputs and registers each VM in the `Sales Demo VMs` inventory —
`linuxweb` with SSH vars, `windemo` with WinRM vars.
## Verify against the cluster, not the recap
```
mcp__openshift-<env>__resources_list kubevirt.io/v1 VirtualMachine
namespace: sales-demos-<env>
mcp__openshift-<env>__resources_list kubevirt.io/v1 VirtualMachineInstance
namespace: sales-demos-<env>
```
`apply` returning does **not** mean the guest is up: the default StorageClass is
`WaitForFirstConsumer`, so the disk clones only when the VM first schedules.
Expect the VM to reach `Running` roughly 45s after apply on a warm environment.
Then confirm AAP can actually reach it — that is what `Linux Day 1 - 5 Check and Gather Facts`
is for, and it is the difference between "a VM exists" and "the demo will work".
## Notes worth having before you debug
- **AAP reaches the VMs over in-cluster DNS on port 22.** It runs on the same
cluster and each VM has a headless Service. No bastion, and no `virtctl` —
that is the laptop path, and the execution environment does not ship it.
- **From a laptop, use `virtctl ssh`.** The `ssh_command` Terraform output gives
you the exact line, and the Provision, Configure and Check job logs print it
too. If a line you kept from before #49 fails with `unknown flag:
--local-ssh`, that flag was removed in virtctl v1.x — drop it, and keep the
`vm/` prefix on the target.
- **`demo_ssh_public_key` must not be empty.** cloud-init then emits
`ssh_pwauth: true` with no key *and* no password, and the guest has no
credentials at all. It writes authorized keys on **first boot only**, so a VM
created that way must be re-created, not restarted.
- **State lives in `sales-demos-tfstate`**, a long-lived namespace of its own,
keyed per environment by `secret_suffix`. Never delete it.
## `Error acquiring the state lock`
The kubernetes backend holds a lock for the length of an apply or destroy and
releases it when terraform exits. A job that is **cancelled**, times out, or has
its pod evicted never gets there, so the lock outlives the run that took it and
every later run fails to acquire it. Cancelling a job is a normal thing to do —
this is not an edge case, and it stays invisible until the next demo.
The playbook detects this and fails with the lock ID and the command (#46).
Reading the message:
- **`Who:` is a lie in AAP.** It shows something like
`1000770000@automation-job-92-qswfk`. That pod is gone. It reads like a run in
progress, which encourages waiting — waiting never clears it.
- **Nothing was changed.** The lock is taken before any work starts, so a locked
run created and destroyed nothing.
**The playbook's failure message prints the exact commands.** Use them verbatim —
they already carry the correct `ocpvirt_state_suffix` for the run that failed.
If you need to check manually, the Lease name is
`lock-tfstate-default-<env>-<os>-<role>` (e.g. `lock-tfstate-default-sandbox-linux-web`).
Legacy `lock-tfstate-default-<env>` Leases still exist with an empty holder —
checking those returns nothing and falsely confirms "no lock is held" (#402).
```
mcp__openshift-<env>__resources_get coordination.k8s.io/v1 Lease lock-tfstate-default-<env>-<os>-<role>
namespace: sales-demos-tfstate
# Check spec.holderIdentity — empty means no lock is held
```
A held lock shows the same value as the `ID:` line in the error. Empty output on
the **correct** Lease means no lock is held and the failure is something else.
To clear it, first confirm in AAP that no Provision or Teardown job is genuinely
running — force-unlocking a live apply corrupts state. Then:
```bash
cd terraform/ocpvirt
terraform init -reconfigure \
-backend-config=secret_suffix=<env>-<os>-<role> \
-backend-config=namespace=sales-demos-tfstate \
-backend-config=config_path=../../.kube/<env>.kubeconfig \
-backend-config=insecure=true
terraform force-unlock <lock-id>
```
**Nothing force-unlocks by itself, deliberately.** Clearing a stale lock is a
recoverable annoyance; clearing a live one corrupts the state file. Making it
automatic safely would need a liveness check against the AAP job, not the pod
name in the message.