Installs into .claude/skills of the current project.
Are you the author of Terrakube Ops?
Add the live security badge to your README. It updates with every re-scan.
[](https://www.skillsdirectory.com/skills/dryvist-terrakube-ops)
---
name: terrakube-ops
description: Operate self-hosted Terrakube (remote OpenTofu/Terraform plan/apply, state, locking) — login/plan/apply flow, why targeted apply is dangerous, lock recovery, offline-mirror gotcha, token rotation. Use for a plan/apply, stuck lock, new workspace.
---
# Terrakube operations
[Terrakube](https://terrakube.io) is a self-hosted remote-execution platform
for OpenTofu/Terraform — plans and applies run on its own executor, not on
the operator's machine or in CI. This skill is the generic operating model;
your own workspace names, hostnames, and credential source stay in your own
inventory.
## Canonical workflow
```bash
tofu login <terrakube-hostname> # once per machine — opens a browser SSO flow
tofu plan # runs remotely; CLI streams the live plan log
tofu apply # same remote workspace, same run
tofu state list # inspect state / outputs afterward
tofu output
```
`tofu plan`/`apply` from the CLI and starting a run from the Terrakube web UI
both drive the same remote executor and the same job log — pick whichever is
convenient; they're interchangeable.
> Local `tofu init -backend=false` + `tofu validate` catches syntax errors
> without contacting any provider — run that first, it's free and instant
> compared to a remote plan.
## Never use a targeted apply
A targeted apply (`-target=...`) can leave the real infrastructure and any
downstream-published inventory describing two different worlds — whatever
consumes that inventory (configuration management, DNS, monitoring) now
disagrees with reality. Apply the reviewed **whole-workspace** plan. For
state surgery, prefer `moved`/`removed` blocks over `-target` or manual state
edits.
## Workspace locking and recovery
Terrakube owns one run queue and one lock per workspace — a second run simply
queues behind the active one; independent workspaces don't share a lock.
- Cancel a genuinely stuck run from the workspace UI, not by force-unlocking
blind.
- Confirm the executor has actually stopped before force-cancelling or
unlocking — cancelling a run that's still writing state is how corruption
happens.
- Prefer fixing forward (revert the config change in git, plan/apply again)
over restoring an older state version. When state genuinely is corrupted,
preserve the *current* state version before restoring an older one — you
may need to diff them later.
## The offline-mirror gotcha
An executor eagerly initializes its Terraform-compatible provider downloader
even on a workspace that only ever runs **OpenTofu**. If your platform points
release/provider resolution at an internal mirror for offline operation,
**both** the Terraform-compatibility URL and the OpenTofu release URL need to
point at that mirror — leaving either one pointed at a public release service
reintroduces an internet dependency and breaks offline recovery exactly when
you need it not to.
## Token rotation invalidates every session
Rotating whatever secret Terrakube uses to sign its own issued tokens (a
`PAT_SECRET`/`INTERNAL_SECRET`-style value) invalidates **every** previously
issued personal access token at once. Plan for a `tofu login` re-auth on
every machine that touches the platform right after that rotation — it isn't
a sign anything else is broken.
## Never judge an apply by its exit code alone
A remote apply can exit 0 while its log still contains `Error:` lines and no
"Apply complete" summary — a partial failure that looks clean from the exit
status. Grep the apply output for both `Error:` and the completion summary
(`Apply complete` / `Resources:`); the summary's absence is the real signal,
not the process exit code.
## Importing existing infrastructure into an empty workspace
Onboarding a root that already manages **live** infrastructure into a
workspace whose state starts empty is not a normal apply, and the ordinary
0-destroy review is not enough to make it safe.
> A plan against an empty workspace reads "create everything, 0 destroy."
> That falsely passes a naive 0-destroy gate — the apply then hits an
> already-exists error and duplicates live infrastructure. For an
> empty-state workspace the real gate is a **refresh-only plan showing
> 0 add / 0 change / 0 destroy** — proof state already matches reality. A
> create-everything plan on a workspace that should already hold resources
> means the state was never migrated.
Recipe: migrate the existing state into the workspace first (a state push,
or an init with state migration against the new backend), run a
refresh-only plan and require exactly 0/0/0, reconcile any benign drift with
`moved`/`removed` blocks rather than destroy-and-recreate, and only then is
the ordinary 0-destroy apply gate meaningful.
## Verification
1. The run reaches `applied` (CLI exits 0; UI shows a green Applied state)
**and** the log shows a completion summary with no `Error:` lines — see
above.
2. `tofu output` reflects the expected values.
3. Re-running `tofu plan` immediately after a clean apply reports "No
changes" — if it doesn't, something outside the applied config drifted,
or the apply didn't fully land.
## Related
- **proxmox-cluster-ops** (this plugin) — the guest-level operations this
IaC layer should be the only thing changing.
- **infrastructure-standards** (infra-standards) — the inventory contract a
producer workspace's outputs typically feed.