Skip to content
Back to skills

Testing Mcp With Cli Agents

ASecurity

Test an MCP server by driving real CLI agents (Claude, Codex, Cursor, Gemini, Grok, agy, opencode) against it, using isolated tmux sockets and send-keys instead of trusting unit tests alone. Use this whenever verifying MCP-server behavior end-to-end, checking that a local branch or checkout works across installed agent CLIs, comparing trunk-vs-branch MCP behavior, driving an interactive agent TUI to exercise approval flows or cancellation, reproducing a tmux-MCP bug through a live client, or ...

  • 13 stars
  • 0 votes
  • 0 copies
  • 0 views
  • Added September 26, 2026
ai-agentsrustgoshellnodeawstestinggitapi

Works with

  • cursor
  • cli
  • api
  • mcp

Security analysis

A100/100

Pro scans all 2 files and shows the line behind each finding

Scanned September 26, 2026

npx -y skills add tmux-python/libtmux-mcp --skill testing-mcp-with-cli-agents --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Testing Mcp With Cli Agents?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for Testing Mcp With Cli Agents
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/tmux-python-testing-mcp-with-cli-agents/badge)](https://www.skillsdirectory.com/skills/tmux-python-testing-mcp-with-cli-agents)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
SKILL.md
---
name: testing-mcp-with-cli-agents
description: >-
  Test an MCP server by driving real CLI agents (Claude, Codex, Cursor, Gemini,
  Grok, agy, opencode) against it, using isolated tmux sockets and send-keys
  instead of
  trusting unit tests alone. Use this whenever verifying MCP-server behavior
  end-to-end, checking that a local branch or checkout works across installed
  agent CLIs, comparing trunk-vs-branch MCP behavior, driving an interactive
  agent TUI to exercise approval flows or cancellation, reproducing a tmux-MCP
  bug through a live client, or wiring a checkout into the CLIs with mcp_swap.
  Reach for it even when the user only says "test the MCP", "does the branch
  work in the agents", "drive the CLI to call the tool", or "check it across
  Codex/Gemini/Cursor" without naming tmux or sockets explicitly.
---

# Testing an MCP server through real CLI agents

Unit tests prove the server's internals; they don't prove a real agent can
discover a tool, get past its approval gate, call it, and survive cancelling it
mid-flight. This skill exercises that whole path by pointing installed CLI
agents at a checkout and driving them — with the tmux tool surface (libtmux-mcp)
as the running example, though the shape generalizes to any MCP server.

## The core idea: three tmux servers, never one

The single biggest mistake is running everything on one tmux socket. Keep three
distinct servers, each on its own socket, and the whole thing becomes safe and
observable:

| Role | Socket | Who touches it | Why separate |
|---|---|---|---|
| Your real session | default | you, interactively | must never be mutated by a test |
| Harness | `tmux -L cli-harness` | the driver: `send-keys` prompts in, `capture-pane` render out | isolates the CLI TUI you're driving |
| MCP-target | `tmux -L mcp-target`, exported as `LIBTMUX_SOCKET=mcp-target` to the server | the MCP server, when the agent calls tmux tools | independent ground-truth; a destructive tool here can't kill the agent you're driving |

The server resolves its socket from `LIBTMUX_SOCKET` and runs `tmux -L <name>`
(`src/libtmux_mcp/_utils.py` builds argv with `-L server.socket_name`, defaulted
from that env var). Exporting `LIBTMUX_SOCKET=mcp-target` into the server's
config env fully sandboxes every tmux tool call onto a scratch server.
Pre-create it so the agent has something to see:

```console
$ tmux -L mcp-target new-session -d -s scratch
```

Keeping **harness** separate from **MCP-target** is the load-bearing part: if the
agent's own TUI pane lived on the socket its MCP server mutates, one
`kill-server` / `kill-session` tool call would tear down the agent mid-test, and
your captures would be polluted by the agent's own UI redraws.

## Climb only as high as the question needs — three fidelity layers

### Layer 0 — Direct MCP smoke, no CLI at all

Fastest and most deterministic. Drive the server over stdio from a tiny FastMCP
client against an isolated socket and assert the wire contract directly: the
tool list, a couple of representative calls, an error path. Use this to answer
"is the tool surface and result shape correct?" before spending a CLI on it.

Shape-normalization gotchas seen in practice — normalize before asserting:
- Match the current `LIBTMUX_TOOLSETS` selection: excluded toolsets are hidden
  when omitted from the selection, so don't assert they're visible.
- List-returning tools surface `structuredContent` as `{"result": [...]}`;
  `capture_pane` output can come back under `result` as a string. Don't assume a
  top-level `count`.

### Layer 1 — Headless CLI one-shot

Proves the real client can discover and call the tools, scriptably, with no
send-keys. Every CLI has a non-interactive mode. A cheap discovery proof (does
the client *see* the server?) is worth running before spending a model turn —
but the cheapest proof differs sharply per CLI: grok's `mcp doctor` does a real
handshake, codex's `mcp get` only parses config, and agy has no proof short of a
model call. `references/cli-matrix.md` has the verified per-CLI invocation,
isolation lever, and approval-bypass flag for all six. Two things that surprise
people: some `mcp list`/`list-tools` subcommands read the *ambient* config and
ignore your isolated one, and a write-capable tool call needs a per-CLI
approval-bypass flag or it hangs on a no-TTY prompt. Flags drift — re-verify with
`--help`.

### Layer 2 — Interactive, driven by tmux send-keys

The high-fidelity path, and the only one that exercises approval flows, live
streaming, multi-turn, and cancellation. The loop:

```console
# 1. launch the agent TUI in a WIDE harness pane (so TUI text isn't wrapped)
$ tmux -L cli-harness new-session -d -s codex -x 220 -y 50
$ tmux -L cli-harness send-keys -t codex 'cd /repo && LIBTMUX_SOCKET=mcp-target codex' Enter

# 2. wait for the prompt to render, THEN type — never type blind
$ tmux -L cli-harness capture-pane -p -t codex | tail -5     # poll until ready
$ tmux -L cli-harness send-keys -t codex 'Use the tmux MCP to create a window named probe, then list windows' Enter

# 3. handle the approval gate — the #1 hang source (see below)
$ tmux -L cli-harness send-keys -t codex 'y' Enter           # or the mapped key

# 4. observe what the AGENT rendered (harness socket)
$ tmux -L cli-harness capture-pane -p -t codex | tail -30

# 5. assert GROUND TRUTH independently (mcp-target socket)
$ tmux -L mcp-target list-windows -t scratch                 # expect a 'probe' window
```

Step 5 is the whole point: it separates *"the agent said it worked"* from *"the
tool actually mutated tmux."* Layers 0 and 1 can be fooled by a hallucinated
success line; the target socket cannot.

## Two failure modes that waste the most time

**Approval gates hang naive harnesses.** The first tool use pops an approval
dialog. A driver that types the prompt and immediately waits for output waits
forever. Either pre-approve with the CLI's trust/approval flags (see the
matrix), or detect the approval prompt via `capture-pane` and answer its
keystroke *before* waiting for the result.

**Sleeping instead of waiting is flaky.** Poll `capture-pane` for a stable
completion marker — the prompt glyph returning, a known output line — rather than
`sleep N`. If you drive with libtmux-mcp itself, `wait_for_text` with a `stop`
list is the right primitive, but point it at the *harness* socket, never the
server under test.

**Submit as separate events.** Send the prompt text and `Enter` as two distinct
`send-keys` calls — then one Enter submits. Batching text and `Enter` into a
single `send-keys` is what leaves the prompt sitting unsent and makes it look
like you need a second Enter. Also mind PATH: a CLI launched inside a `-L`
harness pane runs in a non-login shell that lacks your mise/node/uv shims, so
`export` the needed bin dirs before launching it or you'll just get `command not
found`.

## High-value test: cancellation / teardown

Cancellation is invisible to the tool list, and it is exactly where tmux-MCP
servers leak. It is reachable at **Layer 0**, but only by one specific move — and
the obvious move is not it.

**Cancelling the client's asyncio task does not cancel the server.** Tee both
directions of the stdio pipe and run exactly that manoeuvre: the client sends
`initialize`, `notifications/initialized`, `tools/list`, `tools/call` — and
nothing more. `task.cancel()` unwinds the caller; the server never hears about
it and keeps working the abandoned call — a `wait_for_text` kept streaming
`notifications/progress` past the cancel and only stopped when the client
process exited and closed the pipe. A bare `task.cancel()` measures the *client*
giving up, not the server's reap path.

**Emit the notification explicitly.** FastMCP's client has
`await client.cancel(request_id)`, which puts a real `notifications/cancelled`
on the wire. With it the same run behaves as intended: the progress stream stops
at the instant of the cancel, and the server answers the in-flight `tools/call`
with `Request cancelled`. `call_tool` does not hand you the request id — **read
it off your capture**, do not count. The numbering depends on which requests the
session actually issues: measured runs put `tools/call` at both id 1 and id 2
depending on whether a `tools/list` went out first. Getting it wrong fails
silently — `client.cancel()` accepts an id that matches no in-flight request,
returns without error, and the call runs to completion. That is the difference
between measuring a child reaped in 0.1 s and one that lingers 17 s, from the
same script. Only with the notification landing on the *right* id is the
server's cancellation path under test.

Layer 0 remains the right layer for wire-contract questions, cancellation
included — it just has to send the frame. Reach for Layer 2 when the question is
the *client's* cancellation semantics (approval gate, streaming, what `Esc`
actually sends) rather than the server's reap path. Layer 2 reproduction:

1. Prompt the agent to call a `wait_for_text` that will never match (long timeout).
2. While it's mid-call — the TUI shows a "working / esc to interrupt" state — send
   `Escape` to that pane. `Esc` during the working phase cancels the in-flight
   tool call **while keeping the MCP server subprocess alive**, which is the exact
   client-cancellation the reap path guards; `Esc` *after* a turn finishes just
   enters edit-previous mode. (Killing the pane instead tears down the whole
   server, which tests a different thing.)
3. On the **mcp-target** socket, assert no orphaned `tmux`/child process survives
   (`pgrep -f mcp-target`) and the server stays healthy. Verified through Codex:
   the `wait_for_text` call returned `Error: interrupted`, the server survived,
   and no child leaked.

A server that reaps its child on cancel passes; one that orphans it hangs
interpreter shutdown. This is a behavioral difference you cannot see from schemas.

## Comparing two versions (trunk vs a branch)

Two worktrees, two target sockets, same prompt:

```console
$ git worktree add ../mcp-trunk  origin/main
$ git worktree add ../mcp-branch <branch>
# point one CLI's config at ../mcp-trunk  (LIBTMUX_SOCKET=mcp-trunk)
# point another at ../mcp-branch (LIBTMUX_SOCKET=mcp-branch)
```

Diff three things: the **tool surface** (`mcp list-tools` or a Layer-0 tool dump,
diffed), the **rendered agent behavior** for the same prompt (capture-pane
transcripts), and the **ground-truth socket state** after the run.

## Wiring a checkout into the CLIs: mcp_swap

`scripts/mcp_swap.py` rewrites each CLI's config to `uv --directory <repo> run
<entry>` and preserves existing env on replacement. It covers eight CLIs; the
two newest are not yet driven through this harness, so
`references/cli-matrix.md` has no verified row for them:

- **opencode** — `$XDG_CONFIG_HOME/opencode/opencode.jsonc`. JSONC, so
  comments survive a swap; the entry packs argv into one `command` array
  under a top-level `mcp` key, and its env table is spelled `environment`.
  A scalar `command` there is a decode error that stops opencode starting.
- **pi** — `~/.pi/agent/mcp.json`. pi ships no MCP client of its own; that
  file is read by the third-party `pi-mcp-adapter` extension, so a swap
  does nothing until it is installed. `detect` reports this.

```console
$ uv run scripts/mcp_swap.py detect                        # which CLIs are present
$ uv run scripts/mcp_swap.py doctor --server tmux           # effective environment + footguns
$ uv run scripts/mcp_swap.py status --server tmux           # current entries
$ uv run scripts/mcp_swap.py use-local --server tmux --env LIBTMUX_SOCKET=mcp-target --dry-run
$ uv run scripts/mcp_swap.py use-local --server tmux --env LIBTMUX_SOCKET=mcp-target
$ uv run scripts/mcp_swap.py revert --cli codex --dry-run   # scope it — a bare revert unwinds EVERY recorded swap
```

Run `doctor` first — it reports which server name each CLI points at (and warns
when the repo is registered under a name other than the derived default),
un-reverted swaps and orphaned backups, missing backups (revert would fail), and
auth-overriding env vars like `OPENAI_API_KEY`. `--env` injects the isolated
`LIBTMUX_SOCKET` into the entry at swap time, so you don't hand-edit the config.

The short version: pass `--server tmux` (the real registration key on this
machine is `tmux`, not the derived `libtmux`); mcp_swap preserves env but does
not add new keys, so inject `LIBTMUX_SOCKET=mcp-target` via each CLI's native
`mcp add ... -e ...` or a post-swap edit; `mcp_swap use-local` mutates the user's
real CLI configs, so dry-run first, record the pre-existing swap state, and
revert only what you swapped.

**Prefer no-config-write isolation for a test.** mcp_swap is for a swap you *want*
to persist. To just exercise a checkout, use each CLI's throwaway config-home /
project-config lever instead — `references/cli-matrix.md` gives the verified one
per CLI (codex `CODEX_HOME` or `-c` overrides, grok `GROK_HOME`, agy
`--gemini_dir` — config only, credentials do not follow it — cursor/gemini
project config, claude `--mcp-config --strict-mcp-config`). Five of the six
(codex, claude, cursor, grok, agy) were driven all the way to a real
model-issued tool call this way, with no swap state touched and every *MCP*
config file left as found; gemini is vendor-blocked (`IneligibleTierError`), not
harness-blocked. Isolating the MCP config does not stop a CLI writing its own
ambient state — a claude `-p` run still grows `~/.claude.json`, a gemini run
still appends `~/.gemini/projects.json` — so diff what you care about
afterwards. Note the machine may already carry an un-reverted swap (all CLIs
pointing at a local checkout); a bare `revert` would unwind that too, so read
the swap state file before running one.

**Copy credentials into a throwaway config home; never symlink them.** This is
the most expensive lesson here. A symlink isolates reads, not writes: agy
refreshed its OAuth token straight *through* the symlink and overwrote the
user's real `~/.gemini/antigravity-cli/antigravity-oauth-token` — the real
file's mtime moved. Repeated with a *copy*, the sandbox copy was refreshed and
the real file was untouched. Both directions were measured. Copy the credential
files in, and diff the originals when you're done.

## Cleanup checklist

```console
$ tmux -L cli-harness kill-server 2>/dev/null
$ tmux -L mcp-target  kill-server 2>/dev/null
```

Scratch sockets vanish with their server.

**Revert only a swap you made.** `revert` works off the swap state file
(`$XDG_STATE_HOME/libtmux-mcp-dev/swap/state.json`, i.e. under `~/.local/state/`
by default), not off whichever backup is newest on disk: each recorded
`(cli, scope)` entry names the config it swapped and the backup that restores
it, and revert writes that backup back, deletes it, and drops the entry. The
hazard is therefore *scope*, not staleness — a bare `revert` targets every
recorded entry for every CLI, so a swap an earlier session left behind is
unwound alongside yours. Name what you swapped with `--cli` (plus `--scope` for
Claude, which has two layers) and preview with `--dry-run`. State is keyed by
`(cli, scope)`, so two swaps of the same layer collapse into one entry — there
is no chain to unwind one step at a time. Re-run `status`/`doctor` and compare
against the state you recorded before starting. If you stayed on the
no-config-write path there is nothing to revert.

## When NOT to reach for the full harness

If the question is purely "is the tool surface correct?" or "does this result
shape parse?", stay at Layer 0 — booting six CLIs to answer a wire-contract
question is wasted effort. Escalate to Layers 1 and 2 only when the client's
discovery, approval, streaming, or cancellation behavior is what's actually in
doubt.

Files in this skill

  • SKILL.md15.5 KB
  • references/cli-matrix.md15.5 KB

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…