Skip to content
Back to skills

Text Generation Webui Expert

ASecurity

Run oobabooga's textgen with multiple backends, OpenAI/Anthropic-compatible API, characters, training tab, and portable installer. Use when deploying, running, or configuring local LLM inference with text generation webui.

  • 8 stars
  • 0 votes
  • 0 copies
  • 1 view
  • Added September 8, 2026
developmentgobashtestinggitapibackend

Works with

  • cli
  • api

Security analysis

A96/100
  • mediumUses curl or wget to download content

Pro shows the line behind each finding and how to fix it

Scanned September 8, 2026

npx -y skills add anubhavg-icpl/vibe --skill text-generation-webui-expert --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Text Generation Webui Expert?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for Text Generation Webui Expert
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/anubhavg-icpl-text-generation-webui-expert/badge)](https://www.skillsdirectory.com/skills/anubhavg-icpl-text-generation-webui-expert)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
SKILL.md
---
name: text-generation-webui-expert
description: "Run oobabooga's textgen with multiple backends, OpenAI/Anthropic-compatible API, characters, training tab, and portable installer. Use when deploying, running, or configuring local LLM inference with text generation webui."
license: CC-BY-NC-SA-4.0
metadata:
  risk: unknown
  source: community
  kind: mode
  category: local-llm
  tags: [local-llm, oobabooga, text-generation-webui, textgen, openai-api, characters, training, lora]
---

# oobabooga text-generation-webui Expert Mode

You are an expert in oobabooga's `textgen` (formerly text-generation-webui), the OG local LLM interface that supports llama.cpp, ik_llama.cpp, Transformers, ExLlamaV3, and TensorRT-LLM behind one UI. You configure backends, run the OpenAI/Anthropic-compatible API, build character cards, train LoRAs, and deploy the portable Electron build for non-developers.

## Core Capabilities

- Install via the portable installer (bundles llama.cpp binaries and Electron)
- Switch backends without restart: llama.cpp, ik_llama.cpp, Transformers, ExLlamaV3, TensorRT-LLM
- Enable OpenAI / Anthropic compatible API (`--api`, `--extensions openai`)
- Configure Cloudflare public URL with `--public-api` (testing only)
- Build/load character cards (JSON or PNG with embedded metadata)
- Train LoRAs from the Training tab (multi-turn chat or raw text)
- Tune sampler / generation defaults via CLI flags or `CMD_FLAGS.txt`
- Use multimodal & tool calling on supported backends
- Tensor-parallel llama.cpp via `--split-mode tensor`

## Approach

1. **Install the portable build** for desktops. Source/Conda only when extending.
2. **Persist flags in `user_data/CMD_FLAGS.txt`** — survives upgrades.
3. **Pick one model + backend per process**; switch via UI when iterating.
4. **Enable `--api`** for client integration; pin a port with `--api-port`.
5. **For LoRA training**, use the Training tab with `text-and-format` for chat or `Raw text file` for pretraining-style.

## Key Patterns

### Portable launch (Linux/macOS/Windows)

```bash
# After downloading and extracting the portable release
./textgen --listen --api --auto-launch
```

`textgen.bat` on Windows. `--listen` exposes on `0.0.0.0`.

### CMD_FLAGS.txt persistent flags

`user_data/CMD_FLAGS.txt`:

```text
--api
--api-port 5000
--listen
--extensions openai
--temperature 0.7
--top-k 40
--repetition-penalty 1.1
--chat-template-file user_data/templates/llama3.jinja
```

### Launch with a specific GGUF + llama.cpp backend

```bash
./textgen \
  --model qwen2.5-7b-instruct-q4_k_m.gguf \
  --loader llama.cpp \
  --n-gpu-layers 99 \
  --n_ctx 16384 \
  --split-mode tensor \
  --api --api-port 5000
```

### ExLlamaV3 backend

```bash
./textgen \
  --model Llama-3.1-8B-EXL3-4.0bpw \
  --loader ExLlamaV3 \
  --api
```

### Hit the OpenAI-compatible API

```bash
curl http://localhost:5000/v1/chat/completions \
  -H "Authorization: Bearer YOUR_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "qwen2.5-7b-instruct",
    "messages": [{"role":"user","content":"Hi"}]
  }'
```

API key is set via `--api-key` (or env). Without it, the API is open if reachable.

### Headless mode (no Electron window)

```bash
./textgen --nowebui --api --listen
```

### Character card (chat-instruct mode)

Place a `.json` or character `.png` (with embedded card) under `user_data/characters/`:

```json
{
  "name": "Ada",
  "context": "Ada is a curious researcher who explains things briefly.",
  "greeting": "Hi! What's on your mind?",
  "example_dialogue": "User: ...\nAda: ..."
}
```

Select in the Chat tab; persona is prepended to every message.

### Train a LoRA (Training tab)

1. Models tab → load a base model with the Transformers loader (LoRA training requires Transformers).
2. Training tab → select dataset:
   - **Formatted dataset** (JSON) for chat
   - **Raw text file** for pretraining-style
3. Set `LoRA Rank`, `LoRA Alpha`, `Cutoff Length`, `Learning Rate`.
4. Start training; resume supported via the same name.

Output saves under `user_data/loras/<name>`. Apply at load time: select the LoRA in the Models tab.

### Cloudflare quick-share (development only)

```bash
./textgen --api --public-api
```

Prints a `https://....trycloudflare.com` URL. Do not use for production.

## Common Pitfalls

- **Public API accidentally on `0.0.0.0` with no `--api-key`** — anyone on the network can hit it. Always set a key or bind to `127.0.0.1`.
- **`--public-api`** is convenient but exposes the model publicly via Cloudflare quick tunnel; not for prod.
- **Wrong loader for the model file** — GGUF needs llama.cpp/ik_llama.cpp; safetensors with `EXL3` metadata needs ExLlamaV3.
- **Training tab requires Transformers loader** — try with llama.cpp loader and you cannot train.
- **Old UI muscle memory** — the project renamed to `textgen` in v4.0 and switched to portable Electron; old `start_*.sh` scripts are gone.
- **Chat template missing** → outputs leak role tokens. Pass `--chat-template-file` or use a model with the template baked into metadata.
- **Disk filling on Windows** — the portable bundle is large; pin to a SSD.

## Hardware/Resource Sizing

- Same backend rules as upstream: llama.cpp `-ngl`, ExLlamaV3 EXL3 bpw, Transformers FP16/BF16
- 8GB GPU: 7B GGUF Q4 / 8B EXL3 4.0
- 24GB GPU: 30B GGUF Q4 / 70B EXL3 with tensor split
- LoRA training: 7B Transformers fp16 fits on 24GB with rank 16 on cutoff 1024

## When to Use This Mode

- Single-user power-tinkerer who wants every backend in one UI
- Persona-driven chat with character cards
- Local LoRA training from a UI
- Quick Cloudflare share of a local model for collaborators (dev only)
- Use **llama-cpp-server-expert** / **vllm-local-deploy-expert** for production serving
- Use **lm-studio-expert** or **jan-ai-expert** for a more polished UX

## Sources

- [textgen (oobabooga) GitHub](https://github.com/oobabooga/textgen)
- [textgen wiki home](https://github.com/oobabooga/text-generation-webui/wiki)
- [textgen OpenAI API wiki](https://github.com/oobabooga/text-generation-webui/wiki/12-%E2%80%90-OpenAI-API)
- [textgen Training Tab wiki](https://github.com/oobabooga/text-generation-webui/wiki/05-%E2%80%90-Training-Tab)
- [textgen v4.0 release](https://newreleases.io/project/github/oobabooga/textgen/release/v4.0)
- [textgen portable getting started (DeepWiki)](https://deepwiki.com/oobabooga/textgen/1.1-getting-started-and-installation)

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…