Skip to content
Back to skills

Sglang Expert

BSecurity

Serve LLMs with SGLang's RadixAttention, structured outputs (compressed FSM), tensor parallel, DP-attention, and PD disaggregation. Use when deploying, running, or configuring local LLM inference with sglang.

  • 8 stars
  • 0 votes
  • 0 copies
  • 1 view
  • Added September 8, 2026
ai-agentspythongobashnodegitfrontendbackendperformance

Works with

  • cli

Security analysis

B88/100
  • mediumUses curl or wget to download content
  • mediumInstalls packages at runtime which could introduce malicious dependencies

Pro shows the line behind each finding and how to fix it

Scanned September 8, 2026

npx -y skills add anubhavg-icpl/vibe --skill sglang-expert --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Sglang Expert?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for Sglang Expert
[![Security: B — Skills Directory](https://www.skillsdirectory.com/api/skills/anubhavg-icpl-sglang-expert/badge)](https://www.skillsdirectory.com/skills/anubhavg-icpl-sglang-expert)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
SKILL.md
---
name: sglang-expert
description: "Serve LLMs with SGLang's RadixAttention, structured outputs (compressed FSM), tensor parallel, DP-attention, and PD disaggregation. Use when deploying, running, or configuring local LLM inference with sglang."
license: CC-BY-NC-SA-4.0
metadata:
  risk: unknown
  source: community
  kind: mode
  category: local-llm
  tags: [local-llm, sglang, radixattention, structured-outputs, tensor-parallel, dp-attention, pd-disaggregation]
---

# SGLang Expert Mode

You are an SGLang expert. SGLang is a high-performance serving framework with **RadixAttention** for KV-prefix reuse across requests, **compressed-FSM** for fast structured outputs (JSON / regex / function calling), tensor parallelism, DP-attention, and prefill-decode (PD) disaggregation. You launch `sglang.launch_server` for OpenAI-compat serving and use the Python frontend DSL for multi-step LLM programs.

## Core Capabilities

- Launch the server: `python -m sglang.launch_server --model-path ...`
- Tune `--tp` (tensor parallel), `--dp-size` + `--enable-dp-attention`
- PD disaggregation: separate prefill / decode nodes for higher throughput
- RadixAttention auto-shares KV across requests with shared prefixes
- Structured outputs: JSON, regex, EBNF via compressed finite state machine
- OpenAI-compat `/v1/chat/completions` + native generate
- Reasoning models: `--reasoning-parser` for Qwen3 / DeepSeek-R1 / GLM
- Day-1 support for new models (DeepSeek V3/R1, Qwen3, Llama 4 Scout)

## Approach

1. **Default to TP per visible GPU** for large models; single GPU for ≤7B.
2. **Always enable RadixAttention** (default) for multi-turn chat — it gives huge speedups when system prompts repeat.
3. **Use compressed-FSM JSON** instead of post-hoc parsing. Up to 3× decoding speedup vs unconstrained.
4. **PD-disaggregation** only at scale (≥8 GPUs); over-engineered below.
5. **DP-attention** when you have many GPUs and want better KV utilization.
6. **Pin a release** (`pip install sglang==<ver>`) — the engine moves quickly.

## Key Patterns

### Install + launch

```bash
pip install -U "sglang[all]"

python -m sglang.launch_server \
  --model-path Qwen/Qwen2.5-7B-Instruct \
  --host 0.0.0.0 --port 30000 \
  --mem-fraction-static 0.85
```

### Tensor parallel across 4 GPUs

```bash
python -m sglang.launch_server \
  --model-path meta-llama/Llama-3.1-70B-Instruct \
  --tp 4 \
  --port 30000 \
  --context-length 16384
```

### Qwen3 with reasoning parser

```bash
python -m sglang.launch_server \
  --model-path Qwen/Qwen3-32B \
  --tp 4 \
  --reasoning-parser qwen3 \
  --mem-fraction-static 0.85 \
  --context-length 65536
```

### DP attention (data parallel attention) — large multi-GPU

```bash
python -m sglang.launch_server \
  --model-path deepseek-ai/DeepSeek-V3 \
  --tp 8 --dp-size 8 \
  --enable-dp-attention
```

`--dp-size` must be > 1 and `--enable-dp-attention` must be present.

### PD disaggregation (separate prefill / decode)

Prefill node:

```bash
python -m sglang.launch_server \
  --model-path Llama-3.1-70B-Instruct \
  --disaggregation-mode prefill --tp 4
```

Decode node:

```bash
python -m sglang.launch_server \
  --model-path Llama-3.1-70B-Instruct \
  --disaggregation-mode decode --tp 4 --dp 4 --enable-dp-attention
```

### OpenAI-compat call

```bash
curl http://localhost:30000/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "qwen",
    "messages": [{"role":"user","content":"Hi"}]
  }'
```

### JSON structured output via compressed FSM

```bash
curl http://localhost:30000/v1/chat/completions \
  -d '{
    "model": "qwen",
    "messages": [{"role":"user","content":"Pick a color"}],
    "response_format": {
      "type": "json_schema",
      "json_schema": {
        "name": "color",
        "schema": {
          "type":"object",
          "properties":{"color":{"type":"string"}},
          "required":["color"]
        }
      }
    }
  }'
```

### Regex-constrained generation

```python
from sglang import RuntimeEndpoint, function, gen
endpoint = RuntimeEndpoint("http://localhost:30000")

@function
def phone(s):
    s += "Phone: " + gen("p", regex=r"\d{3}-\d{3}-\d{4}")

state = phone.run(backend=endpoint)
print(state["p"])
```

### SGLang frontend DSL — multi-step program

```python
from sglang import function, gen, set_default_backend, RuntimeEndpoint

set_default_backend(RuntimeEndpoint("http://localhost:30000"))

@function
def multi_turn(s, question):
    s += "User: " + question + "\nAssistant: "
    s += gen("answer", max_tokens=128)
    s += "\n\nRate that answer 1-5: "
    s += gen("rating", regex=r"[1-5]")

state = multi_turn.run(question="What's RAG?")
print(state["answer"], state["rating"])
```

## Common Pitfalls

- **`--tp` larger than visible GPUs** → init failure.
- **`--dp-size 1` with `--enable-dp-attention`** → flag silently ignored; needs > 1.
- **Forgetting `--reasoning-parser`** for Qwen3 / DeepSeek-R1 → reasoning trace leaks into the answer.
- **`--mem-fraction-static` too high** → OOM during prefill burst. 0.85 is a safe ceiling.
- **PD-disaggregation overkill** for small clusters — adds operational complexity for marginal gains under 8 GPUs.
- **Using the OpenAI client without `model` matching the loaded model id** — SGLang tolerates anything, but logs are confusing.
- **Mixing `response_format=json_object` with no schema** → unconstrained JSON, no FSM speedup. Use `json_schema`.
- **RadixAttention misses** when system prompts vary slightly — small differences break the prefix cache.

## Hardware/Resource Sizing

- 1× 24GB consumer: 7-13B FP16, 30-34B AWQ
- 2× 24GB: 70B AWQ TP=2
- 4-8× A100/H100: 70B FP16/FP8 long-ctx, MoE models
- 8+ × H100 with PD disaggregation: DeepSeek V3/R1-class workloads
- For long shared system prompts (RAG), RadixAttention can multiply effective throughput

## When to Use This Mode

- Workloads with **shared prefixes** (RAG, agent tool calling, system prompts) — RadixAttention dominates here
- Strict structured outputs at high throughput (JSON/regex) — compressed-FSM
- Reasoning models with `<think>` tags — first-class parser support
- Very large multi-GPU deploys with PD disaggregation
- Compare: **vllm-local-deploy-expert** for general-purpose throughput; **tgi-huggingface-expert** for HF-native; **llama-cpp-server-expert** for static binary

## Sources

- [SGLang GitHub](https://github.com/sgl-project/sglang)
- [SGLang docs](https://sgl-project.github.io/)
- [SGLang server arguments](https://sgl-project.github.io/advanced_features/server_arguments.html)
- [DP, DPA, SGLang DP Router](https://sgl-project.github.io/advanced_features/dp_dpa_smg_guide.html)
- [PD disaggregation](https://docs.sglang.io/docs/advanced_features/pd_disaggregation)
- [SGLang RadixAttention paper / blog (LMSYS)](https://www.lmsys.org/blog/2024-01-17-sglang/)
- [SGLang on Qwen](https://qwen.readthedocs.io/en/latest/deployment/sglang.html)

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…