Skip to content
Back to skills

Baichuan2

ASecurity

"Route Baichuan2 chat inference, deployment/quantization, and

  • 247 stars
  • 0 votes
  • 0 copies
  • 1 view
  • Added September 8, 2026
toolspythonrustbashflaskapibackend

Works with

  • terminal
  • cli
  • api

Security analysis

A100/100

Pro scans all 20 files and shows the line behind each finding

Scanned September 8, 2026

npx -y skills add VectorSpaceLab/AREX-Skill --skill baichuan2 --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Baichuan2?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for Baichuan2
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/vectorspacelab-baichuan2/badge)](https://www.skillsdirectory.com/skills/vectorspacelab-baichuan2)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
SKILL.md
---
name: baichuan2
description: "Route Baichuan2 chat inference, deployment/quantization, and
  fine-tuning workflows."
disable-model-invocation: true
metadata:
  disco-role: operating
license: Apache 2.0
---

# Baichuan2

Use this repo skill when a task is about Baichuan2 model selection, chat/base inference, local chat demos, OpenAI-compatible serving, quantized deployment, CPU deployment, checkpoint conversion, or supervised fine-tuning.

Baichuan2 is a model-family repository rather than an importable Python package. The operating guidance is therefore organized around public workflows and bundled helpers, not a package module tree.

## Start here

1. If the task may depend on repo freshness, read [`references/repo-provenance.md`](references/repo-provenance.md).
2. If the user has not chosen a checkpoint, read [`references/model-overview.md`](references/model-overview.md).
3. Prepare the smallest environment for the requested workflow using [`references/installation.md`](references/installation.md).
4. Run a no-weight dependency check before loading 7B/13B weights:

```bash
python scripts/check_baichuan2_env.py --workflow inference
python scripts/check_baichuan2_env.py --workflow all --require-cuda
```

Use `--workflow inference`, `deployment`, `fine-tuning`, or `all` to match the route. These checks do not download model weights.

## Route by user request

| User request signals | Read | Why |
| --- | --- | --- |
| `model.chat`, Python inference, Base-model generation, terminal demo, Streamlit web UI, OpenAI-compatible `/v1/chat/completions` | [`sub-skills/inference/SKILL.md`](sub-skills/inference/SKILL.md) | Owns Baichuan2 Chat/Base inference, CLI/web demos, and the Flask chat-completions helper. |
| 4-bit/8-bit quantization, BitsAndBytes, CPU-only loading, `load_in_8bit`, `quantize(4)`, `lm_head.weight`, Baichuan1 optimization migration | [`sub-skills/deployment/SKILL.md`](sub-skills/deployment/SKILL.md) | Owns memory-reduction, CPU deployment, and checkpoint-conversion workflows. |
| supervised fine-tuning, DeepSpeed, LoRA, `fine-tune.py`, hostfile, `W_pack`, `ds_config`, training data schema | [`sub-skills/fine-tuning/SKILL.md`](sub-skills/fine-tuning/SKILL.md) | Owns SFT data validation, ZeRO-3 launch planning, LoRA, and post-training loading. |
| installation/import/backend failure before a route is clear | [`references/troubleshooting.md`](references/troubleshooting.md) | Covers shared CUDA, dependency, model-access, and optional-extra problems. |
| license, commercial-use constraints, or citation | [`references/license-and-citation.md`](references/license-and-citation.md) | Distills the repo's license/citation section. |

## Core operating facts

- Released model ids include `baichuan-inc/Baichuan2-7B-Base`, `baichuan-inc/Baichuan2-7B-Chat`, `baichuan-inc/Baichuan2-13B-Base`, `baichuan-inc/Baichuan2-13B-Chat`, and published 4-bit Chat variants.
- Chat checkpoints expose `model.chat(tokenizer, messages)` through Hugging Face `trust_remote_code=True`. Base checkpoints use `model.generate(...)` instead.
- The repository demos are GPU-oriented by default. CPU loading is a separate float32 deployment branch and is expected to be slow.
- Quantization and training are CUDA-sensitive. Do not use a CPU import check as proof that those workflows work.
- The bundled scripts intentionally include `--help`, `--dry-run`, or validation modes so future agents can inspect plans without downloading large weights.

## Do not do this

- Do not tell the user to open or run original checkout scripts; use the bundled helpers and references in this skill.
- Do not force Base checkpoints through Chat-only CLI/web/API wrappers.
- Do not add `device_map="auto"` to online quantization; the deployment sub-skill explains why.
- Do not launch DeepSpeed training until the data validator and dry-run plan pass.
- Do not import this skill into live repo-skill storage unless a later verification/import workflow is explicitly approved. This construction run was requested as **not import**.

Files in this skill

  • SKILL.md3.9 KB
  • references/installation.md2.9 KB
  • references/license-and-citation.md1.4 KB
  • references/model-overview.md2.5 KB
  • references/repo-provenance.md1.4 KB
  • references/repo-routing-metadata.json453 B
  • references/troubleshooting.md3.1 KB
  • scripts/check_baichuan2_env.py6.5 KB
  • sub-skills/deployment/SKILL.md3.6 KB
  • sub-skills/deployment/references/conversion.md3.3 KB
  • sub-skills/deployment/references/cpu-deployment.md2.6 KB
  • sub-skills/deployment/references/quantization.md5.1 KB
  • sub-skills/deployment/references/troubleshooting.md4.7 KB
  • sub-skills/deployment/scripts/normalize_lm_head.py7.3 KB
  • sub-skills/deployment/scripts/quantize_model.py10 KB
  • sub-skills/fine-tuning/SKILL.md3.1 KB
  • sub-skills/fine-tuning/references/data-format.md4.4 KB
  • sub-skills/fine-tuning/references/deepspeed-config.md4.4 KB
  • sub-skills/fine-tuning/references/installation.md3.4 KB
  • sub-skills/fine-tuning/references/troubleshooting.md5.3 KB

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…