Skip to content
Back to skills

Language Models

ASecurity

"Routes AXLearn GPT-family trainer catalogs, tokenizer variants,

  • 247 stars
  • 0 votes
  • 0 copies
  • 1 view
  • Added September 8, 2026
toolspythonbashgcp

Works with

  • cli

Security analysis

A100/100

Pro scans all 4 files and shows the line behind each finding

Scanned September 8, 2026

npx -y skills add VectorSpaceLab/AREX-Skill --skill language-models --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Language Models?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for Language Models
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/vectorspacelab-language-models/badge)](https://www.skillsdirectory.com/skills/vectorspacelab-language-models)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
SKILL.md
---
name: language-models
description: "Routes AXLearn GPT-family trainer catalogs, tokenizer variants,
  MoE configs, and flash-attention workflows."
disable-model-invocation: true
metadata:
  disco-role: operating
license: Apache 2.0
---

# language-models

Use this sub-skill for AXLearn's decoder-only language-model catalogs and related tokenizer/model-family helpers.

Typical triggers:

- GPT, C4, Pajama, Fuji, Gala, Honeycrisp, Qwen, or MoE model names.
- Long-context, flash-attention, RoPE, ALiBi, or mesh-shape questions in `axlearn.experiments.text.gpt`.
- Tokenizer files such as `bpe_32k.json`, `bpe_128k.json`, `Llama-3-tokenizer.json`, or Fuji v3 vocabulary helpers.
- Queries about `tokamax`, `qwix`, `FlashAttention`, or the model-specific trainer catalogs.

If the user is only asking about the shared trainer runtime, `config_for_function`, or fake-data smoke checks, use `../training-core/` first.
If the user is asking about cloud launch or GCP job execution, use `../cli-cloud/`.

## What to read

- `references/overview.md` for the GPT catalog structure and major model families.
- `references/troubleshooting.md` for optional dependency and tokenizer-path failures.
- `scripts/inspect_gpt_configs.py` for a safe config-inspection helper.

## Families covered

- `c4_trainer` for C4-based training catalogs.
- `fuji`, `gala`, `honeycrisp`, `gspmd`, and `qwen` for model-family builders and trainer variants.
- `deterministic_trainer` and the Pajama configs for dataset-specific catalog variants.
- `vocabulary_fuji_v3` for tokenizer compatibility and Llama-3-style tokenizer files.
- `gala_sigmoid` for sigmoid-attention-related config manipulation.

## Typical workflows

### Inspect exported config names

Use the bundled helper to list the named trainer configs for a module and to inspect one resolved config:

```bash
python scripts/inspect_gpt_configs.py --module axlearn.experiments.text.gpt.gala --config 7B
```

### Understand tokenizer wiring

The GPT helpers read tokenizer files from the configured data directory. When `DATA_DIR=FAKE`, they fall back to the packaged repository data under `axlearn/data/tokenizers/`.

### Route around optional dependencies

Some GPT-family modules pull in extra MoE or flash-attention dependencies at import time. If the import fails, check `references/troubleshooting.md` before assuming the catalog is unavailable.

## Decision points

- Use this sub-skill when the user names a concrete GPT-family architecture or tokenizer file.
- Keep reusable trainer mechanics in `training-core`.
- Do not send vision or ASR questions here just because they also use trainer configs.

Files in this skill

  • SKILL.md2.6 KB
  • references/overview.md3.3 KB
  • references/troubleshooting.md1.9 KB
  • scripts/inspect_gpt_configs.py3.8 KB

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…