Skip to content
Back to skills

Distributed Engines Backends

ASecurity

"Plan and debug AReaL distributed engines, inference backends,

  • 247 stars
  • 0 votes
  • 0 copies
  • 1 view
  • Added September 8, 2026
businesspythongobashnodeapibackend

Works with

  • cli
  • api

Security analysis

A100/100

Pro scans all 5 files and shows the line behind each finding

Scanned September 8, 2026

npx -y skills add VectorSpaceLab/AREX-Skill --skill distributed-engines-backends --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Distributed Engines Backends?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for Distributed Engines Backends
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/vectorspacelab-distributed-engines-backends/badge)](https://www.skillsdirectory.com/skills/vectorspacelab-distributed-engines-backends)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
SKILL.md
---
name: distributed-engines-backends
description: "Plan and debug AReaL distributed engines, inference backends,
  allocation strings, weight sync, LoRA/FP8, and GPU/Ray/Slurm backend
  failures."
disable-model-invocation: true
metadata:
  disco-role: operating
license: Apache 2.0
---

# distributed-engines-backends

Use this sub-skill when the task is about AReaL backend selection or backend failure modes: FSDP2, Megatron, Archon, SGLang, vLLM, backend strings, GPU allocation, parallelism dimensions, weight-update modes, LoRA, FP8, CUDA/NCCL hangs, OOM, Ray/Slurm placement, or backend install variants.

## Route first

- If the user needs a full experiment command, config migration, algorithm recipe, or which training script/workflow to run, route to sibling sub-skill `post-training-experiments` and return here only for the backend fields.
- If the user needs to start/stop/register/debug AReaL v2 services, gateways, workers, sessions, or CLI process lifecycle, route to sibling sub-skill `services-cli-operations`; return here only for worker backend, allocation, and weight-sync constraints.
- If the user is authoring datasets, reward functions, `RolloutWorkflow`, or agent workflow code, route to sibling sub-skill `custom-data-rewards-workflows`.
- Never claim that CPU import or CLI help proves GPU backend behavior. It only proves import/config surface availability.

## Operating workflow

1. Establish the user's target roles (`rollout`, `actor`, optional `critic`, `ref`, `teacher`), cluster shape, backend strings, install variant, weight-update mode, LoRA/FP8 flags, and whether actor/rollout are separated or colocated.
2. Parse backend strings and compute GPU demand with [`scripts/check_backend_plan.py`](scripts/check_backend_plan.py):

   ```bash
   python scripts/check_backend_plan.py \
     rollout.backend=sglang:d2t4 actor.backend=fsdp:d8 \
     cluster.n_nodes=2 cluster.n_gpus_per_node=8 \
     actor.weight_update_mode=xccl
   ```

   Add `--probe-env` only for safe CUDA visibility/import facts; it is still not a backend runtime proof.
3. Use [`references/backend-planning.md`](references/backend-planning.md) for backend syntax, install variants, parallelism capability, Ray/Slurm placement, LoRA, and FP8 planning.
4. Use [`references/engine-api-and-weight-sync.md`](references/engine-api-and-weight-sync.md) for train/inference engine contracts, generation request behavior, weight versioning, and `disk`/`xccl`/`awex` update modes.
5. Use [`references/troubleshooting.md`](references/troubleshooting.md) to debug parse/config errors, optional dependency issues, CUDA/NCCL hangs, OOM, LoRA/FP8 failures, placement mistakes, and checkpoint/recovery mismatches.

## Safe outputs to give users

- Backend field diffs or config snippets such as `rollout.backend=sglang:d2t4`, `actor.backend=megatron:(attn:d1p4t2c2|ffn:d1p4t1e4)`, `actor.weight_update_mode=disk`, or `actor.megatron.bridge_type=megatron-bridge`.
- A GPU-demand calculation and assumptions: separated roles sum GPU worlds; colocated roles require explicit placement and usually matching planned worlds.
- A validation checklist and one safe checker command.
- A clear skip/block statement for GPU-only, multi-node, model-download, service, or credentialed validation that was not actually run.

## Hard stops

Do not start training, launch SGLang/vLLM/Ray/Slurm services, download models/datasets, run native repo tests, mutate driver/CUDA stacks, or guess cluster-specific placement. Ask for cluster/model/runtime decisions when they are required to choose between incompatible backends or install variants.

Files in this skill

  • SKILL.md3.5 KB
  • references/backend-planning.md14.9 KB
  • references/engine-api-and-weight-sync.md13.7 KB
  • references/troubleshooting.md15.5 KB
  • scripts/check_backend_plan.py25.2 KB

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…