Skip to content
Back to skills

Ray Distributed Trainer

ASecurity

Distributed computing skill using Ray for parallel training, hyperparameter search, and resource management.

  • 1,760 stars
  • 0 votes
  • 0 copies
  • 1 view
  • Added September 2, 2026
ai-agentsjavascriptjavabashnode

Security analysis

A100/100

Pro scans all 2 files and shows the line behind each finding

Scanned September 2, 2026

npx -y skills add a5c-ai/babysitter --skill ray-distributed-trainer --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Ray Distributed Trainer?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for Ray Distributed Trainer
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/a5c-ai-ray-distributed-trainer-babysitter/badge)](https://www.skillsdirectory.com/skills/a5c-ai-ray-distributed-trainer-babysitter)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
SKILL.md
---
name: ray-distributed-trainer
description: Distributed computing skill using Ray for parallel training, hyperparameter search, and resource management.
allowed-tools:
  - Read
  - Write
  - Bash
  - Glob
  - Grep
graph:
  domains: [domain:data-science]
  specializations: [specialization:data-science-ml]
  skillAreas: [skill-area:machine-learning-frameworks, skill-area:hyperparameter-tuning-experiment-management]
  roles: [role:ml-engineer, role:ml-ops-engineer]
  workflows: [workflow:ml-model-lifecycle]

---

# ray-distributed-trainer

## Overview

Distributed computing skill using Ray for parallel training, hyperparameter search, and resource management across clusters.

## Capabilities

- Ray Train for distributed training
- Ray Tune for hyperparameter search at scale
- Cluster resource management
- Fault tolerance and checkpointing
- Actor-based parallelism
- Integration with PyTorch and TensorFlow
- Elastic training support
- Multi-node orchestration

## Target Processes

- Distributed Training Orchestration
- AutoML Pipeline Orchestration
- Model Training Pipeline

## Tools and Libraries

- Ray
- Ray Train
- Ray Tune
- Ray Cluster

## Input Schema

```json
{
  "type": "object",
  "required": ["mode", "config"],
  "properties": {
    "mode": {
      "type": "string",
      "enum": ["train", "tune", "cluster"],
      "description": "Ray operation mode"
    },
    "config": {
      "type": "object",
      "properties": {
        "numWorkers": { "type": "integer" },
        "useGpu": { "type": "boolean" },
        "resourcesPerWorker": {
          "type": "object",
          "properties": {
            "cpu": { "type": "number" },
            "gpu": { "type": "number" }
          }
        }
      }
    },
    "trainConfig": {
      "type": "object",
      "properties": {
        "trainerPath": { "type": "string" },
        "framework": { "type": "string", "enum": ["pytorch", "tensorflow", "xgboost"] },
        "scalingConfig": { "type": "object" }
      }
    },
    "tuneConfig": {
      "type": "object",
      "properties": {
        "searchSpace": { "type": "object" },
        "scheduler": { "type": "string" },
        "numSamples": { "type": "integer" },
        "metric": { "type": "string" },
        "mode": { "type": "string", "enum": ["min", "max"] }
      }
    }
  }
}
```

## Output Schema

```json
{
  "type": "object",
  "required": ["status", "results"],
  "properties": {
    "status": {
      "type": "string",
      "enum": ["success", "error", "partial"]
    },
    "results": {
      "type": "object",
      "properties": {
        "bestConfig": { "type": "object" },
        "bestMetric": { "type": "number" },
        "numTrials": { "type": "integer" },
        "completedTrials": { "type": "integer" }
      }
    },
    "checkpointPath": {
      "type": "string"
    },
    "clusterStatus": {
      "type": "object",
      "properties": {
        "numNodes": { "type": "integer" },
        "totalCpu": { "type": "number" },
        "totalGpu": { "type": "number" }
      }
    },
    "trainingTime": {
      "type": "number"
    }
  }
}
```

## Usage Example

```javascript
{
  kind: 'skill',
  title: 'Distributed hyperparameter tuning',
  skill: {
    name: 'ray-distributed-trainer',
    context: {
      mode: 'tune',
      config: {
        numWorkers: 4,
        useGpu: true,
        resourcesPerWorker: { cpu: 2, gpu: 1 }
      },
      tuneConfig: {
        searchSpace: {
          lr: { type: 'loguniform', min: 1e-5, max: 1e-1 },
          batchSize: { type: 'choice', values: [16, 32, 64] }
        },
        scheduler: 'asha',
        numSamples: 100,
        metric: 'val_loss',
        mode: 'min'
      }
    }
  }
}
```

Files in this skill

  • README.md1.1 KB
  • SKILL.md3.6 KB

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…