Skip to content
Back to skills

Choose Model Lab Workflow

ASecurity

Route language-model training, data, evaluation, checkpoint, representation, steering, ablation, jailbreak, tool-calling, Apple-runtime, and benchmark requests. Use when the primary workflow or Socket owner is unclear.

  • 7 stars
  • 0 votes
  • 0 copies
  • 2 views
  • Added September 5, 2026
ai-agentspythonswiftsecurity

Security analysis

A100/100

Pro scans all 2 files and shows the line behind each finding

Scanned September 5, 2026

npx -y skills add gaelic-ghost/socket --skill choose-model-lab-workflow --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Choose Model Lab Workflow?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for Choose Model Lab Workflow
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/gaelic-ghost-choose-model-lab-workflow/badge)](https://www.skillsdirectory.com/skills/gaelic-ghost-choose-model-lab-workflow)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
SKILL.md
---
name: choose-model-lab-workflow
description: Route language-model training, data, evaluation, checkpoint, representation, steering, ablation, jailbreak, tool-calling, Apple-runtime, and benchmark requests. Use when the primary workflow or Socket owner is unclear.
---

# Choose Model Lab Workflow

## Outcome

Select one primary workflow, name any supporting workflows, and make the evidence boundary explicit before work begins.

## Route The Request

| Requested outcome | Primary skill |
| --- | --- |
| Define a hypothesis, controls, budget, and artifacts | `design-model-experiment` |
| Curate, transform, split, or document examples | `prepare-language-model-dataset` |
| Run SFT, LoRA, QLoRA, or a full parameter update | `fine-tune-language-model` |
| Measure capability, behavior, quality, or safety | `evaluate-language-model` |
| Decide which checkpoint is better and why | `compare-model-checkpoints` |
| Choose Core AI, Core ML, MLX, ExecuTorch, or Foundation Models | `choose-apple-model-runtime` |
| Locate or test internal representations | `research-model-representations` |
| Apply activation or weight-space behavior steering | `steer-language-model-behavior` |
| Remove or suppress a refusal direction | `ablate-refusal-representations` |
| Measure jailbreak or prompt-injection robustness | `evaluate-jailbreak-resilience` |
| Measure tool selection, arguments, execution, or recovery | `evaluate-tool-calling-model` |
| Compare latency, memory, energy, throughput, or artifact size | `benchmark-model-runtime` |
| Run preference optimization such as DPO/ORPO | Keep the experiment and eval here; use the supported TRL workflow through `fine-tune-language-model` until a stable standalone skill is earned |
| Pretrain or continue pretraining a foundation model | Do not collapse it into fine-tuning; define the distributed/corpus contract and treat `train-language-model` as a deferred skill candidate |
| Merge adapters or quantize/package an artifact | Use `compare-model-checkpoints` around the exact transformation; use the project-native tool and evaluate the deployable output |
| Evaluate an agent skill, plugin, or host harness rather than a model protocol | Hand off to `agent-engineering-skills` and `agent-portability-skills` |

## Respect Ownership Boundaries

- Use `cloud-inference-skills` for provider, GPU, endpoint, cost, and teardown decisions.
- Use `python-skills` for Python packaging and environment repair.
- Use `apple-dev-skills` for Swift/Xcode application integration after the runtime has been chosen.
- Use `agent-engineering-skills` for evaluating agent skills, prompts, or plugin packages rather than model checkpoints.
- Use `cybersecurity-skills` when an authorized evaluation targets a deployed system instead of a model artifact.

## Return A Routing Contract

State:

1. the primary skill;
2. supporting skills in execution order;
3. the controlled variable;
4. the artifact or metric that proves completion;
5. any paid compute, data-access, or deployment authorization required.

Do not silently turn research planning into a paid run, model publication, or production-system test.

Files in this skill

  • SKILL.md3.1 KB
  • agents/openai.yaml272 B

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…