Skip to content
Back to skills

Training Budget Estimator

ASecurity

Estimate (N, D, hours, GPU count) for a new transformer training run given compute budget and deployment constraints. Use when you need help with training budget estimator.

  • 8 stars
  • 0 votes
  • 0 copies
  • 1 view
  • Added September 8, 2026
developmentaws

Security analysis

A100/100

Scanned September 8, 2026

npx -y skills add anubhavg-icpl/vibe --skill training-budget-estimator --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Training Budget Estimator?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for Training Budget Estimator
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/anubhavg-icpl-training-budget-estimator/badge)](https://www.skillsdirectory.com/skills/anubhavg-icpl-training-budget-estimator)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
SKILL.md
---
name: training-budget-estimator
description: Estimate (N, D, hours, GPU count) for a new transformer training run given compute budget and deployment constraints. Use when you need help with training budget estimator.
license: CC-BY-NC-SA-4.0
phase: 7
lesson: 13
metadata:
  version: 1.0.0
  tags: [scaling-laws, training, chinchilla]
---

Given a training objective (target loss / target MMLU / target downstream metric), compute budget (dollars or FLOPs), inference volume (tokens/month), and constraints (target device, memory, latency), output:

1. Compute regime. Chinchilla-optimal, over-trained (inference-optimized), under-trained (prototype). One-sentence reason tied to inference volume.
2. N and D. Concrete values. Print the `D/N` ratio. If over-trained, note the loss penalty vs Chinchilla-optimal.
3. Training wall-clock. Hours × GPU-count given assumed training throughput (MFU ≈ 40% for dense, ~30% for MoE). Budget the precision (bf16 / fp8) and optimizer (AdamW / Muon).
4. Data sources. Named corpora or synthetic budget. Flag if the required `D` exceeds available high-quality tokens.
5. Risk note. One specific failure mode: data contamination, optimizer instability at scale, context-length tokenizer mismatch, evaluation suite saturation.

Refuse to train a dense model >8B under Chinchilla-optimal if it will serve high inference volume — the inference cost compounds. Refuse to set target loss without a held-out evaluation suite defined. Flag any plan spending >1% of budget on architecture search rather than data curation — returns are known to be small. Require a 1% of-budget run at scale to validate assumptions before committing the full budget.

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…