All authors

Claude Skills by yangwhale
github.com/yangwhale94 skills0 installs118 views
- Launch Nemo RlPlaybook for launching, monitoring, stopping, and debugging NeMo-RL recipes on a Kubernetes cluster via the nrl-k8s CLI. Covers ephemeral vs long-lived RayCluster modes, iterating on runs, and debugging hung or failed training jobs.Votes: 0GitHub stars: 4
- Mcore Create IssueInvestigate a failing GitHub Actions run or job and create a GitHub issue for the failure.Votes: 0GitHub stars: 4
- Mcore Linting And FormattingLinting and formatting for Megatron-LM. Covers running autoformat.sh, tools (ruff, black, isort, pylint, mypy), and code style rules.Votes: 0GitHub stars: 4
- Mcore Run On SlurmHow to launch distributed Megatron-LM training jobs on a SLURM cluster. Covers a minimal sbatch skeleton, environment-variable setup for torch.distributed.run, CUDA_DEVICE_MAX_CONNECTIONS rules across hardware and parallelism modes, container conventions, monitoring, and per-rank failure diagnosis.Votes: 0GitHub stars: 4
- Mcore Split PrSplit a PR into multiple PRs to reduce the number of required CODEOWNERS reviewer groups.Votes: 0GitHub stars: 4
- Mcore TestingTest system for Megatron-LM. Covers test layout, recipe YAML structure, adding and running unit and functional tests, golden values, marker filters, and CI parity.Votes: 0GitHub stars: 4
- Nemo Automodel Distributed TrainingGuide for selecting and configuring distributed training strategies in NeMo AutoModel, including FSDP2, Megatron FSDP, DDP, and parallelism settings.Votes: 0GitHub stars: 4
- Nemo Automodel Launcher ConfigConfigure NeMo AutoModel job launches for interactive runs, Slurm clusters, and SkyPilot cloud execution.Votes: 0GitHub stars: 4
- Nemo Automodel Model OnboardingGuide for onboarding new model architectures into NeMo AutoModel, including architecture discovery, implementation patterns, registration, and validation.Votes: 0GitHub stars: 4
- Nemo Automodel Recipe DevelopmentCreate and modify NeMo AutoModel training and evaluation recipes, including YAML structure, builders, and execution flow.Votes: 0GitHub stars: 4
- Nemo Mbridge Mlm Bridge TrainingRun Megatron-LM (MLM) and Megatron Bridge training with mock or real data. Covers correlation testing, available recipes, and multi-GPU examples.Votes: 0GitHub stars: 4
- Nemo Mbridge Multi Node SlurmConvert single-node scripts to multi-node Slurm sbatch jobs and debug common multi-node failures. Covers srun-native vs uv run torch.distributed approaches, container setup, NCCL timeouts, OOM sizing for MoE models, and interactive allocation.Votes: 0GitHub stars: 4
- Nemo Mbridge Perf Activation RecomputeValidate and use selective and full activation recompute in Megatron Bridge to reduce GPU memory usage at the cost of extra compute.Votes: 0GitHub stars: 4
- Nemo Mbridge Perf Cpu OffloadingValidate and use CPU offloading in Megatron Bridge, including layer-level activation offloading and fractional optimizer state offloading with HybridDeviceOptimizer.Votes: 0GitHub stars: 4
- Nemo Mbridge Perf Cuda GraphsValidate and use CUDA graph capture in Megatron Bridge, including local full-iteration graphs and Transformer Engine scoped graphs for attention, MLP, and MoE modules.Votes: 0GitHub stars: 4
- Nemo Mbridge Perf Expert Parallel OverlapValidate and use MoE expert-parallel communication overlap in Megatron-Bridge, including overlap_moe_expert_parallel_comm, delay_wgrad_compute, and flex dispatcher backends such as DeepEP and HybridEP.Votes: 0GitHub stars: 4
- Nemo Mbridge Perf Hierarchical Context ParallelOperational guide for enabling hierarchical context parallelism in Megatron-Bridge, including config knobs, code anchors, pitfalls, and verification.Votes: 0GitHub stars: 4
- Nemo Mbridge Perf Megatron FsdpOperational guide for enabling Megatron FSDP in Megatron-Bridge, including config knobs, code anchors, pitfalls, and verification.Votes: 0GitHub stars: 4
- Nemo Mbridge Perf Memory TuningTechniques for reducing peak GPU memory in Megatron Bridge — expandable segments, parallelism resizing, activation recompute, CPU offloading constraints, and common OOM fixes.Votes: 0GitHub stars: 4
- Nemo Mbridge Perf Moe Comm OverlapMoE expert-parallel communication overlap in Megatron Bridge. Covers dispatch/combine overlap, flex dispatcher backends, and expert wgrad scheduling.Votes: 0GitHub stars: 4
- Nemo Mbridge Perf Moe Dispatcher SelectionChoose the right MoE token dispatcher (`alltoall`, DeepEP, or HybridEP) for the hardware, EP degree, and optimization stage. Summarizes patterns from DSV3, Qwen3, Qwen3-Next, and VLM bring-up work.Votes: 0GitHub stars: 4
- Nemo Mbridge Perf Moe Hardware ConfigsRepresentative MoE training playbooks by hardware platform and model family. Summarizes rounded throughput bands, parallelism patterns, and common tuning stacks.Votes: 0GitHub stars: 4
- Nemo Mbridge Perf Moe Long ContextLong-context MoE training guidance for Megatron Bridge. Covers CP sizing, selective recompute, dispatcher choices, and practical patterns from DSV3, Qwen3, and Qwen3-Next long-context experiments.Votes: 0GitHub stars: 4
- Nemo Mbridge Perf Moe Optimization WorkflowSystematic workflow for MoE training optimization in Megatron Bridge, based on the Megatron-Core MoE paper. Covers the Three Walls framework, parallel folding, recompute strategy, dispatcher choice, and CUDA-graph bring-up.Votes: 0GitHub stars: 4
- Nemo Mbridge Perf Moe Vlm TrainingPractical guidance for training MoE VLMs in Megatron Bridge. Compares FSDP and 3D-parallel approaches, using rounded lessons from Qwen3-VL, Qwen3-Next, and other multimodal experiments.Votes: 0GitHub stars: 4
- Nemo Mbridge Perf Parallelism StrategiesOperational guide for choosing and combining parallelism strategies in Megatron Bridge, including sizing rules, hardware topology mapping, and combined parallelism configuration.Votes: 0GitHub stars: 4
- Nemo Mbridge Perf Sequence PackingValidate and use packed sequences and long-context training in Megatron-Bridge, distinguishing offline packed SFT for LLMs from in-batch packing for VLMs, and applying the right CP constraints.Votes: 0GitHub stars: 4
- Nemo Mbridge Perf Tp Dp Comm OverlapOperational guide for enabling TP, DP, and PP communication overlap in Megatron-Bridge, including config knobs, code anchors, pitfalls, and verification.Votes: 0GitHub stars: 4
- Nemo Mbridge Recipe RecommenderRecommend and customize Megatron Bridge recipes for a user's model, GPU count, and training goal. Indexes library recipes (pretrain/SFT/PEFT) and performance recipes.Votes: 0GitHub stars: 4
- Nemo Mbridge ResiliencyResiliency features in Megatron Bridge including fault tolerance, straggler detection, in-process restart, preemption, and re-run state machine.Votes: 0GitHub stars: 4
- Nemo Rl Auto ResearchAutonomous NeMo-RL research agent workflow for directed hypothesis testing and open-ended discovery. Guides agents through the full experiment lifecycle: understanding recipes and environments, wiring RL or NeMo-gym runs, launching reproducible baselines and iterations, analyzing results, preserving human oversight, and using git plus TSV logs as the research ledger. Do NOT use for: bug fixes, code review, documentation, refactoring, dependency updates, or single-file changes.Votes: 0GitHub stars: 4
- Nemo Rl Brev EtiquetteBrev instance operating guidance for NeMo-RL agents working in /home/ubuntu/RL with limited workspace disk, a larger /ephemeral volume, and optional /home/ubuntu/RL/.env secrets. Use when running nemo-rl-auto-research campaigns, experiments, training jobs, model or dataset downloads, shared cache-heavy commands, log-producing runs, checkpoint generation, W&B or Hugging Face authenticated workflows, or any workflow that may create large files on Brev.Votes: 0GitHub stars: 4
- Nemo Rl DocsDocumentation conventions for NeMo-RL. Covers docs/index.md updates and docstring format. Do NOT use for: bug fixes, test fixes, dependency bumps, refactoring, CI/CD changes, performance tuning, or any task that does not involve writing or updating documentation.Votes: 0GitHub stars: 4
- Nemo Rl Session MemoryManage durable working-session memory for coding agents. Use when a user asks to preserve or recover agent context across disconnects, VS Code restarts, long-running work, handoffs, or any session where important state should be written periodically under the repo's session directory. Do NOT use for: simple questions, short tasks, one-off commands, linting, or code review.Votes: 0GitHub stars: 4
- Tilegym Adding Cutile KernelAdd a new cuTile GPU kernel operator to TileGym. Covers dispatch registration in ops.py, cuTile backend implementation, __init__.py exports, test creation, and benchmark in tests/benchmark. Use when adding, creating, or implementing a new cuTile operator/kernel in TileGym, or when asking how to register a new cuTile op.Votes: 0GitHub stars: 4
- Tilegym Converting Cutile To JuliaConverts cuTile Python GPU kernels (@ct.kernel) to cuTile.jl Julia equivalents. Handles kernel syntax translation, 0-indexed to 1-indexed conversion, broadcasting differences, memory layout (row-major to column-major), type system mapping, and launch API differences. Use when converting, porting, or translating cuTile Python kernels to Julia cuTile.jl, or debugging/optimizing existing Julia cuTile translations.Votes: 0GitHub stars: 4
- Tilegym Converting Cutile To TritonConverts cuTile GPU kernels (@ct.kernel) to Triton (@triton.jit). Handles standard in-repo conversion, debugging (cudaErrorIllegalAddress, shape mismatch, numerical mismatch), and mapping cuTile idioms (ct.load/ct.store, ct.Constant, ct.launch) to Triton equivalents. Covers dual-kernel layout flags (e.g. transpose=True/False + autotune grid via META) per translations/advanced-patterns.md. Use when converting, porting, or translating cuTile kernels to Triton, or debugging existing Triton trans...Votes: 0GitHub stars: 4
- Tilegym Cutile AutotuningUse when adding, modifying, optimizing, or debugging CuTile autotuning code. Trigger signals: `exhaustive_search` / `replace_hints` / `hints_fn` / `cuda.tile.tune` in code, `autotune` in filenames, or correctness/performance issues in autotuned CuTile kernels. Covers: tune-once/cache/launch pattern, per-architecture configs (sm80–sm120), parameter space design (tile sizes, occupancy, num_ctas), and 7 common pitfalls with solutions.Votes: 0GitHub stars: 4
- Tilegym Cutile PythonExpert cuTile programming assistant. Write high-performance GPU kernels using cuTile's tile-based programming model with proper validation and optimization. Supports deep agent orchestration for complex multi-kernel tasks.Votes: 0GitHub stars: 4
- Tilegym Improve Cutile Kernel PerfIteratively optimize cuTile kernel performance through systematic profiling, bottleneck analysis, IR comparison, and targeted tuning. Covers tile sizes, occupancy, autotune configs, TMA, latency hints, persistent scheduling, num_ctas, flush_to_zero, and IR-level debugging. Use when asked to "optimize cutile kernel", "improve kernel perf", "tune cutile performance", "make kernel faster", or iteratively benchmark and refine a cuTile GPU kernel in the TileGym project.Votes: 0GitHub stars: 4
- Tilegym Monkey Patch Kernels To TransformersIntegrate TileGym kernels into Hugging Face `transformers` models by replacing the library's submodule(s) and certain class(es)' implementations, and patching certain class(es)' init/forward/load weight methods prior to instantiating models. Used when the user requires integrating TileGym kernels into `transformers` models.Votes: 0GitHub stars: 4
- Agent Teams创建和管理 Agent Teams 并行任务团队。当用户说"建个团队"、"拉个队伍"、"并行跑"、"create team"、"spawn agents"、"组队"时触发。Votes: 0GitHub stars: 4
- Bot Config管理 CloseCrab bot 配置(Firestore)。添加新 bot、给已有 bot 添加 channel、切换 channel、查看配置。当用户说"添加一个 bot"、"新建 bot"、"加个飞书 channel"、"加个 discord channel"、"切换到飞书"、"bot 配置"、"add bot"、"add channel"等关键词时触发。Votes: 0GitHub stars: 4
- Browser CliDrive the already-logged-in Chrome on Chris's cloudtop from the command line (agent-browser over CDP), instead of chrome-devtools-mcp. Use for ANY interactive browser work on internal/SSO-gated sites — gHire, Buganizer UI, Google Chat, internal dashboards, form filling, clicking through a flow. ~40-50x cheaper than the MCP (350 tokens vs 15-20K per snapshot). Also provides multi-round Google Chat conversation — send a message and block waiting for the peer's reply. Trigger on "操作浏览器", "点开网页",...Votes: 0GitHub stars: 4
- Bwrap Bypass绕过 Claude Code bwrap (bubblewrap) sandbox 限制。bwrap sandbox 会重新挂载文件系统,导致 Bash 命令无法访问特定路径、凭证文件、Unix socket 等资源。当遇到以下情况时使用:(1) Bash 命令报 Permission denied 但 sandbox 外正常 (2) LOAS2/Kerberos 凭证不可用 (3) 新机器/Bot 部署后需要配置 sandbox bypass (4) 命令因 sandbox 文件系统隔离而失败 (5) 用户说"装 bwrap bypass"、"配 sandbox"。关键词:"bwrap"、"sandbox"、"Operation not permitted"、"Permission denied"、"credentials-cache"。Votes: 0GitHub stars: 4
- Cc Pages BackupCC Pages 增量备份到 GitHub private repoVotes: 0GitHub stars: 4
- Chat Style聊天平台消息格式化规则。当 Claude Code 通过聊天平台(Discord、飞书等)运行时自动适配消息风格,避免渲染问题。当用户说"聊天风格"、"chat风格"、"消息格式"时触发。Votes: 0GitHub stars: 4
- Chinese DictationGenerate Chinese (traditional/simplified) dictation practice — extract words+pinyin from a screenshot, produce an HTML page with audio that reads each word aloud with configurable pause between words. Use when user says "听写"、"聽寫"、"语文听写"、"听写练习"、"dictation"、"做个听写"、"听写卡"、"听写音频".Votes: 0GitHub stars: 4
- Code Wiki ReconFirst-pass codebase reconnaissance using Google Code Wiki (codewiki.google). Use when the user is opening a new repository for the first time, asks for a quick architectural overview of an unfamiliar codebase, or needs to build a mental model before diving into code. Pulls the auto-generated TOC, architecture diagram, and section summaries from Code Wiki to bootstrap understanding. Strictly a recon tool, not a Q&A oracle — Code Wiki cannot answer deep runtime/optimization questions, only desc...Votes: 0GitHub stars: 4
- Deck BuilderBuild polished, on-brand slide decks (python-pptx → Google Slides) and richly formatted Google Docs, with a reusable design system, matplotlib charts, AI-generated imagery (Nano Banana), a render-verify loop, and link-stable Drive uploads (PATCH same fileId so the share link never changes). Use when the user wants to create or iterate a presentation / briefing deck / exec one-pager / formatted Google Doc — especially customer or executive briefings that go into Google Drive. Also covers extra...Votes: 0GitHub stars: 4