Skip to content
Back to skills

Coreweave Performance Tuning

ASecurity

'Optimize CoreWeave GPU inference latency and throughput.

  • 2,688 stars
  • 0 votes
  • 0 copies
  • 1 view
  • Added May 30, 2026
ai-agentsgobashkubernetesapiperformancedocumentation

Works with

  • claude code
  • api

Security analysis

A100/100

Scanned May 30, 2026

npx -y skills add jeremylongshore/claude-code-plugins-plus-skills --skill coreweave-performance-tuning --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Coreweave Performance Tuning?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for Coreweave Performance Tuning
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/jeremylongshore-coreweave-performance-tuning/badge)](https://www.skillsdirectory.com/skills/jeremylongshore-coreweave-performance-tuning)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
SKILL.md
---
name: coreweave-performance-tuning
description: 'Optimize CoreWeave GPU inference latency and throughput.

  Use when reducing inference latency, maximizing GPU utilization,

  or tuning batch sizes and concurrency.

  Trigger with phrases like "coreweave performance", "coreweave latency",

  "coreweave throughput", "optimize coreweave inference".

  '
allowed-tools: Read, Write, Edit, Bash(kubectl:*)
version: 1.0.0
license: MIT
author: Jeremy Longshore <jeremy@intentsolutions.io>
tags:
- saas
- gpu-cloud
- kubernetes
- inference
- coreweave
compatibility: Designed for Claude Code
---
# CoreWeave Performance Tuning

## GPU Selection by Workload

| Workload | Recommended GPU | Why |
|----------|----------------|-----|
| LLM inference (7-13B) | A100 80GB | Good balance of memory and cost |
| LLM inference (70B+) | 8xH100 | NVLink for tensor parallelism |
| Image generation | L40 | Good for diffusion models |
| Training (large models) | 8xH100 SXM5 | Fastest interconnect |
| Batch processing | A100 40GB | Cost-effective |

## Inference Optimization

```yaml
# Continuous batching with vLLM
containers:
  - name: vllm
    args:
      - "--model=meta-llama/Llama-3.1-8B-Instruct"
      - "--max-num-batched-tokens=8192"
      - "--max-num-seqs=256"
      - "--gpu-memory-utilization=0.90"
      - "--enable-prefix-caching"
      - "--dtype=float16"
```

## Autoscaling Tuning

```yaml
# HPA based on GPU utilization
apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
metadata:
  name: inference-hpa
spec:
  scaleTargetRef:
    apiVersion: apps/v1
    kind: Deployment
    name: inference-server
  minReplicas: 2
  maxReplicas: 10
  metrics:
    - type: Pods
      pods:
        metric:
          name: DCGM_FI_DEV_GPU_UTIL
        target:
          type: AverageValue
          averageValue: "70"
```

## Performance Benchmarks

| Metric | A100-80GB | H100-80GB |
|--------|-----------|-----------|
| Llama-8B tokens/sec | ~2,000 | ~4,500 |
| Llama-70B tokens/sec | ~200 (4x) | ~500 (4x) |
| Cold start (vLLM) | 30-60s | 20-40s |

## Resources

- [CoreWeave Inference](https://www.coreweave.com/solutions/ai-inference)
- [vLLM Documentation](https://docs.vllm.ai)

## Next Steps

For cost optimization, see `coreweave-cost-tuning`.

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…