Use when cost optimization patterns for LLM API usage — model routing by task complexity, budget tracking, retry logic, and prompt caching. Triggers on \"cost-aware-llm-pipeline\", \"cost aware llm pipeline\", \"pipeline\".
Installs into .claude/skills of the current project.
Are you the author of Cost Aware Llm Pipeline?
Add the live security badge to your README. It updates with every re-scan.
[](https://www.skillsdirectory.com/skills/majinmagros-cost-aware-llm-pipeline)
---
name: cost-aware-llm-pipeline
description: "Use when cost optimization patterns for LLM API usage — model routing by task complexity, budget tracking, retry logic, and prompt caching. Triggers on \"cost-aware-llm-pipeline\", \"cost aware llm pipeline\", \"pipeline\"."
metadata:
origin: ECC
---
# Cost-Aware LLM Pipeline
Patterns for controlling LLM API costs while maintaining quality. Combines model routing, budget tracking, retry logic, and prompt caching into a composable pipeline.
## When to Activate
- Building applications that call LLM APIs (Claude, GPT, etc.)
- Processing batches of items with varying complexity
- Need to stay within a budget for API spend
- Optimizing cost without sacrificing quality on complex tasks
## Core Concepts
### 1. Model Routing by Task Complexity
Automatically select cheaper models for simple tasks, reserving expensive models for complex ones.
```python
MODEL_SONNET = "claude-sonnet-4-6"
MODEL_HAIKU = "claude-haiku-4-5-20251001"
_SONNET_TEXT_THRESHOLD = 10_000 # chars
_SONNET_ITEM_THRESHOLD = 30 # items
def select_model(
text_length: int,
item_count: int,
force_model: str | None = None,
) -> str:
"""Select model based on task complexity."""
if force_model is not None:
return force_model
if text_length >= _SONNET_TEXT_THRESHOLD or item_count >= _SONNET_ITEM_THRESHOLD:
return MODEL_SONNET # Complex task
return MODEL_HAIKU # Simple task (3-4x cheaper)
```
### 2. Immutable Cost Tracking
Track cumulative spend with frozen dataclasses. Each API call returns a new tracker — never mutates state.
```python
from dataclasses import dataclass
@dataclass(frozen=True, slots=True)
class CostRecord:
model: str
input_tokens: int
output_tokens: int
cost_usd: float
@dataclass(frozen=True, slots=True)
class CostTracker:
budget_limit: float = 1.00
records: tuple[CostRecord, ...] = ()
def add(self, record: CostRecord) -> "CostTracker":
"""Return new tracker with added record (never mutates self)."""
```
## Intent-Based Routing (Batch 16, #46)
Route by capability name, not model name. The app asks for
"text-summarizer"; the gateway maps it to the contracted model with
fallback/retry/timeout policies. Developers stop tracking which model is
"best this week" - models are commodities and the contract owner swaps
them. Decide each routing change with the latency x quality x cost
tradeoff written down (e.g. +5pp accuracy for +50% cost per 1M tokens is
worth it only when errors strangle the business).
## Preco por hora de agente (Batch 17a, #52)
Preco/token engana entre tiers: Fable gastou $200 vs Opus $91 vs Sonnet
$55 no mesmo bench e "perdeu" no token — mas a metrica que importa e
preco por hora de agente inteligente. Modelos Mythos-class so se pagam
em specs grandes e complexas; em task pequena o caro e desperdicio.
Meca sempre na sua carga antes de orcar.