Skip to content
Back to skills

Embeddings

ASecurity

Generate, store, and search vector embeddings with provider selection, chunking strategies, and similarity search optimization.

  • 17 stars
  • 0 votes
  • 0 copies
  • 1 view
  • Added September 6, 2026
ai-agentspythongoapidatabase

Works with

  • cli
  • api

Security analysis

A100/100

Pro scans all 6 files and shows the line behind each finding

Scanned September 6, 2026

npx -y skills add clawic/skills --skill embeddings --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Embeddings?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for Embeddings
[![Security: A โ€” Skills Directory](https://www.skillsdirectory.com/api/skills/clawic-embeddings/badge)](https://www.skillsdirectory.com/skills/clawic-embeddings)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
SKILL.md
---
name: Embeddings
slug: embeddings
version: 1.0.0
description: Generate, store, and search vector embeddings with provider selection, chunking strategies, and similarity search optimization.
homepage: https://clawic.com/skills/embeddings
metadata:
  clawdbot:
    emoji: ๐Ÿงฎ
    displayName: Embeddings
---

## When to Use

User wants to convert text/images to vectors, build semantic search, or integrate embeddings into applications.

## Quick Reference

| Topic | File |
|-------|------|
| Provider comparison & selection | `providers.md` |
| Chunking strategies & code | `chunking.md` |
| Vector database patterns | `storage.md` |
| Search & retrieval tuning | `search.md` |

## Core Capabilities

1. **Generate embeddings** โ€” Call provider APIs (OpenAI, Cohere, Voyage, local models)
2. **Chunk content** โ€” Split documents with overlap, semantic boundaries, token limits
3. **Store vectors** โ€” Insert into Pinecone, Weaviate, Qdrant, pgvector, Chroma
4. **Similarity search** โ€” Query with top-k, filters, hybrid search
5. **Batch processing** โ€” Handle large datasets with rate limiting and retries
6. **Model comparison** โ€” Evaluate embedding quality for specific use cases

## Decision Checklist

Before recommending approach, ask:
- [ ] What content type? (text, code, images, multimodal)
- [ ] Volume and update frequency?
- [ ] Latency requirements? (real-time vs batch)
- [ ] Budget constraints? (API costs vs self-hosted)
- [ ] Existing infrastructure? (cloud provider, database)

## Critical Rules

- **Same model everywhere** โ€” Query embeddings MUST use identical model as document embeddings
- **Normalize before storage** โ€” Most similarity metrics assume unit vectors
- **Chunk with overlap** โ€” 10-20% overlap prevents context loss at boundaries
- **Batch API calls** โ€” Never embed one item at a time in production
- **Cache embeddings** โ€” Regenerating is expensive; store with source hash
- **Monitor dimensions** โ€” Higher isn't always better; 768-1536 is usually optimal

## Provider Quick Selection

| Need | Provider | Why |
|------|----------|-----|
| Best quality, any cost | OpenAI `text-embedding-3-large` | Top benchmarks |
| Cost-sensitive | OpenAI `text-embedding-3-small` | 5x cheaper, 80% quality |
| Multilingual | Cohere `embed-multilingual-v3` | 100+ languages |
| Code/technical | Voyage `voyage-code-2` | Optimized for code |
| Privacy/offline | Local (e5, bge, nomic) | No data leaves machine |
| Images | OpenAI CLIP, Cohere multimodal | Cross-modal search |

## Common Patterns

```python
# Batch embedding with retry
def embed_batch(texts, model="text-embedding-3-small"):
    results = []
    for chunk in batched(texts, 100):  # API limit
        response = client.embeddings.create(input=chunk, model=model)
        results.extend([e.embedding for e in response.data])
    return results

# Similarity search with filter
results = index.query(
    vector=query_embedding,
    top_k=10,
    filter={"category": "technical"},
    include_metadata=True
)
```

Files in this skill

  • SKILL.md3 KB
  • _meta.json172 B
  • chunking.md3.7 KB
  • providers.md3.2 KB
  • search.md4.4 KB
  • storage.md3.8 KB

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading commentsโ€ฆ