Skip to content
Back to skills

Embeddings

ASecurity

Work with text embeddings: model selection, vectorization, similarity, clustering, and caching for search and retrieval. Use for semantic representation of text.

  • 2 stars
  • 0 votes
  • 0 copies
  • 0 views
  • Added October 1, 2026
ai-agentspythongobashapi

Works with

  • api

Security analysis

A96/100
  • mediumInstalls packages at runtime which could introduce malicious dependencies

Pro scans all 2 files and shows the line behind each finding

Scanned October 1, 2026

npx -y skills add ssrjkk/agent-skills --skill embeddings --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Embeddings?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for Embeddings
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/ssrjkk-embeddings/badge)](https://www.skillsdirectory.com/skills/ssrjkk-embeddings)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
SKILL.md
---
name: embeddings
description: "Work with text embeddings: model selection, vectorization, similarity, clustering, and caching for search and retrieval. Use for semantic representation of text."
category: ai
tags: [embeddings, vectors, similarity, semantic-search, nlp, representation]
models: [sonnet, opus, gpt-6, gemini-3, glm-5]
version: 1.0.0
created: 2026-09-29
updated: 2026-09-29
author: ssrjkk
---
# Embeddings

> Representing text as vectors for semantic similarity and retrieval.

## Quick Start
```bash
pip install sentence-transformers numpy scikit-learn
python -c "from sentence_transformers import SentenceTransformer; m=SentenceTransformer('all-MiniLM-L6-v2'); print(m.encode('hello').shape)"
```

## When to Use
- Semantic search and RAG retrieval
- Clustering and deduplication of texts
- Recommender systems by content similarity
- Anomaly detection over text corpora

## Best Practices

### Model Selection
- Match model size to data volume and latency budget
- Use retrieval-tuned models (e5, bge, embedding-3) for search
- Normalize vectors when using cosine similarity
- Freeze and version the embedding model

### Vectorization
- Encode in batches for throughput
- Cache embeddings to avoid recomputation
- Chunk long documents before encoding
- Keep a metadata map: id -> source text

### Similarity & Retrieval
- Use cosine similarity for normalized vectors
- Index with ANN (HNSW) for scale
- Combine with keyword search for hybrid retrieval
- Threshold scores for confidence

### Evaluation
- Test retrieval recall on a labeled set
- Compare candidate models on your domain
- Watch for domain drift in vocabulary
- Measure latency vs quality trade-offs

## Dependencies
```bash
pip install sentence-transformers numpy scikit-learn
# API-based: pip install openai
```

## Examples
```python
from sentence_transformers import SentenceTransformer

model = SentenceTransformer("all-MiniLM-L6-v2")
docs = ["Alpha release notes", "Beta API reference", "Gamma user guide"]
emb = model.encode(docs, normalize_embeddings=True)
print(emb.shape)  # (3, 384)
```
```python
# Cosine similarity matrix
import numpy as np

def cosine(a, b):
    return float(np.dot(a, b))

query = model.encode(["how to deploy?"], normalize_embeddings=True)[0]
scores = [cosine(query, e) for e in emb]
best = int(np.argmax(scores))
print(f"best: {docs[best]} score={scores[best]:.3f}")
```
```python
# Batch encoding with caching
import json, pathlib

cache_path = pathlib.Path("embeddings.json")
cache = json.loads(cache_path.read_text()) if cache_path.exists() else {}

def get_emb(texts: list[str]) -> list[list[float]]:
    missing = [t for t in texts if t not in cache]
    if missing:
        for t, v in zip(missing, model.encode(missing, normalize_embeddings=True)):
            cache[t] = v.tolist()
        cache_path.write_text(json.dumps(cache))
    return [cache[t] for t in texts]
```
```python
# Clustering with k-means
from sklearn.cluster import KMeans

kmeans = KMeans(n_clusters=3, n_init=10, random_state=0)
labels = kmeans.fit_predict(emb)
for i, label in enumerate(labels):
    print(docs[i], "-> cluster", label)
```

## Step-by-Step
1. Choose an embedding model sized to your data and latency.
2. Encode your corpus in batches; cache results.
3. Normalize vectors for cosine similarity.
4. Index with an ANN store (or numpy for small sets).
5. Build similarity queries with thresholds.
6. Add hybrid keyword retrieval where recall matters.
7. Evaluate recall@k on a labeled set.
8. Version the model and re-encode on upgrades.

## Validation
1. Similar documents score higher than dissimilar ones
2. recall@k meets the target on the eval set
3. Encoding latency is within budget
4. Caching avoids duplicate work
5. Results are reproducible with a fixed seed

## Troubleshooting
- Poor similarity: switch model or normalize vectors properly.
- Slow encoding: batch larger or downsize the model.
- Domain drift: fine-tune embeddings or add keyword boost.

Files in this skill

  • SKILL.md3.9 KB
  • SKILL.ru.md5.3 KB

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…