Skip to content
Back to skills

Embeddings

ASecurity

Set up embeddings and a vector store for semantic search or retrieval

  • 3 stars
  • 0 votes
  • 0 copies
  • 1 view
  • Added September 3, 2026
ai-agentsgo

Security analysis

A100/100

Scanned September 3, 2026

npx -y skills add black141312/ada --skill embeddings --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Embeddings?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for Embeddings
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/black141312-embeddings/badge)](https://www.skillsdirectory.com/skills/black141312-embeddings)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
SKILL.md
---
name: embeddings
description: Set up embeddings and a vector store for semantic search or retrieval
category: data-ml
---

# Embeddings

Use when you need semantic search, clustering, dedup, or retrieval and must turn text (or other data) into vectors stored in an index.

1. Pick an embedding model matched to the domain and language; note its dimensionality, max input length, and cost.
2. Preprocess and chunk inputs to fit the model's token limit, keeping a stable id and metadata for each item.
3. Batch the embedding calls, handle rate limits/retries, and normalize vectors if your similarity metric expects it (cosine).
4. Store vectors in a vector store (FAISS, pgvector, or a managed index) with the chosen distance metric and an ANN index sized to your recall/latency target.
5. Query by embedding the query the same way, retrieving top-k, and verifying results make semantic sense on real queries.
6. Persist the model name and version with the index so you can detect and handle re-embedding needs.

## Rules
- Embed queries and documents with the exact same model and preprocessing — mismatches silently wreck recall.
- Match the index's distance metric to the model (most use cosine/dot on normalized vectors); don't mix metrics.
- Re-embedding is required when you change the model, dimension, or chunking — version the index accordingly.
- Store source ids and metadata alongside vectors so retrieved hits can be traced back and filtered.
- Benchmark recall and latency at your real data size; small-sample behavior doesn't predict ANN tradeoffs.

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…