Skip to content
Back to skills

Retrieval Rag And Data Pipelines

ASecurity

"Document preprocessing, LocalDB persistence, retriever indexing,

  • 247 stars
  • 0 votes
  • 0 copies
  • 0 views
  • Added September 8, 2026
ai-agentsbackend

Works with

  • cli

Security analysis

A100/100

Pro scans all 7 files and shows the line behind each finding

Scanned September 8, 2026

npx -y skills add VectorSpaceLab/AREX-Skill --skill retrieval-rag-and-data-pipelines --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Retrieval Rag And Data Pipelines?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for Retrieval Rag And Data Pipelines
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/vectorspacelab-retrieval-rag-and-data-pipelines/badge)](https://www.skillsdirectory.com/skills/vectorspacelab-retrieval-rag-and-data-pipelines)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
SKILL.md
---
name: retrieval-rag-and-data-pipelines
description: "Document preprocessing, LocalDB persistence, retriever indexing,
  and RAG context assembly."
disable-model-invocation: true
metadata:
  disco-role: operating
license: MIT
---

# Retrieval, RAG, and Data Pipelines

Use this sub-skill when the task is to prepare documents, split text, persist local data, build retriever indexes, convert retrieval output into context, or assemble a retrieval-first RAG flow.

## Route here for

- `Document` modeling for text corpora and chunked passages.
- `TextSplitter` chunking rules, separator choices, overlap settings, and invalid chunk-shape troubleshooting.
- `ToEmbeddings` and `RetrieverOutputToContextStr` as retrieval-side data transforms.
- `LocalDB` load/extend/transform/save/load flows for documents and chunked corpora.
- Retriever base contracts and concrete retrieval backends: `BM25Retriever`, `FAISSRetriever`, `LanceDBRetriever`, `QdrantRetriever`, and `PostgresRetriever`.
- Retrieval-oriented RAG assembly patterns: split, embed, persist, retrieve, deduplicate, and pass context downstream.

## Do not handle here

- Provider/client selection, embedding model setup, or generator configuration: route to the model-client-and-generator-workflows sub-skill.
- Metrics, evaluator loops, prompt optimization, or retriever scoring analysis: route to the evaluation-and-optimization sub-skill.
- Agentic RAG, tool use, runner orchestration, or streaming event handling: route to the agents-tools-and-streaming sub-skill.
- Model training or optimization of retriever/generator parameters: route to the evaluation-and-optimization sub-skill.

## Operating workflow

1. Normalize inputs into `Document` objects and keep raw metadata with the corpus.
2. Choose the lightest valid preprocessing path:
   - `TextSplitter` for chunking plain text.
   - `ToEmbeddings` when chunks need vectors.
   - `LocalDB` when you need persistence, transform reuse, or staged filtering.
3. Pick the retrieval backend:
   - `BM25Retriever` for lexical or keyword-first retrieval.
   - `FAISSRetriever` for local vector search.
   - `LanceDBRetriever`, `QdrantRetriever`, or `PostgresRetriever` when the index lives in an optional store.
4. Use `RetrieverOutputToContextStr` to assemble retrieved chunks into a downstream context string.
5. Build the RAG boundary with a retrieval step, a context builder, and a downstream answer component from another sub-skill.
6. When retrieval fails, consult [troubleshooting](references/troubleshooting.md) before changing the corpus or the retrieval contract.

## Bundled scripts

- [`scripts/text_splitter_smoke.py`](scripts/text_splitter_smoke.py): deterministic `TextSplitter` smoke check with a tiny `Document`.
- [`scripts/localdb_smoke.py`](scripts/localdb_smoke.py): `LocalDB` load/transform/save/load smoke using a safe bundled transform.

Start with [data pipelines](references/data-pipelines.md) for document and transform contracts, [retrievers](references/retrievers.md) for backend selection, and [RAG recipes](references/rag-recipes.md) for assembly patterns.

Files in this skill

  • SKILL.md3 KB
  • references/data-pipelines.md5.7 KB
  • references/rag-recipes.md6 KB
  • references/retrievers.md6.2 KB
  • references/troubleshooting.md5.2 KB
  • scripts/localdb_smoke.py1.8 KB
  • scripts/text_splitter_smoke.py1 KB

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…