Skip to content
Back to skills

Ai Rag Patterns

ASecurity

Use when building features that answer questions from private data, documents, policies, or time-sensitive information — RAG architecture, chunking, hybrid search, re-ranking, vector databases, evaluation, agentic and multimodal RAG, and bounded iterative retrieval for context-starved subagents.

  • 28 stars
  • 0 votes
  • 0 copies
  • 2 views
  • Added May 28, 2026
developmentpythongoawsdatabasebackendsecurity

Security analysis

A100/100

Pro scans all 7 files and shows the line behind each finding

Scanned October 1, 2026

npx -y skills add peterbamuhigire/skills-web-dev --skill ai-rag-patterns --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Ai Rag Patterns?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for Ai Rag Patterns
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/peterbamuhigire-ai-rag-patterns/badge)](https://www.skillsdirectory.com/skills/peterbamuhigire-ai-rag-patterns)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
SKILL.md
---
name: ai-rag-patterns
description: Use when building features that answer questions from private data, documents, policies, or time-sensitive information — RAG architecture, chunking, hybrid search, re-ranking, vector databases, evaluation, agentic and multimodal RAG, and bounded iterative retrieval for context-starved subagents.
metadata:
  portable: true
  compatible_with:
  - claude-code
  - codex
---

# RAG Patterns — Retrieval-Augmented Generation

## Operating contract

## Inputs

| Input | Required | Purpose |
|---|---|---|
| Domain evidence | yes | question set, authoritative corpus, tenancy and access rules, freshness target, citation needs, and evaluation set |

## Outputs

- Produce: retrieval architecture, ingestion/chunking policy, index and filter contract, grounded-answer schema, and evaluation report.

## Capability and permission boundaries

Default to read-only analysis. Read only scoped records; redact secrets and regulated data. Writes, execution, network calls, production configuration, customer communication, billing changes, and delegation require explicit authority and an identified owner. Never widen tenant, time-window, or system scope implicitly.

## Degraded mode

When required telemetry, evidence, execution, network access, or write authority is unavailable, return a partial result with each unassessed item labelled, preserve the safest existing state, and state the evidence or approval needed to continue. Never convert missing evidence into a pass.

## Decision rules

| Condition | Action |
|---|---|
| Scope, owner, or threshold is missing | Stop the affected decision and request it |
| Evidence is incomplete but read-only analysis is safe | Produce a qualified partial result and gap list |
| A mutation exceeds authority or tenant boundary | Block it and route for approval |
| Evidence meets the stated threshold | Issue the output with provenance and owner |

## Anti-Patterns

- Treating absent evidence as success. Fix: mark the check unassessed and name the missing source.
- Expanding one tenant or workflow to all tenants. Fix: enforce supplied scope at every query and action.
- Performing a production write during analysis. Fix: emit a reviewed change plan until authority is explicit.
- Reporting a metric without population, window, or source. Fix: attach all three.
- Hiding a failed threshold inside an average. Fix: report failure slices and the remediation owner.

Acknowledgement: Shared by Peter Bamuhigire, techguypeter.com, +256 784 464178.

<!-- dual-compat-start -->
## Use When

- Use when building features that answer questions from private data, documents, policies, or time-sensitive information — RAG architecture, chunking strategies, hybrid search, re-ranking, vector databases, evaluation, agentic RAG, multimodal RAG...

## Evidence Produced

| Category | Artifact | Format | Example |
|----------|----------|--------|---------|
| Correctness | RAG retrieval evaluation report | Markdown doc covering recall / precision / answer-quality on a fixed eval set | `docs/ai/rag-eval-2026-04-16.md` |
| Data safety | Index ingestion + tenancy isolation note | Markdown doc covering chunking, source filtering, and per-tenant index segregation | `docs/ai/rag-tenancy-note.md` |

## References

- Use the `references/` directory for deep detail after reading the core workflow below.
- [references/curated-corpus-worked-example.md](references/curated-corpus-worked-example.md): small curated corpus with BM25 ranking, per-domain score floors, calibrated abstention and a graded, fingerprinted golden set.
<!-- dual-compat-end -->
## Overview

RAG solves the core LLM limitation: they only know what they were trained on. Use RAG to inject private data (invoices, menus, policies, reports) into every AI response.

**Core principle:** RAG = look up a database + LLM synthesises the results. The LLM never needs to "know" your data.

---

## When to Use RAG

| Condition | Action |
|---|---|
| Knowledge base < 200K tokens (~500 pages) | Include everything in context — no RAG needed |
| Knowledge base > 200K tokens | Use RAG |
| Data changes frequently (menus, prices, stock) | RAG (update documents, not model) |
| Data is private/confidential | RAG (keeps data out of training pipelines) |
| Need source citations | RAG (chunks are traceable to source) |
| Model needs brand voice / domain jargon | Fine-tune instead |

---

## RAG vs Fine-Tuning

| Factor | RAG | Fine-Tuning |
|---|---|---|
| Up-to-date content | ✅ Yes (add docs anytime) | ❌ Stale until retrained |
| Hallucinations | ✅ Lower (document-grounded) | ❌ Higher |
| Source citations | ✅ Yes | ❌ No |
| Brand voice control | ❌ Weak | ✅ Strong |
| Domain jargon | ❌ Weak | ✅ Strong |
| Up-front cost | ✅ Lower | ❌ High |

**Default: start with RAG.** Fine-tune only when RAG + prompt engineering cannot deliver the required tone or vocabulary.

---

## Additional Guidance

Guidance is split across two reference files so this entrypoint stays compact.

**[references/skill-deep-dive.md](references/skill-deep-dive.md)** — architecture, chunking, retrieval, schema:

- `Pipeline Architecture`
- `Chunking Strategies`
- `Embedding Model Selection`
- `Vector Database Selection`
- `Retrieval Algorithms`
- `Re-Ranking`
- `Full RAG Query Algorithm`
- `Query Rewriting (Multi-Turn)`
- `RAG Schema (Multi-Tenant)`
- `Evaluation Framework`
- `Production Patterns`
- `Agentic RAG`
- `Multimodal RAG`, `Edge Cases`, `Cost Optimisation`, `Sources`

**[references/production-rag.md](references/production-rag.md)** — the progression from draft to production and the gates before shipping:

- `RAG Maturity Model` — Naive → Advanced → Modular
- `Query Transformation` — HyDE, Multi-Query, Step-Back
- `Contextual Compression`
- `Self-RAG`
- `RAGAS Evaluation` — 4 metrics with production thresholds
- `Embedding Pipeline` — batching, upserts, re-embed triggers, $/1M-token table
- `Cost Management Decision Tree` — concrete dollar figures per branch
- `Failure Mode Playbook` — empty, irrelevant, hallucinated, stale
- `Gates Before Shipping`

**[references/iterative-retrieval.md](references/iterative-retrieval.md)** — load when a context-starved subagent must explore a codebase or workspace before working: bounded DISPATCH-EVALUATE-REFINE-LOOP (max 3 cycles), relevance banding, explicit gap lists, and terminology learning.

Load the production file when building a RAG system that has to pass evaluation gates, survive multi-tenant review, or hit a cost budget under load.

When answers depend on structured records about a specific customer, order or account (not documents), or when embedding, prompt and model versions must be managed like an ML system, load `python-ml-predictive` reference `ml-system-architecture-fti.md` (feature/training/inference pipelines, entity-keyed retrieval, version coupling).
## Multi-Tenant Addendum

This skill describes RAG patterns in general. When the RAG feature ships inside a multi-tenant SaaS, the production answer is `ai-rag-multi-tenant` — per-tenant ingestion pipelines, vector store partitioning, tier-specific chunking and embedding models, defence-in-depth retrieval security, and citation grounding tied to live sources.

Cross-references:
- `ai-rag-multi-tenant` — multi-tenant RAG end-to-end.
- `ai-tenant-isolation-patterns` — vector-store partitioning tradeoffs and data-bleed tests.
- `ai-on-saas-architecture` — KB service as a control-plane service.
- `ai-hallucination-slo-and-grounding` — citation grounding + faithfulness SLO.
- `ai-model-gateway` — gateway-mediated retrieval calls.
- `saas-tenant-data-portability-and-erasure` — KB erasure cascade for embeddings.
- [Catalog metadata for AI](../../backend-databases/database-design-engineering/references/metadata-catalog-and-lineage-design.md) — load when retrieval or an agent draws on data-catalog or data-contract metadata: classification-first filtering, ODCS `context` block, provenance.
## Consolidated Child References

- Load [references/routing.md](references/routing.md) to map retired AI child skill slugs to their reference modules.

Files in this skill

  • SKILL.md6.7 KB
  • references/ai-rag-multi-tenant/entrypoint.md11 KB
  • references/ai-rag-multi-tenant/references/per-tenant-ingestion-pipeline.md4.8 KB
  • references/ai-rag-multi-tenant/references/retrieval-security-patterns.md4.2 KB
  • references/production-rag.md14.7 KB
  • references/routing.md390 B
  • references/skill-deep-dive.md15 KB

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…