All authors

Claude Skills by claude-dev-suite
github.com/claude-dev-suite740 skills14 installs1,174 views
- JacocoJaCoCo Java code coverage tool USE WHEN: user mentions "JaCoCo", "Java coverage", "code coverage", asks about "coverage threshold", "jacoco-maven-plugin", "coverage report", "LINE coverage", "BRANCH coverage" DO NOT USE FOR: JavaScript/TypeScript coverage - use Vitest skill, SonarQube analysis - use `sonarqube` skill, test execution - use testing skillsVotes: 0GitHub stars: 31
- Java QualityJava code quality with Checkstyle, SpotBugs, PMD, and SonarJava. Covers static analysis, code style, and best practices. USE WHEN: user works with "Java", "Spring Boot", "Maven", "Gradle", asks about "Checkstyle", "SpotBugs", "PMD", "Java code smells", "Java best practices" DO NOT USE FOR: SonarQube generic - use `sonarqube` skill, testing - use Spring Boot test skills, security - use `java-security` skillVotes: 0GitHub stars: 31
- Kotlin QualityKotlin code quality toolchain — detekt (static analysis with rule sets, baseline, custom rules), ktlint (formatter + linter for the Kotlin style guide), and Compose-specific lint rules. Covers Gradle integration (KMP-aware), pre-commit hooks, CI integration, baseline files for legacy code, and config patterns for production apps including Jetpack Compose specifics (detekt-compose-rules). USE WHEN: user mentions "detekt", "ktlint", "ktlint-gradle", "detekt-gradle", "kotlin static analysis", "...Votes: 0GitHub stars: 31
- Osv ScannerOSV-Scanner — language-agnostic vulnerability scanner by Google, queries the OSV.dev database (aggregates RustSec, GHSA, PyPA, npm advisories, Go vulndb, etc.). Scans Cargo.lock, package-lock.json, pnpm-lock.yaml, Pipfile.lock, go.sum, Gradle dependency tree, container images, and Git submodules. Replaces multiple language-specific scanners with one tool. CI integration with SARIF output. USE WHEN: user mentions "osv-scanner", "osv.dev", "OSV scanner", "OSV database", "GHSA", "language-agnos...Votes: 0GitHub stars: 31
- Php QualityPHP code quality: static analysis, coding standards and best-practice enforcement for PHP codebases. USE WHEN: working with PHP codebases requiring code quality analysis, static analysis, or best practices enforcement DO NOT USE FOR: security-specific issues (use `php-security`), runtime debugging, or deploymentVotes: 0GitHub stars: 31
- Python QualityPython code quality with Ruff, Black, mypy, and Pylint. Covers linting, formatting, type checking, and best practices. USE WHEN: user works with "Python", "Django", "FastAPI", "Flask", asks about "Ruff", "Black", "mypy", "Pylint", "Python linting", "Python type hints" DO NOT USE FOR: SonarQube - use `sonarqube` skill, testing - use pytest skills, security - use `python-security` skillVotes: 0GitHub stars: 31
- Rust QualityRust code quality with Clippy, rustfmt, and cargo tools. Covers linting, formatting, and idiomatic Rust patterns. USE WHEN: user works with "Rust", "Cargo", "Actix", "Axum", asks about "Clippy", "rustfmt", "Rust linting", "Rust best practices", "idiomatic Rust" DO NOT USE FOR: SonarQube - use `sonarqube` skill, security - use `rust-security` skillVotes: 0GitHub stars: 31
- Rust Supply ChainRust supply-chain & quality toolchain — cargo-deny (license + advisory + bans), cargo-audit (RustSec advisory DB), cargo-nextest (faster test runner with retry/parallel), cargo-tarpaulin / llvm-cov (code coverage), cargo-machete (unused deps), cargo-outdated (version gap detection), cargo-vet (third-party audit attestations). CI integration patterns and policies for production Rust apps. USE WHEN: user mentions "cargo-deny", "cargo-audit", "RustSec", "cargo-nextest", "cargo-tarpaulin", "llvm...Votes: 0GitHub stars: 31
- SonarqubeSonarQube/SonarCloud integration for continuous code quality. Setup, configuration, quality gates, and CI/CD integration. USE WHEN: user mentions "SonarQube", "SonarCloud", "quality gates", asks about "code coverage", "technical debt", "code smells", "sonar-project.properties", "SonarScanner" DO NOT USE FOR: ESLint/Biome - use linting skills, OWASP security - use security skills, testing tools - use Vitest/Playwright skillsVotes: 0GitHub stars: 31
- Typescript Eslinttypescript-eslint - ESLint plugin for TypeScript USE WHEN: user mentions "typescript-eslint", "TypeScript linting", "type-aware rules", asks about "no-floating-promises", "TypeScript ESLint config", "@typescript-eslint rules" DO NOT USE FOR: ESLint 9 full setup - use `eslint-biome` skill, Biome - use `eslint-biome`, general quality - use `quality-common`Votes: 0GitHub stars: 31
- Typescript QualityTypeScript/JavaScript code quality with ESLint, Biome, and strict TypeScript. Covers linting, formatting, type safety, and best practices. USE WHEN: user works with "TypeScript", "JavaScript", "ESLint", "Biome", asks about "TS strict mode", "type safety", "linting rules", "code formatting" DO NOT USE FOR: SonarQube - use `sonarqube` skill, testing - use Vitest/Jest skills, security - use security skillsVotes: 0GitHub stars: 31
- CanopyPinecone's Canopy framework — a pre-packaged RAG stack on top of Pinecone. Covers Canopy CLI, REST server (OpenAI-compatible), chunking, embedding, retrieval, chat-history management, and customization hooks (custom chunkers/records). USE WHEN: user mentions "Canopy", "Pinecone Canopy", "canopy start", "canopy new", "Pinecone out-of-the-box RAG" DO NOT USE FOR: custom pipelines on non-Pinecone stores - use `rag-architecture`; Postgres-based stacks - use `r2r`; LangChain/LlamaIndex abstracti...Votes: 0GitHub stars: 31
- DspyDSPy 2.5+ for programming (not prompting) LMs. Covers Signatures, Modules (Predict, ChainOfThought, ReAct, Retrieve), compilers/optimizers (BootstrapFewShot, MIPROv2, BootstrapFinetune), retrievers, evaluation, and production deployment of compiled programs. USE WHEN: user mentions "DSPy", "dspy.Module", "dspy.Signature", "MIPROv2", "BootstrapFewShot", "ChainOfThought optimization", "program > prompt" DO NOT USE FOR: manual prompt templates - use `langchain`; LlamaIndex query engines - use ...Votes: 0GitHub stars: 31
- HaystackHaystack 2.x pipeline architecture for RAG and LLM apps. Covers Components, Pipeline as DAG, document stores (InMemory, Elasticsearch, Weaviate, Pinecone, Qdrant), Embedders, Retrievers (BM25, embedding, hybrid), Generators, PromptBuilder, conditional routing, and evaluation components. USE WHEN: user mentions "Haystack", "deepset", "Haystack pipeline", "Component DAG", "DocumentStore", "BM25Retriever", "ConditionalRouter" DO NOT USE FOR: LlamaIndex specifics - use `llamaindex`; LangChain -...Votes: 0GitHub stars: 31
- Langgraph RagLangGraph state machines for RAG and agentic flows. Covers typed state, conditional edges for routing (answer/clarify/retrieve/rewrite), checkpointing with SqliteSaver/PostgresSaver, human-in-the-loop interrupts, multi-agent supervisor patterns, Self-RAG and CRAG as explicit graphs, combining with LangChain retrievers. USE WHEN: user mentions "LangGraph", "StateGraph", "agentic RAG", "conditional edges", "checkpointer", "human in the loop", "Self-RAG", "CRAG", "corrective RAG", "supervisor a...Votes: 0GitHub stars: 31
- LlamaindexLlamaIndex 0.12+ for RAG and agent applications. Covers Document/Node model, IngestionPipeline, NodeParser variants, VectorStoreIndex, query engines, sub-question decomposition, router engines, Property Graph Index, LlamaParse integration, and observability callbacks. USE WHEN: user mentions "LlamaIndex", "llama_index", "VectorStoreIndex", "IngestionPipeline", "PropertyGraphIndex", "LlamaParse", "SubQuestionQueryEngine" DO NOT USE FOR: LangChain specifics - use `langchain`; RAG architecture...Votes: 0GitHub stars: 31
- R2rR2R (RAG to Riches) by SciPhi — a production-ready RAG engine with built-in hybrid search, automatic knowledge graph construction, agentic workflows, multi-tenant support, REST + Python SDK, and self-hosted or cloud deployment. USE WHEN: user mentions "R2R", "RAG to Riches", "SciPhi", "R2R SDK", "R2R knowledge graph", "R2R ingestion pipeline" DO NOT USE FOR: DIY retrieval pipelines - use `rag-architecture`; Pinecone-specific stacks - use `canopy`; lightweight embeddings on edge - use `txtai`Votes: 0GitHub stars: 31
- RagatouilleRAGatouille — a high-level wrapper around ColBERTv2 for late-interaction retrieval. Covers RAGPretrainedModel, index creation with PLAID, training custom ColBERT checkpoints with hard negatives, serving as a retriever, integration with LangChain/LlamaIndex. USE WHEN: user mentions "RAGatouille", "ColBERT", "ColBERTv2", "PLAID", "late interaction", "token-level retrieval", "RAGPretrainedModel" DO NOT USE FOR: dense single-vector retrieval - use `rag-architecture`; sparse BM25/SPLADE - use `h...Votes: 0GitHub stars: 31
- Txtaitxtai — lightweight, SQLite-backed embeddings + semantic search + pipeline framework. Covers Embeddings index, pipelines (summarization, translation, QA), workflows, graph support, agents, and edge/on-device deployment with small models. USE WHEN: user mentions "txtai", "SQLite embeddings", "lightweight RAG", "on-device RAG", "edge RAG", "txtai pipeline", "txtai workflow" DO NOT USE FOR: heavy Pinecone stacks - use `canopy`; full KG/agent platforms - use `r2r`; large LangChain/LlamaIndex ap...Votes: 0GitHub stars: 31
- Batch InferenceBatch RAG for high-volume ingest and bulk query scenarios. Covers OpenAI Batch API (50% discount, 24h SLA), Anthropic Message Batches API, Voyage AI and Cohere batch embeddings, ingestion-time vs query-time batching, async/Ray parallelism, parallel writes to vector DBs, and rate-limit coordination across workers. USE WHEN: user mentions "batch API", "OpenAI batch", "Anthropic batches", "bulk embedding", "Ray embeddings", "parallel ingest", "batch RAG" DO NOT USE FOR: streaming single-query ...Votes: 0GitHub stars: 31
- Cost AllocationPer-tenant, per-feature, per-query RAG cost tracking. Covers token counting (tiktoken, Anthropic count_tokens), structured metadata logging, aggregation in BigQuery/Snowflake/ClickHouse, dashboards (Grafana, Metabase), LangSmith and Langfuse native cost reports, and budget alerts. Schema + example queries. USE WHEN: user mentions "RAG cost", "cost per tenant", "cost per query", "token counting", "chargeback", "showback", "LangSmith cost", "Langfuse cost", "budget alerts" DO NOT USE FOR: red...Votes: 0GitHub stars: 31
- Llm GatewayLLM gateways in front of RAG stacks. Covers Portkey (caching, fallbacks, retries, observability), OpenRouter (300+ model routing), LiteLLM Proxy, Kong AI Gateway, semantic caching at gateway layer, cost-based routing (cheap model for easy queries), rate-limit handling, and unified API across providers. Config examples. USE WHEN: user mentions "LLM gateway", "Portkey", "OpenRouter", "LiteLLM", "Kong AI Gateway", "AI gateway", "semantic cache gateway", "provider fallback", "unified LLM API" D...Votes: 0GitHub stars: 31
- Multi RegionMulti-region RAG deployments for latency and resilience. Covers geo-replicated vector stores (Pinecone multi-region, Qdrant cluster, MongoDB Atlas Global, Weaviate), per-region embedding/rerank pools, LLM routing to nearest provider region (Anthropic, OpenAI, Bedrock), eventual-consistency strategies for index updates, and failover patterns. USE WHEN: user mentions "multi-region RAG", "global RAG", "geo replication", "RAG failover", "latency routing", "regional LLM endpoint", "disaster recov...Votes: 0GitHub stars: 31
- Tei Triton ServingHigh-performance serving of embedding and reranker models in production. Covers HuggingFace Text Embeddings Inference (TEI) on GPU, Docker deployment, ONNX/FP16 quantization, dynamic batching, NVIDIA Triton for co-serving embedding + reranker + LLM, and NVIDIA NIM as a managed alternative. Benchmarks and configs. USE WHEN: user mentions "TEI", "Text Embeddings Inference", "Triton", "Triton Inference Server", "NIM", "NVIDIA NIM", "embedding server", "reranker server", "ONNX embedding", "serve...Votes: 0GitHub stars: 31
- Advanced RetrievalRetrieval strategies beyond top-K similarity. Parent-document, small-to-big, multi-vector, contextual compression, sentence-window, auto-merging, RAPTOR, and hierarchical indexing. LangChain + LlamaIndex code with tradeoffs. USE WHEN: user mentions "parent document retriever", "small-to-big", "multi-vector", "sentence window", "auto-merging", "RAPTOR", "hierarchical index", "contextual compression" DO NOT USE FOR: chunking the source docs - use `chunking-strategies`; query rewriting - use `...Votes: 0GitHub stars: 31
- Agentic RagAgent-driven RAG patterns. Self-RAG, Corrective RAG (CRAG) with web fallback, Adaptive RAG with routing classifier, ReAct with retrieval tool, multi-hop retrieval, plan-and-execute, LangGraph state machines for RAG. USE WHEN: user mentions "agentic RAG", "Self-RAG", "Corrective RAG", "CRAG", "Adaptive RAG", "multi-hop retrieval", "LangGraph RAG", "ReAct RAG", "plan and execute RAG" DO NOT USE FOR: static retrieval pipelines - use `rag-architecture`; query rewriting only - use `query-transfo...Votes: 0GitHub stars: 31
- Ares FrameworkARES (Stanford) automated RAG evaluation: synthetic query generation, fine-tuned classifiers for faithfulness and relevance, prediction-powered inference (PPI) for unbiased estimates from a small gold set. Compared to RAGAS. USE WHEN: user mentions "ARES", "Stanford ARES", "prediction-powered inference", "PPI for RAG", "ares-ai", "fine-tuned RAG judges", "automated RAG eval with small gold set" DO NOT USE FOR: LLM-as-judge only workflow - use `rag-evaluation` (RAGAS); Giskard tooling - use ...Votes: 0GitHub stars: 31
- Cdc Streaming IngestionReal-time RAG ingestion. CDC (Debezium, Postgres logical replication), Kafka/ Pulsar topics for doc events, stream processing (Flink, Kafka Streams) to embedding service, exactly-once semantics, late-arriving updates, tombstones (deletes), upsert to vector DB, schema evolution. Full Debezium + Kafka -> vector DB example. USE WHEN: user mentions "CDC RAG", "Debezium RAG", "Kafka RAG", "real-time embeddings", "streaming ingestion", "Flink embeddings", "Pulsar RAG", "logical replication RAG" D...Votes: 0GitHub stars: 31
- Chunking StrategiesDocument chunking techniques for RAG. Fixed-size, recursive, semantic, token-based, document-aware, proposition, parent-child, sliding window, and Anthropic contextual retrieval. Tradeoff tables, LangChain and LlamaIndex code. USE WHEN: user mentions "chunking", "text splitter", "split documents", "semantic chunking", "contextual retrieval", "parent-child chunks", "proposition chunking" DO NOT USE FOR: retrieval after chunking - use `advanced-retrieval`; query-side transforms - use `query-t...Votes: 0GitHub stars: 31
- Contextual RetrievalAnthropic's Contextual Retrieval technique in depth. Prepend LLM-generated chunk-specific context (Claude Haiku) to each chunk before indexing. Combines contextual BM25 + contextual embeddings + reranking for up to 67% retrieval failure reduction. Full production pipeline with prompt caching (90% cost cut), batch processing, and eval numbers. USE WHEN: user mentions "contextual retrieval", "contextual embeddings", "Anthropic contextual retrieval", "chunk context", "contextual BM25", "49% ret...Votes: 0GitHub stars: 31
- Continuous EvaluationCI/CD for RAG quality: golden dataset fixtures, RAGAS/DeepEval in pytest, regression thresholds, GitHub Actions workflows, merge-blocking gates, weekly scheduled eval, LangSmith/Langfuse in CI. USE WHEN: user mentions "RAG CI", "eval in CI", "regression gate", "golden dataset fixture", "PR quality check", "scheduled RAG evaluation", "LangSmith CI", "Langfuse CI" DO NOT USE FOR: RAGAS metric internals - use `rag-evaluation`; ARES - use `ares-framework`; Giskard internals - use `giskard-rag`;...Votes: 0GitHub stars: 31
- Conversational RagMulti-turn RAG: chat history management, context window compaction (summarization, sliding window, vector memory), query rewriting with coreference resolution, follow-up vs new-query routing, LangChain ConversationalRetrievalChain, LlamaIndex ChatEngine, Redis/Postgres message history stores. USE WHEN: user mentions "conversational RAG", "multi-turn RAG", "chat history", "follow-up question", "chat memory", "ChatEngine", "ConversationalRetrievalChain", "coreference in RAG" DO NOT USE FOR: s...Votes: 0GitHub stars: 31
- Domain TemplatesProduction RAG templates for five domains: customer-support, developer-docs, legal, medical, financial. Each covers corpus shape, chunking, embedding model choice, metadata schema, domain-specific guardrails (hallucination tolerance, PII, compliance), and evaluation criteria. USE WHEN: user mentions "RAG for support", "legal RAG", "medical RAG", "developer docs RAG", "financial RAG", "domain template RAG" DO NOT USE FOR: general architecture - use `rag-architecture`; evaluation methodology ...Votes: 0GitHub stars: 31
- Entity ResolutionEntity resolution (ER) for multi-source RAG and knowledge graphs. Dedup entities across documents ("Acme Corp" = "Acme, Inc." = "Acme Corporation"). Covers blocking, probabilistic linking (Splink), Dedupe.io, embedding-based matching, LLM-assisted resolution with rules, graph-based resolution (merged nodes), canonical IDs, and mapping tables. USE WHEN: user mentions "entity resolution", "record linkage", "dedup", "deduplication", "Splink", "Dedupe.io", "canonicalization", "entity matching", ...Votes: 0GitHub stars: 31
- Feedback LoopsUser feedback signals for RAG improvement: thumbs up/down, click-through, dwell time, explicit ratings. Implicit vs explicit signals, logging schema, feedback -> retraining pipelines (embedding fine-tuning with hard negatives, reranker fine-tuning from CTR), A/B testing RAG variants, Langfuse/LangSmith feedback APIs, feature stores for online learning. USE WHEN: user mentions "user feedback", "thumbs up down", "CTR", "dwell time", "RAG evaluation feedback", "hard negatives", "reranker fine-t...Votes: 0GitHub stars: 31
- Giskard RagGiskard RAGET (RAG Evaluation Toolkit): automatic testset generation (simple / complex / distracting / conversational), component-level scoring (retriever / generator / rewriter), hallucination and bias tests, CI integration. Compared to RAGAS and DeepEval. USE WHEN: user mentions "Giskard", "RAGET", "Giskard RAG toolkit", "automatic testset generation", "component-level RAG scoring", "hallucination test Giskard" DO NOT USE FOR: general RAGAS usage - use `rag-evaluation`; Stanford ARES - us...Votes: 0GitHub stars: 31
- Graph RagKnowledge-graph-augmented retrieval. Entity and triple extraction, graph construction (Neo4j, LlamaIndex PropertyGraphIndex), hierarchical community summarization (Microsoft GraphRAG), personalized PageRank (HippoRAG), multi-hop traversal retrieval, and hybrid graph + vector pipelines. USE WHEN: user mentions "GraphRAG", "HippoRAG", "knowledge graph RAG", "entity extraction", "multi-hop reasoning", "Neo4j RAG", "LlamaIndex property graph", "LangChain graph retriever", "triple extraction", "c...Votes: 0GitHub stars: 31
- Hybrid SearchCombining sparse (BM25, SPLADE) and dense vector retrieval. Reciprocal rank fusion with formula and code, weighted score fusion, alpha tuning, and native hybrid indexes in Pinecone, Qdrant, Weaviate. USE WHEN: user mentions "hybrid search", "BM25", "sparse dense", "RRF", "reciprocal rank fusion", "SPLADE", "learned sparse", "alpha tuning", "sparse dense hybrid" DO NOT USE FOR: rewriting queries before retrieval - use `query-transformations`; reranking after retrieval - use `reranking`; vect...Votes: 0GitHub stars: 31
- Ingestion OrchestrationProduction ingestion pipelines with Airflow, Prefect 3, Dagster. DAG design for RAG: extract -> parse -> chunk -> embed -> index. Retry policies, idempotency, partial failure, monitoring, backfills, incremental vs full refresh, data lineage, upstream dependencies. Full Dagster and Prefect examples. USE WHEN: user mentions "ingestion pipeline", "Airflow RAG", "Prefect RAG", "Dagster RAG", "DAG for embeddings", "backfill embeddings", "incremental ingestion", "idempotent ingestion" DO NOT USE ...Votes: 0GitHub stars: 31
- Knowledge Graph ConstructionBuilding knowledge graphs from unstructured text. LLM-based triple extraction (subject-predicate-object), schema-guided extraction via Pydantic + structured output, REBEL model, OpenIE, entity linking to Wikidata/DBpedia, validation with LLM judges, incremental KG updates, and exporting to Neo4j / Amazon Neptune / TigerGraph. Full pipeline. USE WHEN: user mentions "knowledge graph construction", "KG construction", "triple extraction", "OpenIE", "REBEL", "Pydantic triples", "entity linking Wi...Votes: 0GitHub stars: 31
- Long Context Vs RagDecision framework: when long-context (Gemini 2M, Claude 200k, GPT 128k) beats RAG, hybrid approaches (RAG narrows, long-context reads), cost-quality-latency tradeoffs, lost-in-the-middle / context rot research, needle vs synthesis tasks, prompt caching economics, concrete $ per query math. USE WHEN: user mentions "long context vs RAG", "Gemini 2M", "lost in the middle", "context rot", "when not to use RAG", "stuff the prompt", "prompt caching cost" DO NOT USE FOR: implementing RAG - use `r...Votes: 0GitHub stars: 31
- Multimodal RagRetrieval over images, PDFs with figures and tables, audio, and video. Vision-language embeddings (CLIP, SigLIP, ImageBind, VoyageAI multimodal, BGE-M3), vision-model ingestion (Claude Vision, GPT-4o Vision), table-aware retrieval, Whisper-based audio RAG, keyframe + transcript video RAG. USE WHEN: user mentions "multimodal RAG", "image search", "PDF with figures", "table extraction", "audio RAG", "video RAG", "CLIP", "SigLIP", "VoyageAI multimodal", "BGE-M3", "Claude Vision RAG", "GPT-4o Vi...Votes: 0GitHub stars: 31
- Ontology Guided RetrievalOntology-aware RAG. Leveraging domain ontologies (SNOMED CT, FIBO, schema.org, Gene Ontology, custom OWL) to expand queries with synonyms, hyponyms, and broader concepts; boost retrieval for ontology-matched entities; hybrid SPARQL + vector pipelines; RDFLib + embeddings. Examples with medical and financial ontologies. USE WHEN: user mentions "ontology", "SNOMED", "FIBO", "schema.org", "OWL", "SKOS", "RDF", "SPARQL RAG", "taxonomy-guided", "concept expansion", "hyponym retrieval" DO NOT USE...Votes: 0GitHub stars: 31
- Personalization RagUser-specific retrieval. Per-user namespaces/filters, preference embeddings, collaborative signals, reranking with user context (role, history, favorites), privacy-preserving design (encrypted metadata, differential privacy), GDPR- compliant personalization, long-term user memory with mem0/Zep/Letta, graph- based user-entity memory. USE WHEN: user mentions "personalized RAG", "user-specific RAG", "per-user retrieval", "mem0", "Zep", "Letta", "long-term memory", "user preferences RAG" DO NOT...Votes: 0GitHub stars: 31
- Query TransformationsPre-retrieval query rewriting techniques. HyDE, multi-query, step-back, RAG-fusion with RRF, sub-query decomposition, query routing, and expansion. Full Python code per technique with LangChain and native Anthropic SDK. USE WHEN: user mentions "HyDE", "hypothetical document", "multi-query", "step-back prompting", "RAG-fusion", "query rewriting", "query decomposition", "query routing" DO NOT USE FOR: post-retrieval reranking - use `reranking`; sparse+dense fusion on retrieved docs - use `hyb...Votes: 0GitHub stars: 31
- Rag ArchitectureRAG system architecture and design decisions. Covers naive vs advanced vs agentic RAG, decision trees for RAG vs fine-tuning vs long context, production topology, latency budgets, and component sequencing. USE WHEN: user mentions "RAG architecture", "RAG design", "naive RAG", "advanced RAG", "agentic RAG", "RAG vs fine-tuning", "RAG vs long context", "production RAG" DO NOT USE FOR: chunking details - use `chunking-strategies`; query rewriting - use `query-transformations`; retrieval algori...Votes: 0GitHub stars: 31
- Rag CachingCaching strategies across the RAG stack. Semantic caching with GPTCache and LangChain, Redis-based embedding-similarity cache, cache key design, TTL/invalidation, partial caching (cache retrieval only), provider-native prompt caching (Anthropic, OpenAI), and hierarchical L1/L2 caches. USE WHEN: user mentions "semantic cache", "GPTCache", "LLM cache", "prompt caching", "Redis vector cache", "cache invalidation for RAG", "reduce LLM cost", "latency reduction LLM" DO NOT USE FOR: retrieval acc...Votes: 0GitHub stars: 31
- Rag EvaluationEvaluating RAG systems. RAGAS, DeepEval, TruLens, custom LLM-as-judge, golden datasets, synthetic test generation, A/B testing, retrieval-only metrics (Hit@K, MRR, NDCG), and answer quality metrics. USE WHEN: user mentions "RAGAS", "RAG evaluation", "faithfulness", "answer relevancy", "context precision", "hit rate", "MRR", "NDCG", "LLM as judge", "golden dataset", "synthetic data" DO NOT USE FOR: runtime telemetry only - use standard monitoring; chunking decisions - use `chunking-strategie...Votes: 0GitHub stars: 31
- Rag GuardrailsTrust-and-safety layer for RAG. Hallucination detection (LLM self-check, NLI entailment, TRUE metric), groundedness scoring, forced-citation prompting, out-of-scope refusal, NeMo Guardrails, Guardrails AI, Pydantic structured outputs with source verification, LLM-as-judge validators. USE WHEN: user mentions "hallucination detection", "groundedness", "citation enforcement", "NeMo Guardrails", "Guardrails AI", "LLM-as-judge", "refusal", "faithfulness", "TRUE metric", "answer verification" DO ...Votes: 0GitHub stars: 31
- Rag ObservabilityTracing, evaluation, and alerting for RAG systems. LangSmith, Langfuse, Arize Phoenix, Comet Opik, OpenTelemetry GenAI conventions. What to log (query, chunks+scores, rerank scores, answer, citations), retrieval debugging workflows, alerts for empty/low-score retrieval. USE WHEN: user mentions "LangSmith", "Langfuse", "Phoenix", "Opik", "OpenTelemetry LLM", "LLM tracing", "retrieval debugging", "RAG metrics", "RAG dashboard", "RAG alerting" DO NOT USE FOR: hallucination validators - use `ra...Votes: 0GitHub stars: 31