Skip to content
Back to skills

Llamaindex

ASecurity

LlamaIndex is an open-source Python framework for connecting LLMs to your own data: it loads and chunks documents, indexes them in a vector store, and answers questions over them with retrieval (RAG) and tool-calling agents. Use when ingesting documents, configuring retrieval strategies, building query engines, creating multi-step agents, or evaluating answer quality. Trigger words: llamaindex, llama-index, rag, retrieval augmented generation, vector index, query engine, document loader, know...

  • 142 stars
  • 0 votes
  • 0 copies
  • 4 views
  • Added May 29, 2026
ai-agentspythongobashreactnodegitapidocumentation

Works with

  • terminal
  • cli
  • api

Security analysis

A96/100
  • mediumInstalls packages at runtime which could introduce malicious dependencies

Pro scans all 2 files and shows the line behind each finding

Scanned October 4, 2026

npx -y skills add TerminalSkills/skills --skill llamaindex --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Llamaindex?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for Llamaindex
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/terminalskills-llamaindex/badge)](https://www.skillsdirectory.com/skills/terminalskills-llamaindex)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
SKILL.md
---
name: llamaindex
description: >-
  LlamaIndex is an open-source Python framework for connecting LLMs to your own
  data: it loads and chunks documents, indexes them in a vector store, and
  answers questions over them with retrieval (RAG) and tool-calling agents. Use
  when ingesting documents, configuring retrieval strategies, building query
  engines, creating multi-step agents, or evaluating answer quality. Trigger
  words: llamaindex, llama-index, rag, retrieval augmented generation, vector
  index, query engine, document loader, knowledge base.
license: Apache-2.0
compatibility: "Python 3.10+; an LLM and embedding provider (OpenAI by default, or local models through Ollama)"
metadata:
  author: terminal-skills
  version: "1.1.0"
  category: data-ai
  tags: ["llamaindex", "rag", "llm", "retrieval", "knowledge-base"]
  repository: https://github.com/run-llama/llama_index
---

# LlamaIndex

## Overview

LlamaIndex is a data framework for building RAG pipelines, knowledge assistants, and data-augmented LLM applications. It provides document loaders, chunking strategies, several index types, hybrid retrieval with reranking, tool-calling agents, and evaluation tools for question-answering systems. The Python library is `llama-index-core` plus one pip package per integration (300+). This skill follows the 0.14 API: agents are workflow-based (`FunctionAgent`), and the agent classes from before 0.13 no longer exist.

## Instructions

### Install and configure models

```bash
pip install llama-index                 # llama-index-core + the OpenAI LLM and embedding integrations
pip install llama-index-readers-file    # PDF, DOCX, PPTX, CSV parsing for SimpleDirectoryReader
pip install llama-index-vector-stores-chroma llama-index-retrievers-bm25 llama-index-postprocessor-cohere-rerank
```

The package name spells the import: `llama-index-llms-ollama` is `from llama_index.llms.ollama import Ollama`. Set models once on `Settings`; every index, query engine, and evaluator picks them up.

```python
from llama_index.core import Settings
from llama_index.embeddings.openai import OpenAIEmbedding
from llama_index.llms.openai import OpenAI

# Both read OPENAI_API_KEY from the environment. Left unset, the defaults are the
# older gpt-3.5-turbo and text-embedding-ada-002.
Settings.llm = OpenAI(model="gpt-4o-mini")
Settings.embed_model = OpenAIEmbedding(model="text-embedding-3-small")

# Local models instead: pip install llama-index-llms-ollama llama-index-embeddings-ollama
# from llama_index.llms.ollama import Ollama
# from llama_index.embeddings.ollama import OllamaEmbedding
# Settings.llm = Ollama(model="llama3.1", request_timeout=360.0, context_window=8000)
# Settings.embed_model = OllamaEmbedding(model_name="nomic-embed-text")
```

### Load, index, query

```python
from llama_index.core import SimpleDirectoryReader, StorageContext, VectorStoreIndex, load_index_from_storage

documents = SimpleDirectoryReader(
    "docs/handbook", recursive=True, required_exts=[".md", ".pdf"], filename_as_id=True
).load_data()
index = VectorStoreIndex.from_documents(documents)   # chunks, embeds, stores in memory
index.storage_context.persist(persist_dir="storage")

# Later runs: reload instead of re-embedding
index = load_index_from_storage(StorageContext.from_defaults(persist_dir="storage"))
response = index.as_query_engine(similarity_top_k=4).query("How many vacation days do new hires get?")
print(response)
for hit in response.source_nodes:
    print(hit.score, hit.metadata["file_name"])
```

- When chunking, start with `SentenceSplitter` at 1024 tokens with 200 token overlap (its defaults), use `MarkdownNodeParser` for structured documents, `CodeSplitter` for code (needs `tree_sitter_language_pack`), and adjust based on evaluation results.
- When indexing, use `VectorStoreIndex` as the default for most RAG, `PropertyGraphIndex` for entity relationships (`KnowledgeGraphIndex` is deprecated), and `DocumentSummaryIndex` for per-document summaries.
- When building query engines, use `RetrieverQueryEngine` for standard RAG, `CitationQueryEngine` for responses with source attribution, and `SubQuestionQueryEngine` for complex multi-part queries.
- Set `similarity_top_k` based on context window: 3-5 chunks for large models, 2-3 for smaller models.

### Incremental ingestion into a vector store

```python
import os
import chromadb
from llama_index.core.extractors import TitleExtractor
from llama_index.core.ingestion import IngestionPipeline
from llama_index.core.node_parser import SentenceSplitter
from llama_index.core.storage.docstore import SimpleDocumentStore
from llama_index.vector_stores.chroma import ChromaVectorStore

collection = chromadb.PersistentClient(path="chroma_db").get_or_create_collection("handbook")
vector_store = ChromaVectorStore(chroma_collection=collection)

pipeline = IngestionPipeline(
    transformations=[
        SentenceSplitter(chunk_size=1024, chunk_overlap=200),
        TitleExtractor(),            # per document: one LLM call per chunk (first 5) plus one to merge; adds document_title
        Settings.embed_model,        # must be a stage when a vector store is attached
    ],
    docstore=SimpleDocumentStore(),  # remembers a hash per document id
    vector_store=vector_store,
)
if os.path.exists("pipeline_storage"):
    pipeline.load("pipeline_storage")
nodes = pipeline.run(documents=documents)   # unchanged files are skipped, edited ones are replaced
pipeline.persist("pipeline_storage")
index = VectorStoreIndex.from_vector_store(vector_store)
```

Deduplication keys on the document id, so load with `filename_as_id=True`; otherwise every run generates new ids and re-embeds everything.

### Hybrid retrieval and reranking

```python
from llama_index.core.query_engine import RetrieverQueryEngine
from llama_index.core.retrievers import QueryFusionRetriever
from llama_index.postprocessor.cohere_rerank import CohereRerank
from llama_index.retrievers.bm25 import BM25Retriever

chunks = vector_store.get_nodes(node_ids=None)   # every chunk in the Chroma collection
hybrid = QueryFusionRetriever(
    [index.as_retriever(similarity_top_k=10), BM25Retriever.from_defaults(nodes=chunks, similarity_top_k=10)],
    similarity_top_k=10,
    num_queries=1,               # 1 = no LLM-generated query variations
    mode="reciprocal_rerank",
)
engine = RetrieverQueryEngine.from_args(
    hybrid,
    node_postprocessors=[CohereRerank(top_n=4)],   # reads COHERE_API_KEY
)
```

### Agents

```python
import asyncio
from llama_index.core.agent.workflow import FunctionAgent
from llama_index.core.tools import QueryEngineTool
from llama_index.core.workflow import Context

handbook_tool = QueryEngineTool.from_defaults(
    query_engine=engine,
    name="handbook_search",
    description="Answers questions about the employee handbook: time off, expenses, hardware.",
)
agent = FunctionAgent(
    tools=[handbook_tool],
    llm=Settings.llm,
    system_prompt="Answer HR questions. Search the handbook before answering.",
)

async def main():
    ctx = Context(agent)   # carries chat history between runs
    print(await agent.run("How many vacation days do new hires get?", ctx=ctx))
    print(await agent.run("And how many of them carry over?", ctx=ctx))

asyncio.run(main())
```

`FunctionAgent` relies on the provider's native tool calling. For models without it, use `ReActAgent` from the same module; `AgentWorkflow(agents=[...])` lets several agents hand off to each other. Plain Python functions with type hints and a docstring are valid tools.

### Evaluation

```python
from llama_index.core.evaluation import BatchEvalRunner, FaithfulnessEvaluator, RelevancyEvaluator

async def evaluate(engine, questions):
    judge = OpenAI(model="gpt-4o")
    runner = BatchEvalRunner(
        {"faithfulness": FaithfulnessEvaluator(llm=judge), "relevancy": RelevancyEvaluator(llm=judge)},
        workers=4,
    )
    results = await runner.aevaluate_queries(engine, queries=questions)
    for name, evals in results.items():
        print(name, sum(e.passing for e in evals) / len(evals))
```

Faithfulness checks the answer against the retrieved context and relevancy against the question; neither needs labels. `CorrectnessEvaluator` scores 1-5 against a `reference` answer you supply.

## Examples

### Example 1: Build a RAG pipeline over company documentation

**User request:** "Create a question-answering system over our internal docs"

With `Settings`, `documents`, and the ingestion pipeline above in `ask_handbook.py`, add a citation engine:

```python
import sys
from llama_index.core.query_engine import CitationQueryEngine

engine = CitationQueryEngine.from_args(index, similarity_top_k=4, citation_chunk_size=512)
response = engine.query(sys.argv[1])
print(response)
for i, hit in enumerate(response.source_nodes, start=1):
    print(f"[{i}] {hit.metadata['file_name']}")
```

```bash
python ask_handbook.py "How many vacation days do new hires get, and do unused days carry over?"
```

**Output** (wording varies by model):

```
New hires receive 25 vacation days per calendar year, prorated from the start date [1].
Up to 5 unused days carry over and expire on March 31 [1].
[1] time-off.md
[2] expenses.md
```

The second run of the script embeds no documents (only the question), because the pipeline's document store finds every file unchanged.

### Example 2: Create a multi-source research agent

**User request:** "Build an agent that can search our handbook and our runbooks, and also do date math"

```python
from datetime import date
from llama_index.core.agent.workflow import FunctionAgent, ToolCallResult
from llama_index.core.tools import QueryEngineTool

def days_between(start: str, end: str) -> int:
    """Number of days between two ISO dates (YYYY-MM-DD)."""
    return (date.fromisoformat(end) - date.fromisoformat(start)).days

tools = [
    QueryEngineTool.from_defaults(handbook_index.as_query_engine(), name="handbook_search",
        description="HR policies: time off, expenses, hardware purchases."),
    QueryEngineTool.from_defaults(runbook_index.as_query_engine(), name="runbook_search",
        description="Engineering runbooks: deploys, incident response, on-call."),
    days_between,
]
agent = FunctionAgent(tools=tools, llm=Settings.llm, system_prompt="Use the tools; cite which one you used.")

async def ask(question: str):
    handler = agent.run(question)
    async for event in handler.stream_events():
        if isinstance(event, ToolCallResult):
            print("tool:", event.tool_name, event.tool_kwargs)
    print(await handler)
```

**Output** for `asyncio.run(ask("Vacation carry-over expires March 31. How many days is that from 2026-10-01?"))` (which tools are called, and with what arguments, varies by model): one `tool:` line per call — for example `handbook_search`, then `days_between {'start': '2026-10-01', 'end': '2027-03-31'}` — followed by an answer that uses the returned 181 days.

## Guidelines

- Use `SentenceSplitter` with 1024 token chunks and 200 token overlap as the starting point.
- Install `llama-index-readers-file` before loading PDFs or Office files: since 0.14.18 `pip install llama-index` no longer pulls it in, and without it `SimpleDirectoryReader` only logs a warning and reads those files as raw text.
- Code written for 0.12 or earlier breaks: `ServiceContext` raises (use `Settings`), and `ReActAgent.from_tools`, `AgentRunner`, `FunctionCallingAgent`, `OpenAIAgent`, and `QueryPipeline` were removed in 0.13. Agents are async; call them with `await agent.run(...)`.
- Use hybrid retrieval (vector + keyword) for production; pure vector search misses exact term matches. `QueryFusionRetriever` runs its retrievers asynchronously, so a vector store created with only a sync client (Qdrant, for one) fails there — pass its async client too or set `use_async=False`.
- Add a reranker (`CohereRerank`) after retrieval to improve result relevance for small cost: retrieve 10, keep 4.
- Metadata extractors (`TitleExtractor`, `SummaryExtractor`) improve retrieval but cost LLM calls on every new or changed document; add them once the plain pipeline is measured.
- Evaluate with `BatchEvalRunner` on a fixed question set before deploying and after every change to chunking or retrieval; subjective quality assessment does not scale.
- Use `IngestionPipeline` with a docstore for incremental updates; do not re-embed unchanged documents. Persist both the pipeline state and the vector store.
- Documents and questions are sent to the configured LLM and embedding provider. For data that must stay on the machine, set local models on `Settings` before building any index, and re-embed after changing the embedding model — vectors from different models are not comparable.
- The maintainers now focus on the hosted LlamaParse platform and keep the framework as an open toolkit. For hard PDFs (scans, tables) the built-in readers are basic; for a single prompt over a small file, call the model API directly instead of building an index.

Files in this skill

  • SKILL.md4.1 KB
  • _scores.json2 KB

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…