Skip to content
Back to skills

Open Semantic Search Guide

ASecurity

Self-hosted semantic search and text mining platform

  • 3,639 stars
  • 0 votes
  • 0 copies
  • 1 view
  • Added June 6, 2026
devopspythongobashnodedockergitapi

Works with

  • cli
  • api

Security analysis

A100/100

Scanned June 6, 2026

npx -y skills add brycewang-stanford/Auto-Empirical-Research-Skills --skill open-semantic-search-guide --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Open Semantic Search Guide?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for Open Semantic Search Guide
[![Security: A β€” Skills Directory](https://www.skillsdirectory.com/api/skills/brycewang-stanford-open-semantic-search-guide/badge)](https://www.skillsdirectory.com/skills/brycewang-stanford-open-semantic-search-guide)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
SKILL.md
---
name: open-semantic-search-guide
description: "Self-hosted semantic search and text mining platform"
metadata:
  openclaw:
    emoji: "πŸ”Ž"
    category: "literature"
    subcategory: "search"
    keywords: ["semantic search", "text mining", "self-hosted", "Solr", "NLP", "entity extraction"]
    source: "https://github.com/opensemanticsearch/open-semantic-search"
---

# Open Semantic Search Guide

## Overview

Open Semantic Search is a self-hosted search and text mining platform that combines full-text search (Apache Solr) with semantic analysis β€” entity extraction, named entity recognition, text classification, and knowledge graph building. Process and search across documents (PDF, DOCX, emails) with faceted navigation and visual analytics. Ideal for researchers needing private, on-premise document search over large paper collections.

## Installation

```bash
# Docker deployment (recommended)
git clone https://github.com/opensemanticsearch/open-semantic-search.git
cd open-semantic-search
docker-compose up -d

# Access web UI at http://localhost:8080
# Admin panel at http://localhost:8080/admin
```

## Architecture

```
Documents (PDF, DOCX, HTML, email)
         ↓
   Connector/Crawler (file system, web, IMAP)
         ↓
   ETL Pipeline
   β”œβ”€β”€ Text extraction (Apache Tika)
   β”œβ”€β”€ OCR (Tesseract, for scanned docs)
   β”œβ”€β”€ NER (spaCy, Stanford NER)
   β”œβ”€β”€ Entity linking (knowledge base)
   └── Classification (custom models)
         ↓
   Apache Solr (full-text index + facets)
         ↓
   Web UI (search, browse, visualize)
```

## Indexing Documents

```bash
# Index a directory of papers
curl -X POST "http://localhost:8080/api/index" \
  -H "Content-Type: application/json" \
  -d '{"path": "/data/papers/", "recursive": true}'

# Index single file
curl -X POST "http://localhost:8080/api/index" \
  -H "Content-Type: application/json" \
  -d '{"path": "/data/papers/attention.pdf"}'

# Schedule recurring index
# Add to crontab or use built-in scheduler
```

## Search Features

```markdown
### Full-Text Search
- Boolean queries: "attention mechanism" AND transformer
- Phrase search: "self-attention"
- Wildcard: transform*
- Proximity: "attention transformer"~5 (within 5 words)
- Field-specific: title:"attention" author:"Vaswani"

### Faceted Navigation
- Filter by: author, date, organization, topic, language
- Nested facets for hierarchical browsing
- Date range slider
- Entity type filters (person, organization, location)

### Semantic Features
- Named entity highlighting in results
- Related entity suggestions
- Concept co-occurrence visualization
- Auto-generated tag clouds
```

## Python Client

```python
import requests

SEARCH_URL = "http://localhost:8080/api/search"

def search_papers(query, filters=None, max_results=20):
    """Search indexed documents."""
    params = {
        "q": query,
        "rows": max_results,
        "fl": "title,author,content_type,date,score",
        "hl": "true",        # Highlight matches
        "hl.fl": "content",  # Highlight in content field
        "facet": "true",
        "facet.field": ["author", "organization", "topic"],
    }
    if filters:
        params["fq"] = filters

    resp = requests.get(SEARCH_URL, params=params)
    data = resp.json()

    results = data["response"]["docs"]
    facets = data.get("facet_counts", {}).get("facet_fields", {})

    return results, facets

# Search
results, facets = search_papers(
    "attention mechanism transformer",
    filters='date:[2023-01-01T00:00:00Z TO *]',
)

for doc in results:
    print(f"[{doc.get('date', 'N/A')}] {doc.get('title', 'Untitled')}")
    print(f"  Score: {doc['score']:.2f}")
```

## Entity Extraction Configuration

```json
{
  "ner": {
    "engines": ["spacy", "stanford"],
    "models": {
      "spacy": "en_core_web_lg",
      "stanford": "english.all.3class.caseless"
    },
    "entity_types": [
      "PERSON", "ORG", "GPE", "DATE",
      "WORK_OF_ART", "EVENT"
    ],
    "custom_entities": {
      "METHODOLOGY": ["transformer", "CNN", "RNN", "GAN"],
      "DATASET": ["ImageNet", "CIFAR", "MNIST", "COCO"]
    }
  },
  "classification": {
    "enabled": true,
    "model": "custom_topic_classifier",
    "categories": ["NLP", "CV", "RL", "Theory"]
  }
}
```

## Knowledge Graph

```python
# Query the auto-built knowledge graph
def get_entity_network(entity, depth=2):
    """Get co-occurring entities for a given entity."""
    resp = requests.get(
        f"{SEARCH_URL}/graph",
        params={"entity": entity, "depth": depth},
    )
    graph = resp.json()

    for node in graph["nodes"]:
        print(f"Entity: {node['label']} ({node['type']})")
    for edge in graph["edges"]:
        print(f"  {edge['source']} ↔ {edge['target']} "
              f"(co-occur: {edge['weight']})")

get_entity_network("Transformer")
```

## Use Cases

1. **Paper search**: Full-text search over local paper collections
2. **Literature mining**: Extract entities and relationships from papers
3. **Institutional repository**: Campus-wide document search
4. **Due diligence**: Search across legal/business document archives
5. **Investigative research**: Cross-reference entities across documents

## References

- [Open Semantic Search GitHub](https://github.com/opensemanticsearch/open-semantic-search)
- [Apache Solr](https://solr.apache.org/)
- [Apache Tika](https://tika.apache.org/)

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…