Skip to content
Back to skills

Data Scraper

ASecurity

Extract data from websites, APIs, and files. Supports structured extraction, pagination, and anti-bot bypass.

  • 15 stars
  • 0 votes
  • 0 copies
  • 1 view
  • Added September 12, 2026
ai-agentsbashapi

Works with

  • api

Security analysis

A100/100

Pro scans all 2 files and shows the line behind each finding

Scanned September 12, 2026

npx -y skills add clowlove/Hermes-House --skill data-scraper --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Data Scraper?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for Data Scraper
[![Security: A β€” Skills Directory](https://www.skillsdirectory.com/api/skills/clowlove-data-scraper/badge)](https://www.skillsdirectory.com/skills/clowlove-data-scraper)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
SKILL.md
---
name: data-scraper
description: "Extract data from websites, APIs, and files. Supports structured extraction, pagination, and anti-bot bypass."
triggers:
  - "scrape website"
  - "ηˆ¬θ™«"
  - "extract data"
  - "web scraping"
---

# Data Scraper

Extract data from websites, APIs, and files.

## Features

- 🌐 **Web Scraping** β€” Extract from any website
- πŸ“‘ **API Scraping** β€” GET/POST requests
- πŸ“„ **File Parsing** β€” CSV, Excel, PDF, JSON
- πŸ”„ **Pagination** β€” Auto-handle paginated content
- πŸ›‘οΈ **Anti-bot** β€” Bypass protection
- πŸ“Š **Data Cleaning** β€” Clean extracted data

## Usage

### Scrape Website

```bash
# Simple scrape
hermes scrape "https://example.com/products" \
  --selector ".product-item" \
  --fields name,price,image

# With pagination
hermes scrape "https://example.com/products" \
  --selector ".product" \
  --fields name,price \
  --pages 10
```

### Extract from API

```bash
hermes scrape-api "https://api.example.com/data" \
  --headers "Authorization: Bearer TOKEN" \
  --output data.json
```

### File Parsing

```bash
# Parse CSV
hermes scrape file data.csv --format json

# Parse Excel
hermes scrape file report.xlsx --sheet "Sales"

# Parse PDF
hermes scrape file document.pdf --extract text
```

## Configuration

```yaml
# scraper-config.yml
scraping:
  delay: 1000  # ms between requests
  retries: 3
  timeout: 30
  
anti_bot:
  rotate_user_agent: true
  use_proxy: false
  bypass_cloudflare: true
  
output:
  format: json
  encoding: utf-8
```

## Extractors

### CSS Selector
```bash
hermes scrape --selector ".product .title" --attr text
hermes scrape --selector "img.product" --attr src
```

### JSON Path
```bash
hermes scrape-json --path "$.data[*].name" --file data.json
```

### XPath
```bash
hermes scrape --xpath "//div[@class='product']/span" --file page.html
```

## Pitfalls

1. **Legal Issues** β€” Check robots.txt and terms of service
2. **Rate Limits** β€” Don't overload servers
3. **Data Quality** β€” Verify extracted data accuracy
4. **Anti-bot** β€” Some sites block scrapers

Files in this skill

  • SKILL.md2 KB
  • skill.json498 B

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…