Back to skills
SKILL.md
Data Scraper
ASecurityExtract data from websites, APIs, and files. Supports structured extraction, pagination, and anti-bot bypass.
- 15 stars
- 0 votes
- 0 copies
- 1 view
- Added September 12, 2026
Works with
Security analysis
100/100Pro scans all 2 files and shows the line behind each finding
npx -y skills add clowlove/Hermes-House --skill data-scraper --agent claude-codeAre you the author of Data Scraper?
Add the live security badge to your README. It updates with every re-scan.
[](https://www.skillsdirectory.com/skills/clowlove-data-scraper)---
name: data-scraper
description: "Extract data from websites, APIs, and files. Supports structured extraction, pagination, and anti-bot bypass."
triggers:
- "scrape website"
- "η¬θ«"
- "extract data"
- "web scraping"
---
# Data Scraper
Extract data from websites, APIs, and files.
## Features
- π **Web Scraping** β Extract from any website
- π‘ **API Scraping** β GET/POST requests
- π **File Parsing** β CSV, Excel, PDF, JSON
- π **Pagination** β Auto-handle paginated content
- π‘οΈ **Anti-bot** β Bypass protection
- π **Data Cleaning** β Clean extracted data
## Usage
### Scrape Website
```bash
# Simple scrape
hermes scrape "https://example.com/products" \
--selector ".product-item" \
--fields name,price,image
# With pagination
hermes scrape "https://example.com/products" \
--selector ".product" \
--fields name,price \
--pages 10
```
### Extract from API
```bash
hermes scrape-api "https://api.example.com/data" \
--headers "Authorization: Bearer TOKEN" \
--output data.json
```
### File Parsing
```bash
# Parse CSV
hermes scrape file data.csv --format json
# Parse Excel
hermes scrape file report.xlsx --sheet "Sales"
# Parse PDF
hermes scrape file document.pdf --extract text
```
## Configuration
```yaml
# scraper-config.yml
scraping:
delay: 1000 # ms between requests
retries: 3
timeout: 30
anti_bot:
rotate_user_agent: true
use_proxy: false
bypass_cloudflare: true
output:
format: json
encoding: utf-8
```
## Extractors
### CSS Selector
```bash
hermes scrape --selector ".product .title" --attr text
hermes scrape --selector "img.product" --attr src
```
### JSON Path
```bash
hermes scrape-json --path "$.data[*].name" --file data.json
```
### XPath
```bash
hermes scrape --xpath "//div[@class='product']/span" --file page.html
```
## Pitfalls
1. **Legal Issues** β Check robots.txt and terms of service
2. **Rate Limits** β Don't overload servers
3. **Data Quality** β Verify extracted data accuracy
4. **Anti-bot** β Some sites block scrapersFiles in this skill
- SKILL.md
- skill.json
Attribution
Comments
Loading commentsβ¦