Skip to content
Back to skills

Local Scrape

ASecurity

Extract clean article text from web pages for free with Trafilatura — no API credits, no browser needed. Use this skill whenever the user wants page content and the page is static (articles, docs, blogs, news). Always try this BEFORE Firecrawl: it is the default scraping path on the user's limited free tier. Also use when the user says "scrape", "extract article text", or needs bulk text extraction for research or RAG.

  • 6 stars
  • 0 votes
  • 0 copies
  • 0 views
  • Added September 30, 2026
ai-agentsjavascriptpythonjavabashapibackend

Works with

  • cli
  • api

Security analysis

A92/100
  • mediumInstalls packages at runtime which could introduce malicious dependencies

Pro scans all 2 files and shows the line behind each finding

Scanned September 30, 2026

npx -y skills add aicodedecode/awesome-muse-skills --skill local-scrape --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Local Scrape?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for Local Scrape
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/aicodedecode-local-scrape/badge)](https://www.skillsdirectory.com/skills/aicodedecode-local-scrape)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
SKILL.md
---
name: local-scrape
description: Extract clean article text from web pages for free with Trafilatura — no API credits, no browser needed. Use this skill whenever the user wants page content and the page is static (articles, docs, blogs, news). Always try this BEFORE Firecrawl: it is the default scraping path on the user's limited free tier. Also use when the user says "scrape", "extract article text", or needs bulk text extraction for research or RAG.
---

# Local Scrape (free, zero-credit)

Trafilatura-based article extraction running locally in `~/workspace/.venvs/scrape`. Fetches with urllib (proxy-aware on this VM) and strips boilerplate — nav, footers, ads, cookie banners — leaving clean article text. Verified 2026-09-30: 7,415 clean chars from a JS-free news article.

## The scraping decision tree (credit-saving order)

1. **Static page → this skill** (free). Articles, docs, blogs, most news.
1b. **Fetch blocked or page needs light JS → Jina Reader** (free, no key, no install): `curl -sL "https://r.jina.ai/<url>"` returns clean Markdown with title + published time. Verified working from this VM 2026-09-30. Try this before the browser tools for single pages — it's one HTTP call. (Idea salvaged from evaluating Panniantong/agent-reach, which was declined: its other backends duplicate working capabilities or need desktop login-state absent here.)
2. **JavaScript-heavy or interactive → browser tools** (`browser.open`, `browser.spawn_task`). Still free. Do NOT script Playwright/Chromium from exec — the runtime routes browser work through the browser tools.
3. **Anti-bot blocked, or bulk structured extraction where 1+2 fail → Firecrawl** (costs credits, last resort — see the `firecrawl` skill). The user is on a limited free tier.

## Usage

```bash
# Single page -> JSON on stdout
~/workspace/.venvs/scrape/bin/python ~/workspace/skills/local-scrape/scripts/scrape.py https://example.com/article

# Many pages in parallel -> file
~/workspace/.venvs/scrape/bin/python ~/workspace/skills/local-scrape/scripts/scrape.py \
  $(cat urls.txt) --out /tmp/pages.json --workers 8
```

Output per URL: `{url, title, text, char_count, ok}`. `ok: false` with an `error` string when the fetch fails or the page needs JavaScript.

## Notes

- Trafilatura returns plain text (not markdown). Tables are preserved via `include_tables`.
- If `ok` is false for a page that visibly has content, the page likely needs JavaScript — escalate to the browser tools, not Firecrawl, first.
- `crawl4ai` is also installed in the venv, but driving a browser from exec is restricted in this environment — treat it as unavailable and use the browser tools for rendered pages.
- Keep the venv at `~/workspace/.venvs/scrape`; system pip is blocked (PEP 668), never `pip install` outside a venv.

Files in this skill

  • SKILL.md2.8 KB
  • scripts/scrape.py2.2 KB

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…