Skip to content
Back to skills

Blog Scraper

ASecurity

Scrape blog posts via RSS feeds (free, no API key) with Apify fallback for JS-heavy sites. Use when you need to monitor competitor blogs, track industry content, or aggregate blog posts by keyword.

  • 3 stars
  • 0 votes
  • 0 copies
  • 2 views
  • Added May 29, 2026
content-marketingpythonbashapi

Works with

  • cli
  • api

Security analysis

A92/100
  • mediumInstalls packages at runtime which could introduce malicious dependencies

Pro scans all 3 files and shows the line behind each finding

Scanned May 29, 2026

npx -y skills add levalencia/agent-god-mode --skill blog-scraper --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Blog Scraper?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for Blog Scraper
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/levalencia-blog-scraper/badge)](https://www.skillsdirectory.com/skills/levalencia-blog-scraper)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
SKILL.md
---
name: blog-scraper
description: >
  Scrape blog posts via RSS feeds (free, no API key) with Apify fallback for
  JS-heavy sites. Use when you need to monitor competitor blogs, track industry
  content, or aggregate blog posts by keyword.
---

# Blog Scraper

Scrape blog posts via RSS/Atom feeds (free) with optional Apify fallback for JS-heavy sites.

## Quick Start

For RSS mode (free), only dependency is `pip install requests`. No API key needed.

```bash
# Scrape a blog's RSS feed
python3 skills/blog-scraper/scripts/scrape_blogs.py \
  --urls "https://growthx.ai/blog" --days 30

# Multiple blogs with keyword filter
python3 skills/blog-scraper/scripts/scrape_blogs.py \
  --urls "https://blog1.com,https://blog2.com" --keywords "AI,marketing" --output summary

# Force Apify for JS-heavy sites
python3 skills/blog-scraper/scripts/scrape_blogs.py \
  --urls "https://example.com" --mode apify
```

## How It Works

### Auto Mode (default)
1. For each URL, tries to discover an RSS/Atom feed:
   - Checks HTML `<link rel="alternate">` tags
   - Probes common paths: `/feed`, `/rss`, `/atom.xml`, `/feed.xml`, `/rss.xml`, `/blog/feed`, `/index.xml`
2. Parses discovered feeds (supports RSS 2.0 and Atom)
3. If any URLs fail, falls back to Apify `jupri/rss-xml-scraper` (if token available)
4. Applies date and keyword filtering client-side

### RSS Mode
Only tries RSS feeds, no Apify fallback.

### Apify Mode
Uses Apify actor directly, skipping RSS discovery.

## CLI Reference

| Flag | Default | Description |
|------|---------|-------------|
| `--urls` | *required* | Blog URL(s), comma-separated |
| `--keywords` | none | Keywords to filter (comma-separated, OR logic) |
| `--days` | 30 | Only include posts from last N days |
| `--max-posts` | 50 | Max posts to return |
| `--mode` | auto | `auto` (RSS + fallback), `rss` (RSS only), `apify` (Apify only) |
| `--output` | json | Output format: `json` or `summary` |
| `--token` | env var | Apify token (only needed for Apify mode/fallback) |
| `--timeout` | 300 | Max seconds for Apify run |

## Cost

- **RSS mode:** Free (no API, no tokens)
- **Apify mode:** Uses `jupri/rss-xml-scraper` — minimal Apify credits

Files in this skill

  • SKILL.md2.1 KB
  • scripts/scrape_blogs.py14.7 KB
  • skill.meta.json246 B

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…