Skip to content
Back to skills

Crawlee

ASecurity

Build reliable web scrapers and crawlers with Crawlee — Apify's open-source framework for structured web scraping. Use when someone asks to "scrape a website", "build a crawler", "Crawlee", "web scraping at scale", "scrape JavaScript-rendered pages", "crawl with Playwright/Puppeteer", or "extract data from websites reliably". Covers HTTP crawling, browser crawling, request queues, proxy rotation, and data export.

  • 142 stars
  • 0 votes
  • 0 copies
  • 5 views
  • Added May 27, 2026
developmentjavascripttypescriptgojavabashreactvuenodegitapi

Works with

  • terminal
  • api

Security analysis

A96/100
  • mediumInstalls packages at runtime which could introduce malicious dependencies

Pro scans all 2 files and shows the line behind each finding

Scanned October 4, 2026

npx -y skills add TerminalSkills/skills --skill crawlee --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Crawlee?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for Crawlee
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/terminalskills-crawlee/badge)](https://www.skillsdirectory.com/skills/terminalskills-crawlee)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
SKILL.md
---
name: crawlee
description: >-
  Build reliable web scrapers and crawlers with Crawlee — Apify's open-source
  framework for structured web scraping. Use when someone asks to "scrape a
  website", "build a crawler", "Crawlee", "web scraping at scale", "scrape
  JavaScript-rendered pages", "crawl with Playwright/Puppeteer", or "extract
  data from websites reliably". Covers HTTP crawling, browser crawling,
  request queues, proxy rotation, and data export.
license: Apache-2.0
compatibility: "Node.js 16+ (20+ recommended). Optional: Playwright or Puppeteer for JS-rendered pages."
metadata:
  author: terminal-skills
  version: "1.1.0"
  repository: https://github.com/apify/crawlee
  category: data-ai
  tags: ["scraping", "crawling", "crawlee", "apify", "playwright"]
---

# Crawlee

## Overview

Crawlee is a web scraping and crawling library that handles the hard parts — request queuing, retries, proxy rotation, browser fingerprinting, and rate limiting. Use Cheerio for fast HTML-only scraping or Playwright/Puppeteer for JavaScript-rendered pages. Built-in storage for datasets, request queues, and key-value stores. Scales from single pages to millions of URLs.

## When to Use

- Scraping data from websites (product prices, job listings, articles)
- Crawling entire sites for content or link analysis
- JavaScript-rendered pages (SPAs, React/Vue sites)
- Scraping at scale with proxy rotation and anti-blocking
- Structured data extraction with automatic retries

## Instructions

### Setup

```bash
npx crawlee create my-crawler        # official template, or add to an existing project:
npm install crawlee                  # CheerioCrawler / HttpCrawler only
npm install crawlee playwright       # browser crawling (Playwright or Puppeteer are separate installs)
npx playwright install chromium
```

Use ES modules (`"type": "module"` in package.json) so top-level `await crawler.run(...)` works. Checked against crawlee 3.18.

### HTTP Crawling with Routing (Fast, No Browser)

Give start requests a `label` and route them to handlers with a router; `enqueueLinks` stays on the same hostname by default and deduplicates URLs.

```typescript
// scraper.ts
import { CheerioCrawler, Dataset, createCheerioRouter } from "crawlee";

const router = createCheerioRouter();

router.addHandler("LISTING", async ({ enqueueLinks }) => {
  await enqueueLinks({ selector: "article.product_pod h3 a", label: "DETAIL" });
  await enqueueLinks({ selector: "li.next a", label: "LISTING" }); // pagination
});

router.addHandler("DETAIL", async ({ request, $, pushData }) => {
  await pushData({
    url: request.loadedUrl,
    title: $("h1").text().trim(),
    price: $(".price_color").first().text().trim(),
    scrapedAt: new Date().toISOString(),
  });
});

const crawler = new CheerioCrawler({
  requestHandler: router,
  maxConcurrency: 10,
  maxRequestRetries: 3,
  maxRequestsPerMinute: 300,       // politeness cap
  maxRequestsPerCrawl: 500,        // safety limit while developing
  respectRobotsTxtFile: true,      // skip URLs disallowed by robots.txt
  failedRequestHandler({ request, log }) {
    log.error(`Failed: ${request.url} after ${request.retryCount} retries`);
  },
});

await crawler.run([{ url: "https://books.toscrape.com/", label: "LISTING" }]);

const dataset = await Dataset.open();
await dataset.exportToCSV("books");   // saved in the key-value store as books.csv
```

Run it with `npx tsx scraper.ts`. Results land in `./storage/datasets/default/*.json` and the CSV in `./storage/key_value_stores/default/books.csv`. Set `CRAWLEE_STORAGE_DIR` to move the folder.

### Browser Crawling (JavaScript-Rendered Pages)

```typescript
// browser-scraper.ts
import { PlaywrightCrawler } from "crawlee";

const crawler = new PlaywrightCrawler({
  maxConcurrency: 5,           // browsers are heavy
  headless: true,              // set false to watch while developing

  async requestHandler({ page, request, pushData, infiniteScroll, log }) {
    await page.waitForSelector(".product-card");   // wait for the rendered content
    await infiniteScroll({ timeoutSecs: 20 });     // built-in helper for lazy lists

    const items = await page.$$eval(".product-card", (cards) =>
      cards.map((card) => ({
        name: card.querySelector("h3")?.textContent?.trim(),
        price: card.querySelector(".price")?.textContent?.trim(),
      }))
    );
    log.info(`${items.length} products on ${request.url}`);
    await pushData(items.map((item) => ({ ...item, sourceUrl: request.url })));
  },
});

await crawler.run(["https://shop.acme-outdoors.com/products"]);
```

Crawlee's browser crawlers apply browser fingerprints by default, so extra launch flags to hide automation are rarely needed. Prefer `waitForSelector` or `locator` waits over fixed `waitForTimeout` sleeps.

### Proxy Rotation

```typescript
import { CheerioCrawler, ProxyConfiguration } from "crawlee";

const proxyConfiguration = new ProxyConfiguration({
  proxyUrls: [
    `http://${process.env.PROXY_USER}:${process.env.PROXY_PASS}@proxy-eu.acme-proxies.net:8080`,
    `http://${process.env.PROXY_USER}:${process.env.PROXY_PASS}@proxy-us.acme-proxies.net:8080`,
  ],
});

const crawler = new CheerioCrawler({
  proxyConfiguration,
  useSessionPool: true,   // ties cookies and proxies to sessions; blocked sessions are retired
  async requestHandler({ request, proxyInfo, log }) {
    log.info(`${request.url} via ${proxyInfo?.hostname}`);
  },
});
```

## Examples

### Example 1: Scrape product data from an e-commerce site

**User prompt:** "Scrape all product names, prices, and ratings from books.toscrape.com and export to CSV."

The agent installs crawlee, builds a CheerioCrawler with a LISTING/DETAIL router following the "next" links, runs it, and exports `books.csv` from the default key-value store. The run log ends with `Finished! Total N requests: N succeeded, 0 failed`.

### Example 2: Monitor competitor prices

**User prompt:** "Build a daily scraper that checks competitor prices and alerts when they change."

The agent builds a PlaywrightCrawler for the JS-rendered shop, saves prices into a named dataset (so the default purge does not erase history), compares each run with the previous one, and prints or sends the price changes.

## Guidelines

- **Cheerio for static HTML** - much faster and lighter than a browser; use Playwright only when content needs JavaScript.
- **Routers and labels** - one handler per page type keeps listing and detail logic apart.
- **`pushData` for output** - writes to the default dataset; export with `exportToCSV` / `exportToJSON` or read with `getData()`.
- **Default storage is purged at the start of each run** - copy results out, or use a named storage (`Dataset.open("books")`) to keep data across runs. Do not rely on it for incremental "compare with last run" without a named dataset.
- **Be polite** - keep `respectRobotsTxtFile: true`, set `maxRequestsPerMinute`, and check the site's terms before scraping.
- **Limit while developing** - `maxRequestsPerCrawl` stops a runaway crawl.
- **Handle failures** - `failedRequestHandler` runs after retries are exhausted; check `request.retryCount`.
- **Proxy credentials from the environment** - never hardcode them in source.

Files in this skill

  • SKILL.md6.2 KB
  • _scores.json1.6 KB

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…