</>MCP Agents Market
Skill

node-crawler

by bda-research6.8kTypeScriptUpdated 2026-06-18

Web Crawler/Spider for NodeJS + server-side jQuery ;-)

OpenClawHermesClaude CodeCodex

Node-crawler is an agent skill that teaches AI assistants how to perform large-scale web scraping and crawling using a NodeJS-based crawler with server-side jQuery capabilities. The skill instructs agents on queueing versus direct requests, rate limiting, proxy rotation, Cheerio parsing, and common pitfalls when crawling websites. It integrates with AI agents like OpenClaw, Hermes, Claude Code, and Codex to enable programmatic web data extraction. The skill bundles documentation and best practices for the crawler library, which features configurable connection pools, priority queues, automatic charset detection, and HTTP/2 support.

Key Features

Server-side DOM manipulation with Cheerio (jQuery-like API) for parsing HTML
Configurable connection pooling and automatic retry logic with exponential backoff
Priority queue system with adjustable rate limiting per request or rate limiter group
Proxy rotation support with independent rate limiters per proxy
Automatic charset detection and UTF-8 conversion for international sites
HTTP/2 protocol support alongside traditional HTTP/1.1 requests
Direct request mode and queued crawling with preRequest hooks
Skip duplicate URL detection and homogeneous queue reallocation

Use Cases

  • 01Training AI agents to scrape product data from e-commerce sites with rate limits
  • 02Teaching assistants to crawl documentation sites while respecting server load
  • 03Enabling agents to extract structured data from HTML pages using CSS selectors
  • 04Instructing agents to download binary files (images, PDFs) with proper encoding handling
  • 05Guiding agents to rotate proxies when crawling at scale to avoid IP blocks
  • 06Helping agents implement respectful crawling with configurable delays between requests

Related Skills

View more

node-crawler — FAQ

What is the node-crawler agent skill?+

The node-crawler skill is a knowledge package that teaches AI agents how to use the node-crawler library for web scraping tasks. It includes documentation on queue management, rate limiting, proxy rotation, HTML parsing with Cheerio, and avoiding common crawling pitfalls.

How do I install the node-crawler skill for my AI agent?+

You can install via ClawHub using 'openclaw skills install node-crawler', download the skill bundle from the GitHub releases page and unzip it into your agent's skills directory, or ask your agent to download and install it by providing the release URL. The skill folder should be placed in directories like .openclaw/skills/, .claude/skills/, or wherever your agent runtime loads skills.

Which AI clients and agents work with this skill?+

The skill is compatible with OpenClaw, Hermes, Claude Code, Codex, and other AI agents that support skill loading from a designated skills directory. The agent automatically loads the skill when tasks involve large-scale scraping or crawling.

What are the prerequisites for using node-crawler?+

Node.js version 22 or above is required. The crawler library itself is installed via npm (npm install crawler). No API keys are needed for the crawler, though target websites may have their own authentication requirements.

Is the node-crawler skill free to use?+

Yes, the node-crawler library and agent skill are open source and free to use. The skill is hosted on GitHub and can be downloaded or installed without cost.

Does node-crawler support HTTP/2 and proxy rotation?+

Yes, the crawler supports both HTTP/2 protocol (by setting http2: true) and proxy rotation. You can specify multiple proxies and assign different rate limiters to each proxy for independent request throttling.

How do I install node-crawler?+

Open the source repository on GitHub and follow its README. node-crawler is a skill — MCP Agents Market links you directly to the official repo.

Is node-crawler free?+

node-crawler is an open-source project hosted on GitHub. Check the repository for its license and any usage requirements.

Related searches