node-crawler
Web Crawler/Spider for NodeJS + server-side jQuery ;-)
Node-crawler is an agent skill that teaches AI assistants how to perform large-scale web scraping and crawling using a NodeJS-based crawler with server-side jQuery capabilities. The skill instructs agents on queueing versus direct requests, rate limiting, proxy rotation, Cheerio parsing, and common pitfalls when crawling websites. It integrates with AI agents like OpenClaw, Hermes, Claude Code, and Codex to enable programmatic web data extraction. The skill bundles documentation and best practices for the crawler library, which features configurable connection pools, priority queues, automatic charset detection, and HTTP/2 support.
Key Features
Use Cases
- 01Training AI agents to scrape product data from e-commerce sites with rate limits
- 02Teaching assistants to crawl documentation sites while respecting server load
- 03Enabling agents to extract structured data from HTML pages using CSS selectors
- 04Instructing agents to download binary files (images, PDFs) with proper encoding handling
- 05Guiding agents to rotate proxies when crawling at scale to avoid IP blocks
- 06Helping agents implement respectful crawling with configurable delays between requests
Related Skills
View moresuperpowers
An agentic skills framework & software development methodology that works.
skills
Skills for Real Engineers. Straight from my .agents directory.
skills
Public repository for Agent Skills
ponytail
Makes your AI agent think like the laziest senior dev in the room. The best code is the code you never wrote.
node-crawler — FAQ
What is the node-crawler agent skill?+
The node-crawler skill is a knowledge package that teaches AI agents how to use the node-crawler library for web scraping tasks. It includes documentation on queue management, rate limiting, proxy rotation, HTML parsing with Cheerio, and avoiding common crawling pitfalls.
How do I install the node-crawler skill for my AI agent?+
You can install via ClawHub using 'openclaw skills install node-crawler', download the skill bundle from the GitHub releases page and unzip it into your agent's skills directory, or ask your agent to download and install it by providing the release URL. The skill folder should be placed in directories like .openclaw/skills/, .claude/skills/, or wherever your agent runtime loads skills.
Which AI clients and agents work with this skill?+
The skill is compatible with OpenClaw, Hermes, Claude Code, Codex, and other AI agents that support skill loading from a designated skills directory. The agent automatically loads the skill when tasks involve large-scale scraping or crawling.
What are the prerequisites for using node-crawler?+
Node.js version 22 or above is required. The crawler library itself is installed via npm (npm install crawler). No API keys are needed for the crawler, though target websites may have their own authentication requirements.
Is the node-crawler skill free to use?+
Yes, the node-crawler library and agent skill are open source and free to use. The skill is hosted on GitHub and can be downloaded or installed without cost.
Does node-crawler support HTTP/2 and proxy rotation?+
Yes, the crawler supports both HTTP/2 protocol (by setting http2: true) and proxy rotation. You can specify multiple proxies and assign different rate limiters to each proxy for independent request throttling.
How do I install node-crawler?+
Open the source repository on GitHub and follow its README. node-crawler is a skill — MCP Agents Market links you directly to the official repo.
Is node-crawler free?+
node-crawler is an open-source project hosted on GitHub. Check the repository for its license and any usage requirements.