browser-control
A tiny, fast Rust CLI that drives a real browser over the Chrome DevTools Protocol — built for coding agents.
browser-control is a lightweight Rust CLI that enables AI coding agents to control real browsers through the Chrome DevTools Protocol. It exposes browser capabilities as simple shell commands that return stable element references (@e1, @e2) instead of brittle CSS selectors, making it easy for LLM-based agents to interact with web pages. The tool works with both local Chrome instances and remote cloud browser providers like Browser Use, Steel, Hyperbrowser, and Browserbase, providing a shell-native interface without requiring SDKs or long-running servers.
Key Features
Use Cases
- 01Enabling coding agents to perform web research and data extraction by navigating and interacting with live websites
- 02Automating form filling and user workflows across web applications within agent-driven automation scripts
- 03Testing web interfaces through agent-controlled browsers with full access to DOM events and network activity
- 04Driving cloud-hosted browser sessions from AI agents for scalable web automation without local browser requirements
- 05Debugging agent browser interactions by reviewing automatically captured traces of failed commands
- 06Building reproducible web scraping pipelines where agents can retry and adapt based on network and console feedback
Related Skills
View moresuperpowers
An agentic skills framework & software development methodology that works.
skills
Skills for Real Engineers. Straight from my .agents directory.
skills
Public repository for Agent Skills
ponytail
Makes your AI agent think like the laziest senior dev in the room. The best code is the code you never wrote.
browser-control — FAQ
What is browser-control and what does it do?+
browser-control is a Rust CLI tool that lets AI coding agents drive real browsers over the Chrome DevTools Protocol. It provides shell commands that return stable element references, making it easy for LLMs to interact with web pages without writing complex selectors.
How do I install browser-control?+
Install via Cargo with 'cargo install browser-control-cli', download prebuilt binaries from the GitHub releases page, or build from source. You'll need a recent Chrome or Chromium browser; the tool will auto-start one with 'browser-control launch' or connect to an existing CDP endpoint.
Which AI clients and agents work with browser-control?+
browser-control works with any coding agent that can execute shell commands, including Claude, GPT-based agents, and custom agent frameworks. It's language-agnostic and doesn't require SDK integration—agents simply run subcommands and parse the output.
Do I need API keys to use browser-control?+
No API keys are required for local Chrome usage. Cloud browser providers (Browser Use, Steel, Hyperbrowser, Browserbase) require their respective API keys set as environment variables when using the cloud-start command.
Is browser-control free and open source?+
Yes, browser-control is open source under the MIT license and free to use. The tool itself has no cost, though cloud browser providers charge separately for their hosting services.
How does browser-control differ from Puppeteer or Selenium?+
browser-control is optimized for AI agents with shell-native commands and LLM-friendly element references (@e1, @e2) instead of requiring code integration. It includes built-in observability with event traces and works identically across local and cloud browsers without framework lock-in.
How do I install browser-control?+
Open the source repository on GitHub and follow its README. browser-control is a skill — MCP Agents Market links you directly to the official repo.
Is browser-control free?+
browser-control is an open-source project hosted on GitHub. Check the repository for its license and any usage requirements.