Sep 24, 2026
BrowserAct Hands Your Agent a Browser Built for Anti-Bot Walls
A 6k-star skill gives coding agents a real-browser CLI with stealth fingerprints, CAPTCHA solving, and human handoff — so the scrape that dies at the block page finally finishes.
Headless browsers are cheap until a block page eats your agent's afternoon. BrowserAct Skills — a 6k-star MIT-licensed CLI built for AI agents — wraps real-browser automation in the layers it takes to get past those walls: stealth fingerprints, CAPTCHA solving, residential proxies, and a live human handoff when the machine gets stuck.
Why This Skill Matters
BrowserAct stacks three progressive layers against blocking. The environment layer spoofs fingerprints, rotates TLS, and switches proxies, so most blocks never trigger. The execution layer adds solve-captcha for CAPTCHAs and stealth-extract, which pulls a protected page in one command. The human layer is remote-assist: it generates a live URL you open from any device, you take over in the browser, and the agent continues seamlessly when you are done.
The output is designed for LLM reasoning rather than human scripts: state returns an indexed list of clickable elements, so the next move is click 3 or input 2 "hi" — no DOM parsing, and the README says the compact text format is several times more token-efficient than JSON or HTML.
Installation
The README's install path is a message to your agent:
Install browser-act. Skill source: https://github.com/browser-act/skills/tree/main/browser-act . Verify it works after installation.
Your agent loads the skill, discovers browser state with get-skills, and runs commands from your local environment. The documented compatibility list covers Windows, macOS, and Linux, with Claude Code, Cursor, VS Code, OpenCode, OpenClaw, Codex, and Gemini CLI named as hosts.
Real Workflow: Extract a Protected Page, Then Automate a Flow
- Start with the zero-config path — one command pulls page content through the stealth layer:
browser-act stealth-extract https://example.com
- For a multi-step flow, open a named session and drive it by index:
browser-act --session my-task browser open <id> https://example.com
browser-act --session my-task state # See clickable elements
browser-act --session my-task click 3 # Click by index
browser-act --session my-task input 2 "hi" # Type into a field
- Pick the browser mode to match the scenario:
chromemode reuses your local Chrome login state for authenticated sites, stealth privacy mode gives every session a fresh fingerprint with no residue, and stealth fixed identity keeps one stable fingerprint plus IP for logged-in accounts running in parallel.
Concurrency is built in — separate browsers get independent cookies, fingerprints, and proxies so sites cannot correlate them, while same-browser sessions share login state without blocking each other.
Tips
- Sensitive operations are gated: browser create/delete, profile imports, proxy changes, and privacy toggles require your explicit approval, and the README notes prior approvals do not carry over.
- Almost everything is free. Only managed proxies (Dynamic/Static) and stealth browsers beyond the first 5 are paid.
- Every browser carries a
descfield matched to tasks by meaning — name sessions descriptively and multi-agent runs stay conflict-free. - For a site you scrape repeatedly, look at Skill Forge: it explores the site once and generates a deploy-ready skill package, and the Solutions Catalog already ships 30+ pre-built skills for Amazon, Google Maps, YouTube, and more.
When Not to Use This
If the site offers an official API, use the API — a real browser is the expensive path. If you want zero local setup and lower per-run cost, the README's other mode is fully cloud-managed execution, where BrowserAct builds and runs a reusable scraping bot for you; the local skill is for when you need local Chrome login state or direct integration into your own agent workflow.
See the leaderboard for more skills.