Agent Browser CLI
Token-efficient browser automation for AI agents
What You Can Do
Automate Chrome/Chromium browser interactions using compact accessibility-tree snapshots that use 200-400 tokens instead of raw HTML. Navigate pages, click elements, fill forms, extract data, and take screenshots. Run on local Chrome or AWS Bedrock AgentCore cloud browsers with persistent profiles. Supports parallel sessions, MCP integration, and eve agent embedding.
Features
Get accessibility-tree snapshots with compact @eN refs instead of parsing raw HTML. Reduces token overhead by 50-75% compared to raw DOM inspection
Run on local Chrome/Chromium via CDP or use AWS Bedrock AgentCore for managed cloud sessions with automatic credential resolution
Click, fill, type, select dropdowns, upload files, hover, drag-and-drop, and send keyboard input using snapshot refs or semantic locators
Extract text, HTML, attributes, page titles, and URLs. Capture full-page or cropped screenshots programmatically
Fill complex forms, handle multi-step login flows, maintain session state with persistent browser profiles and localStorage
Run multiple independent browser sessions in parallel with isolated state, cookies, and profiles using named session IDs
Read rendered DOM content, fetch documentation with markdown negotiation, filter by sections, and discover llms.txt endpoints
Expose tools via Model Context Protocol server and @agent-browser/eve extension for seamless AI framework embedding
Example Output
Navigate a Page and Extract Data
agent-browser open https://example.com
agent-browser snapshot -i
agent-browser click @e3
agent-browser screenshot result.png
Opens a webpage, takes an interactive-element snapshot showing clickable refs, clicks an element, and captures a screenshot of the result.
Automate Form Submission
agent-browser fill @e2 "user@example.com"
agent-browser fill @e3 "password123"
agent-browser click @e5
agent-browser wait --load networkidle
Fills form fields with credentials, clicks submit, and waits for network idle to ensure the page has fully loaded.
Run on AWS Bedrock AgentCore
AGENTCORE_REGION=us-east-1 agent-browser -p agentcore open https://example.com
agent-browser snapshot -i
agent-browser click @e1
agent-browser close
Launches a managed cloud browser session on AWS, takes a snapshot, interacts with elements, and closes the session.
What's Included
- agent-browser CLI: Fast command-line tool installed via npm, backed by Chrome DevTools Protocol for local or cloud browser control
- Snapshot & Navigation Commands: snapshot, open, close, navigate, press, scroll, read with markdown parsing and section filtering
- Interaction Commands: click, fill, type, select, check, upload, drag, hover, and element inspection tools (text, html, attr, value)
- AWS Bedrock AgentCore Provider: Cloud browser provider with auto-credential resolution, persistent profiles, live-view monitoring, and multi-region support
- Session & Parallel Execution: Named session isolation, multiple independent browsers, auto-restore, configurable idle timeouts, and worktree-scoped sessions
- MCP & Eve Integration: Stdio MCP server for framework embedding, paginated tool discovery, and @agent-browser/eve extension for agent sandboxes
Who It's For
- AI and ML Engineers
- Automation and Bot Developers
- QA and Test Automation Specialists
- Web Data and API Engineers
- Claude Code and Agent Builders
Best For
- AI agent web task automation
- Web scraping with JavaScript execution
- Automated form filling and submission
- Functional and E2E testing
- Multi-step browser workflow automation