Browser-act uses a two-layer Skill model:
The entry Skill makes the agent aware that Browser-act exists.
The
get-skillsruntime returns current environment state, command guidance, and dynamic instructions.
Install the entry Skill
Ask your agent:
Install the browser-act Skill from: https://github.com/browser-act/skills/tree/main/browser-act
After installation, the agent can recognize Browser-act tasks and call get-skills before using the CLI.
Two layers
flowchart TD User["Browser task"] --> Entry["Entry Skill"] Entry --> Runtime["get-skills runtime"] Runtime --> State["Runtime state"] Runtime --> Rules["Safety rules"] Runtime --> Version["Version guidance"]
Layer 1: entry Skill
The installed Skill file is intentionally small. It:
triggers agent awareness
contains activation language
tells the agent to load runtime content through
get-skillschanges rarely
Layer 2: get-skills
The first command for a Browser-act task should be:
browser-act get-skills core --skill-version <version>
The output gives the agent:
CLI and API key status
available browsers and descriptions
active sessions
core command guidance
dynamic instructions for the current environment
Topics
core
get-skills core --skill-version <v> returns the core workflow, commands, browser state, and safety rules for most tasks.
advanced
get-skills advanced returns proxy, profile import, privacy mode, and advanced browser setup guidance.
main
get-skills main returns the latest Skill content when a version mismatch is detected.
Progressive loading
Tip: Most tasks only need get-skills core. Load advanced topics only when the task requires them. This keeps the agent context focused and reduces irrelevant instructions.
Version compatibility
The --skill-version parameter lets the CLI detect mismatches:
browser-act get-skills core --skill-version 2.0.2
When versions are incompatible, the output includes direct upgrade guidance:
CLI too old
Run uv tool upgrade browser-act-cli.
Skill too old
Run browser-act get-skills main.
Dynamic instructions
get-skills adjusts guidance based on the current state:
Situation | Instruction returned |
Multiple browsers exist | Choose by |
Active sessions exist | Respect session ownership and naming |
API key is missing | Avoid stealth-only features or guide authentication |
Sensitive browser is selected | Ask before opening a |
Agent workflow
flowchart TD Ask["User asks"] --> Activate["Skill activates"] Activate --> Load["Run get-skills"] Load --> Read["Read state"] Read --> Select["Select browser"] Select --> Confirm["Confirm if needed"] Confirm --> Run["Run commands"] Run --> Finish["Close session"]
Agent compatibility
Browser-act can work with agents that can:
load Skill files
run shell commands
read text output
Known compatible environments include Claude Code, GitHub Copilot, Cursor, Windsurf, Gemini CLI, OpenCode, Codex, and similar agent runtimes.
Learn more
Quick Start
Run the first Browser-act automation loop.
Command Reference
Open the full command index.
Skill Forge
Generate reusable task-specific skills.
