Browser use

Agentic AIProtocols and integrationPublished By Simon Budziak

Browser use is the capability that lets an AI agent operate a web browser the way a person does: reading pages, clicking, typing into forms, and navigating across sites to finish a task. It is how agents reach the enormous share of business software that offers a login page but no API.

How does browser use actually work?

The agent perceives the page, through the accessibility tree, the HTML, screenshots, or a mix, decides an action, executes it, and looks again. Modern implementations favor structured page reading over raw pixels because it is cheaper and less brittle; hosted variants run the browser in the provider’s cloud and hand the AI agent element references instead of coordinates. The loop of perceive, act, and re-check is what separates browser use from scripted automation, which replays fixed steps and breaks when the layout shifts.

Why does browser use matter to a business?

Because the long tail of work lives in web apps with no API: supplier portals, government forms, insurance platforms, legacy admin panels. Tool calling covers systems with clean interfaces; browser use covers everything else. The browser is the universal API of last resort, and agents that can drive one can automate processes that were never designed for automation.

Browser use vs computer use: what is the difference?

Scope. Computer use gives an agent the whole desktop, screen, mouse, and keyboard, across any application. Browser use restricts the surface to the browser, which makes it easier to secure, cheaper to run, and accurate enough for most business tasks, since most business software is web software. Most companies should exhaust browser use before granting desktop control, for the same reason you scope any credential tightly.

What are the risks to manage?

An agent in a browser holds real sessions and real permissions, so a page can try to redirect it through hidden instructions, and a wrong click can submit rather than draft. The controls are the usual ones: dedicated scoped accounts, sandboxed execution, human approval on irreversible steps, and logs of every action the agent took.

This entry was drafted with AI assistance.

Frequently asked questions

What is browser use in AI?

An AI agent controlling a real web browser to complete tasks: logging in, filling forms, extracting data, and clicking through flows. The agent reads page structure or screenshots, decides the next action, and adapts when the page changes.

How is browser use different from web scraping?

Scraping extracts data with scripts that break when layouts change. Browser use is goal-driven: the agent interprets the page and adapts, so it can complete multi-step tasks such as submitting a form or reconciling records across two web apps.

Is Browser Use also a tool name?

Yes, an open source project of that name popularized the pattern, and Anthropic, OpenAI, and Google ship browser tools for their agents. The concept matters more than any one tool: a model plus a controlled browser, with permissions set by you.

Summarize this page with

See this working in a system we built