Skip to content

Stage 8 — Agent Interfaces: Browser Use · Computer Use · Sandbox

Earlier stages teach an agent what to think about and which tool to request. This stage teaches a different question: which door should it use to do the work? A larger door reaches more things and creates more risk. Start with the smallest door whose result you can check.

📌 Learning goals

After this stage, you can:

  • Look at a task and choose search, webpage control, full-computer control, or isolated execution.
  • Explain eight core terms in your own words.
  • Draw the allowed sites, actions, and human-confirmation points before an agent acts.
  • Finish a small exercise without signing in, downloading files, or touching a real account.
  • Ask what a benchmark measures, how it scores, and how many steps it permits before trusting one number.

🚪 Entry requirements

If you followed the main path, you can first revisit the previous stage: Stage 7.5 Advanced Agentic Concepts. You only need the Stage 03 loop: model proposes a tool call → code executes it → result returns to the model. Track A may stop after Exercise 1; Track B can continue to Exercise 2.

📚 Required reading

Look at the four official entry points, then read the eight terms and the choice table. For your first pass, only learn what each entry point is for.

Time and environment

Allow 45–90 minutes for the visible path and Exercise 1. Set aside another half day if you will build an executor or sandbox.

Environment: Exercise 1 needs only an isolated browser profile. Exercise 2 needs Python 3.10+, makes no network request, and needs no API key.

Reading order:

  1. Anthropic Computer Use tool: learn that the model proposes actions and the application executes them.
  2. Anthropic Browser Use tool: see how page elements and pixel fallback work together.
  3. OpenAI Computer Use guide: study the GA tool and its safety boundary.
  4. OpenAI Agents SDK Sandbox guide: read this only for mutable workspaces; Sandbox Agents are still Beta.

🔑 Eight core terms

Agent Interface

The door an agent uses to see, control, or execute work. Search, browser, desktop, and isolated execution are doors of different sizes.

Browser Use

Use it when all work stays on webpages. It can read text, buttons, and forms, and fall back to screenshots and coordinates when needed.

Computer Use

Use it when work crosses desktop applications. The model reads a screenshot and proposes mouse or keyboard actions; software you control performs them.

Sandbox

A separate workroom for code. It sees only the files, network, and tools you provide, so a mistake is less likely to harm the host.

Accessibility Tree

A page map prepared for assistive technology. It labels text, buttons, inputs, roles, and states; it is not the entire raw HTML document.

Harness

The control program around the model. It receives actions, checks policy, executes, returns results, limits turns, and keeps inspectable records.

Approval Gate

A brake at the doorway. It always stops for a person before payment, sign-in, sending, deletion, or another hard-to-reverse action.

Prompt Injection

Malicious instructions hidden in webpage content that try to replace the agent's real rules. Treat page text as untrusted input, not higher-priority commands.

🧭 Choose the smallest interface first

Your task Start with Simple reason
Only find or read public information Web Search / Fetch You need data, not screen clicks.
All work stays on webpages Browser Use It understands buttons, fields, and tabs, so the door is smaller than a whole computer.
Work crosses desktop applications Computer Use Only then do you need screenshots, mouse, and keyboard.
Run generated code or change files Sandbox Put the code in an isolated room before inspecting its result.

Prefer a formal API or typed tool. If a service already exposes a clear API, use it first. GUI control is a fallback when necessary, not a smarter shortcut.

How to choose Search, Browser Use, Computer Use, or Sandbox
Open full-size image (new tab)

Read the map by asking what the task truly needs, then choose the smallest door that can finish it. The four cards are choices, not levels you must climb in order.

🖱 Computer Use: complete loop, current tools, and legacy migration

The basic loop is:

  1. The executor captures a screenshot.
  2. The model reads it and returns one action or a batch.
  3. The harness checks allowlists and approvals.
  4. The executor performs allowed actions.
  5. A new screenshot and result go back to the model until completion or a stop condition.

Anthropic's current computer_toolset_20260801 is a client toolset. It supplies screenshot, click, type, and other member tools, but your application executes every call. Official documentation

New OpenAI integrations use the Responses API shape tools=[{"type": "computer"}]. computer-use-preview and computer_use_preview are deprecated and remain only for legacy migration; the current response can include batched actions[]. Official documentation

Do not bind the interface definition to one model ID. Samples and migration tables on the same official page can update at different times. Lock the tool contract and choose a model from the implementation-day documentation.

📏 OSWorld: how to read a Computer Use benchmark

OSWorld 2.0 contains 108 long-horizon workflows. The median human completion time is about 1.6 hours. Under one named model, harness, thinking setting, and 500-step budget, the official primary binary-completion high is 20.6%. Those figures describe that setup, not a permanent ranking for every desktop task.

Ask four questions before comparing:

  • Are the tasks the same? OSWorld 1 and 2.0 have different difficulty, so subtracting their percentages is invalid.
  • How is completion scored? Binary completion and partial score are different metrics.
  • How many steps and tokens are allowed? Different budgets are not directly comparable.
  • Are the executor and environment the same? Model, tool batching, parser, and retries all affect the result.

🌐 Browser Use: page elements, Accessibility Tree, and pixel fallback

Anthropic's current browser_toolset_20260801 is a client toolset. It can read pages, find elements, fill forms, switch tabs, and use screenshots and coordinates. Your application still operates the browser. Official documentation

Do not collapse three signals into one:

Signal What it provides When it helps
DOM Nodes and attributes used by webpage code. Reading structure or using selectors.
Accessibility Tree Human-meaningful roles, names, and states. Finding buttons, fields, and operable elements.
Screenshot / pixel What the page actually looks like. Canvas, images, drag-and-drop, or missing structural signals.

Playwright MCP connects browser control to an MCP-capable client. browser-use helps study or build a full web-agent loop. Neither means safe access to every signed-in website out of the box.

Versus scraping: scraping mainly retrieves data; Browser Use also interacts. Versus traditional RPA: RPA often follows prewritten fixed steps; an agent chooses the next step from page state, which demands tighter limits and verification.

📦 Sandbox: isolation technology, workspaces, and providers
Term Plain meaning Important limit
Container An isolated room that shares the host kernel. Bad configuration can still expose the host or network.
Virtual Machine (VM) A room with its own operating-system kernel. Usually heavier than a container.
microVM A smaller, faster VM design. Not every sandbox uses a microVM.
Firecracker An open-source AWS microVM technology. A technology name is not a complete security policy.
gVisor A user-space kernel layer between a program and host kernel. Compatibility and performance require testing.
Cold start Wait time from no environment to executable. Image, region, and measurement method change it; there is no fixed winner.
Workspace Files the agent can see for this job. Include only task-required files.
Session A live sandbox instance that can continue work. It is not conversational memory.
Snapshot Saved workspace state used to start again later. Remove secrets and temporary files first.

OpenAI Agents SDK separates the agent definition, fresh-workspace contract, and per-run sandbox choice through SandboxAgent, Manifest, and SandboxRunConfig. This area remains Beta. Official documentation

Do not compare only startup speed. Check filesystem boundaries, network policy, secret injection, lifecycle, snapshots, logs, region, price, and cleanup. Modal Sandboxes also documents different network and runtime controls, so providers are not one interchangeable isolation type.

🛡️ Four safety checks

Check Ask before action
1. Isolate Is it in a fresh browser profile, container, or VM?
2. Allowlist Which sites, files, tools, and actions are permitted?
3. Approve Which actions must stop and ask a person?
4. Verify & Log What evidence proves success, and can a failure be traced?
Four safety checks around agent actions
Open full-size image (new tab)

Design all four checks together, but do not treat them as one fixed nested technical stack. Any action may be stopped by one or several checks.

🧭 Track A: choosing a ready-made tool
  • For summaries or finding information, start with built-in search or fetch and leave automation off.
  • For website-only tasks, choose Browser Use with a domain allowlist, action preview, and confirmation.
  • For cross-app work, place Computer Use in a dedicated profile or VM with test data.
  • For long tasks, write stop conditions and completion evidence first; background does not mean unchecked.

Official help still describes Gemini in Chrome as a gradual rollout, so it is not available to everyone. Desktop, mobile, region, language, account, and administrator settings also differ. Google Chrome Help

Do not bypass region, account, or organizational policy when a product is unavailable. Choose another tool at the same layer or return to Search / Fetch.

🧭 Track B: executor, framework, and sandbox paths

Choose one canonical path:

  1. Anthropic Computer Use: read the computer-use demo in claude-quickstarts, including executor and container boundaries.
  2. Web-agent loop: start with browser-use, using a test site and fresh profile.
  3. MCP browser executor: use Playwright MCP, limiting origins and permissions at the client.
  4. Isolated code: use E2B or a container you control, with network off and a narrow workspace first.
  5. Stateful workspace agent: then read OpenAI Sandbox Agents; it remains Beta and its API can change.

Every path still needs action validation, approval, timeout / turn limits, result verification, and cleanup. A framework cannot decide your business risk automatically.

For chapter-length implementation, follow the canonical quickstarts below instead of duplicating an SDK textbook that quickly goes stale.

🛠 Hands-on exercises

Exercise 1 (Track A): open only one safe demo page

Copy this directly into your browser or computer agent:

Open only this page: <https://example.com>
Report the page title and final URL, and attach one screenshot.
Do not sign in, download, or leave example.com.
If the page asks for anything else, stop and tell me.

Check the title, URL, and screenshot yourself. If the agent leaves the allowlist, the exercise failed.

Budget: a local or included-subscription tool may add $0 in API cost; APIs and managed browsers follow provider pricing.

Exercise 2 (Track B): check first, then execute

Copy and run:

from urllib.parse import urlparse

ALLOWED_DOMAINS = {"example.com"}
ALLOWED_SCHEMES = {"https"}
LOW_IMPACT_ACTIONS = {"read", "screenshot"}
HIGH_IMPACT_ACTIONS = {"login", "purchase", "delete", "send"}


def check_action(url: str, action: str) -> str:
    parsed = urlparse(url)
    normalized_action = action.strip().casefold()
    if (
        parsed.scheme not in ALLOWED_SCHEMES
        or parsed.hostname not in ALLOWED_DOMAINS
        or parsed.username is not None
        or parsed.password is not None
    ):
        return "BLOCK"
    if normalized_action in HIGH_IMPACT_ACTIONS:
        return "ASK"
    if normalized_action in LOW_IMPACT_ACTIONS:
        return "ALLOW"
    return "BLOCK"


assert check_action("https://example.com", "read") == "ALLOW"
assert check_action("https://example.com", " Login ") == "ASK"
assert check_action("https://example.com", "upload_credentials") == "BLOCK"
assert check_action("file://example.com/report", "read") == "BLOCK"
assert check_action("https://evil.example", "read") == "BLOCK"
print("policy checks passed")

This is not a complete sandbox. It teaches the outer policy. Only then should an ALLOW action reach the executor, with results and screenshots written to a log.

Budget: this local Python costs $0 and makes no API call.

Exercise 3: isolate code

Put a read-only CSV, output folder, and plotting script into a sandbox without host credentials. Disable unnecessary networking, then retrieve only the image and log. The result is evidence that output came from isolation, not merely a successful process.

Exercise 4: complete action loop

On a test site, connect observe → propose actions → policy check → approve / execute → verify. Send one URL outside the allowlist and prove that it is blocked. Do not use payment, real sign-in, email, or Slack as practice data.

⚠️ Safety cases: indirect prompt injection and protected accounts

Brave's research shows that malicious instructions can hide in content an agent reads. This is not a bug class limited to one browser; any agent that reads untrusted content and can act needs defenses.

Perplexity's BrowseSafe response explains its defense direction, but a provider classifier does not replace isolation, allowlists, approvals, and verification.

The Amazon case should not be reduced to “one browser was banned from Amazon.” The Ninth Circuit opinion dated 2026-08-04 discusses a district-court preliminary injunction concerning password-protected Amazon sections. The district court order gives the fuller scope. This is litigation context, not legal advice or a universal product-availability rule.

Choose only one to start:

  • Desktop loop: Anthropic Computer Use tool.
  • Web agent: Anthropic Browser Use tool or Playwright MCP.
  • Isolated code: OpenAI Sandbox guide or E2B.
  • Research: OSWorld 2.0.
  • Attack surface: Brave indirect prompt injection research.

📚 21 complete learning resources and limits

Checked 2026-08-28 UTC. Stars are this project's teaching ratings, not GitHub stars.

GroupResourceUse it whenLimit / statusRating
Official interface docsAnthropic Computer Use toolUnderstand the desktop action loop.Client toolset; your application supplies the executor.⭐⭐⭐⭐⭐
Anthropic Browser Use toolKeep a task inside webpages.Client toolset; requires a controlled browser.⭐⭐⭐⭐⭐
OpenAI Computer Use guideImplement the GA computer tool.The old preview shape is deprecated.⭐⭐⭐⭐⭐
OpenAI Agents SDK Sandbox guideNeed a stateful workspace.Sandbox Agents are Beta.⭐⭐⭐⭐
Google Chrome Help: Gemini in ChromeCheck whether your account has access.gradual rollout with platform and region limits.⭐⭐⭐
Executor / frameworkanthropics/claude-quickstartsRead the official computer-use demo.Inspect container, credentials, and network boundaries first.⭐⭐⭐⭐⭐
browser-use/browser-useBuild a full web-agent loop.You still own production browser scaling and safety.⭐⭐⭐⭐⭐
microsoft/playwright-mcpConnect a browser to an MCP client.Restrict origins, permissions, and data.⭐⭐⭐⭐⭐
trycua/cuaStudy a cross-platform computer-use stack.Verify the actual backend from current README and releases.⭐⭐⭐⭐
bytedance/UI-TARS-desktopStudy an open desktop agent.Local control is high risk; use a test environment.⭐⭐⭐⭐
Sandbox / runtimee2b-dev/E2BAn agent needs a remote code workspace.Apache-2.0 repo; managed service has separate cost and policy.⭐⭐⭐⭐⭐
cloudflare/sandbox-sdkRun isolated code on Workers and Containers.Apache-2.0; Beta, and APIs may change before v1.0.⭐⭐⭐⭐
Modal SandboxesNeed managed containers and runtime controls.Configure network defaults and Beta / VM features from current docs.⭐⭐⭐⭐
Vercel SandboxAlready build isolated execution in Vercel.Check runtime, region, network, and pricing.⭐⭐⭐⭐
GUI / benchmark / datasetmicrosoft/OmniParserStudy screenshot element parsing.The repository is CC-BY-4.0; do not automatically apply that license to the weights.⭐⭐⭐⭐
OSWorld 2.0Evaluate long-horizon desktop tasks.Read scores with metric, step budget, and harness.⭐⭐⭐⭐⭐
xlang-ai/OSWorldReproduce the original cross-OS benchmark.Its task set differs from 2.0; percentages are not directly comparable.⭐⭐⭐⭐⭐
web-arena-x/webarenaEvaluate self-hosted web tasks.Environment setup and evaluator affect results.⭐⭐⭐⭐
OSU-NLP-Group/Mind2WebStudy demonstrations from real websites.A dataset does not make current sites safe to automate.⭐⭐⭐⭐
Safety research and responseBrave: indirect prompt injectionBuild a browser-agent threat model.A research demo is not proof of every product's current state.⭐⭐⭐⭐
Perplexity BrowseSafeCompare a provider response and defense direction.Read provider claims alongside independent testing.⭐⭐⭐

Read OmniParser weights by version: icon_detect_v3 uses the MIT-licensed YOLOv9 implementation; earlier Ultralytics detectors retain AGPL; caption models use MIT. None of these is synonymous with the repository's CC-BY-4.0 license.

💡 Future interfaces: Voice agents and VLA

A voice agent listens and speaks. VLA (Vision-Language-Action) lets a model see and control a physical machine. They are not the same layer as Browser / Computer / Sandbox, so this stage keeps only three entry points:

The whole-site coherence layer will decide which specialist path owns them. This roadmap does not promise a nonexistent next stage.

✅ Self-check

  • I choose the smallest interface first instead of sending every task to Computer Use.
  • I can explain all eight terms and know Browser Use is not only DOM.
  • I isolate, allowlist, require approval, and verify results and logs.
  • I completed the example.com exercise without leaving the allowed scope.
  • When reading OSWorld scores, I also find the task set, metric, step budget, and harness.

You have now completed the main path. Choose a specialist path: researcher, developer, teacher, knowledge worker, or everyday user.