Stage 8 — Agent Interfaces: Browser Use · Computer Use · Sandbox¶
Earlier stages teach an agent what to think about and which tool to request. This stage teaches a different question: which door should it use to do the work? A larger door reaches more things and creates more risk. Start with the smallest door whose result you can check.
📌 Learning goals¶
After this stage, you can:
- Look at a task and choose search, webpage control, full-computer control, or isolated execution.
- Explain eight core terms in your own words.
- Draw the allowed sites, actions, and human-confirmation points before an agent acts.
- Finish a small exercise without signing in, downloading files, or touching a real account.
- Ask what a benchmark measures, how it scores, and how many steps it permits before trusting one number.
🚪 Entry requirements¶
If you followed the main path, you can first revisit the previous stage: Stage 7.5 Advanced Agentic Concepts. You only need the Stage 03 loop: model proposes a tool call → code executes it → result returns to the model. Track A may stop after Exercise 1; Track B can continue to Exercise 2.
📚 Required reading¶
Look at the four official entry points, then read the eight terms and the choice table. For your first pass, only learn what each entry point is for.
Time and environment
Allow 45–90 minutes for the visible path and Exercise 1. Set aside another half day if you will build an executor or sandbox.
Environment: Exercise 1 needs only an isolated browser profile. Exercise 2 needs Python 3.10+, makes no network request, and needs no API key.
Reading order:
- Anthropic Computer Use tool: learn that the model proposes actions and the application executes them.
- Anthropic Browser Use tool: see how page elements and pixel fallback work together.
- OpenAI Computer Use guide: study the GA tool and its safety boundary.
- OpenAI Agents SDK Sandbox guide: read this only for mutable workspaces; Sandbox Agents are still Beta.
🔑 Eight core terms¶
Agent Interface¶
The door an agent uses to see, control, or execute work. Search, browser, desktop, and isolated execution are doors of different sizes.
Browser Use¶
Use it when all work stays on webpages. It can read text, buttons, and forms, and fall back to screenshots and coordinates when needed.
Computer Use¶
Use it when work crosses desktop applications. The model reads a screenshot and proposes mouse or keyboard actions; software you control performs them.
Sandbox¶
A separate workroom for code. It sees only the files, network, and tools you provide, so a mistake is less likely to harm the host.
Accessibility Tree¶
A page map prepared for assistive technology. It labels text, buttons, inputs, roles, and states; it is not the entire raw HTML document.
Harness¶
The control program around the model. It receives actions, checks policy, executes, returns results, limits turns, and keeps inspectable records.
Approval Gate¶
A brake at the doorway. It always stops for a person before payment, sign-in, sending, deletion, or another hard-to-reverse action.
Prompt Injection¶
Malicious instructions hidden in webpage content that try to replace the agent's real rules. Treat page text as untrusted input, not higher-priority commands.
🧭 Choose the smallest interface first¶
| Your task | Start with | Simple reason |
|---|---|---|
| Only find or read public information | Web Search / Fetch | You need data, not screen clicks. |
| All work stays on webpages | Browser Use | It understands buttons, fields, and tabs, so the door is smaller than a whole computer. |
| Work crosses desktop applications | Computer Use | Only then do you need screenshots, mouse, and keyboard. |
| Run generated code or change files | Sandbox | Put the code in an isolated room before inspecting its result. |
Prefer a formal API or typed tool. If a service already exposes a clear API, use it first. GUI control is a fallback when necessary, not a smarter shortcut.
Read the map by asking what the task truly needs, then choose the smallest door that can finish it. The four cards are choices, not levels you must climb in order.
🖱 Computer Use: complete loop, current tools, and legacy migration
The basic loop is:
- The executor captures a screenshot.
- The model reads it and returns one action or a batch.
- The harness checks allowlists and approvals.
- The executor performs allowed actions.
- A new screenshot and result go back to the model until completion or a stop condition.
Anthropic's current computer_toolset_20260801 is a client toolset. It supplies screenshot, click, type, and other member tools, but your application executes every call. Official documentation
New OpenAI integrations use the Responses API shape tools=[{"type": "computer"}]. computer-use-preview and computer_use_preview are deprecated and remain only for legacy migration; the current response can include batched actions[]. Official documentation
Do not bind the interface definition to one model ID. Samples and migration tables on the same official page can update at different times. Lock the tool contract and choose a model from the implementation-day documentation.
📏 OSWorld: how to read a Computer Use benchmark
OSWorld 2.0 contains 108 long-horizon workflows. The median human completion time is about 1.6 hours. Under one named model, harness, thinking setting, and 500-step budget, the official primary binary-completion high is 20.6%. Those figures describe that setup, not a permanent ranking for every desktop task.
Ask four questions before comparing:
- Are the tasks the same? OSWorld 1 and 2.0 have different difficulty, so subtracting their percentages is invalid.
- How is completion scored? Binary completion and partial score are different metrics.
- How many steps and tokens are allowed? Different budgets are not directly comparable.
- Are the executor and environment the same? Model, tool batching, parser, and retries all affect the result.
🌐 Browser Use: page elements, Accessibility Tree, and pixel fallback
Anthropic's current browser_toolset_20260801 is a client toolset. It can read pages, find elements, fill forms, switch tabs, and use screenshots and coordinates. Your application still operates the browser. Official documentation
Do not collapse three signals into one:
| Signal | What it provides | When it helps |
|---|---|---|
| DOM | Nodes and attributes used by webpage code. | Reading structure or using selectors. |
| Accessibility Tree | Human-meaningful roles, names, and states. | Finding buttons, fields, and operable elements. |
| Screenshot / pixel | What the page actually looks like. | Canvas, images, drag-and-drop, or missing structural signals. |
Playwright MCP connects browser control to an MCP-capable client. browser-use helps study or build a full web-agent loop. Neither means safe access to every signed-in website out of the box.
Versus scraping: scraping mainly retrieves data; Browser Use also interacts. Versus traditional RPA: RPA often follows prewritten fixed steps; an agent chooses the next step from page state, which demands tighter limits and verification.
📦 Sandbox: isolation technology, workspaces, and providers
| Term | Plain meaning | Important limit |
|---|---|---|
| Container | An isolated room that shares the host kernel. | Bad configuration can still expose the host or network. |
| Virtual Machine (VM) | A room with its own operating-system kernel. | Usually heavier than a container. |
| microVM | A smaller, faster VM design. | Not every sandbox uses a microVM. |
| Firecracker | An open-source AWS microVM technology. | A technology name is not a complete security policy. |
| gVisor | A user-space kernel layer between a program and host kernel. | Compatibility and performance require testing. |
| Cold start | Wait time from no environment to executable. | Image, region, and measurement method change it; there is no fixed winner. |
| Workspace | Files the agent can see for this job. | Include only task-required files. |
| Session | A live sandbox instance that can continue work. | It is not conversational memory. |
| Snapshot | Saved workspace state used to start again later. | Remove secrets and temporary files first. |
OpenAI Agents SDK separates the agent definition, fresh-workspace contract, and per-run sandbox choice through SandboxAgent, Manifest, and SandboxRunConfig. This area remains Beta. Official documentation
Do not compare only startup speed. Check filesystem boundaries, network policy, secret injection, lifecycle, snapshots, logs, region, price, and cleanup. Modal Sandboxes also documents different network and runtime controls, so providers are not one interchangeable isolation type.
🛡️ Four safety checks¶
| Check | Ask before action |
|---|---|
| 1. Isolate | Is it in a fresh browser profile, container, or VM? |
| 2. Allowlist | Which sites, files, tools, and actions are permitted? |
| 3. Approve | Which actions must stop and ask a person? |
| 4. Verify & Log | What evidence proves success, and can a failure be traced? |
Design all four checks together, but do not treat them as one fixed nested technical stack. Any action may be stopped by one or several checks.
🧭 Track A: choosing a ready-made tool
- For summaries or finding information, start with built-in search or fetch and leave automation off.
- For website-only tasks, choose Browser Use with a domain allowlist, action preview, and confirmation.
- For cross-app work, place Computer Use in a dedicated profile or VM with test data.
- For long tasks, write stop conditions and completion evidence first; background does not mean unchecked.
Official help still describes Gemini in Chrome as a gradual rollout, so it is not available to everyone. Desktop, mobile, region, language, account, and administrator settings also differ. Google Chrome Help
Do not bypass region, account, or organizational policy when a product is unavailable. Choose another tool at the same layer or return to Search / Fetch.
🧭 Track B: executor, framework, and sandbox paths
Choose one canonical path:
- Anthropic Computer Use: read the computer-use demo in claude-quickstarts, including executor and container boundaries.
- Web-agent loop: start with browser-use, using a test site and fresh profile.
- MCP browser executor: use Playwright MCP, limiting origins and permissions at the client.
- Isolated code: use E2B or a container you control, with network off and a narrow workspace first.
- Stateful workspace agent: then read OpenAI Sandbox Agents; it remains Beta and its API can change.
Every path still needs action validation, approval, timeout / turn limits, result verification, and cleanup. A framework cannot decide your business risk automatically.
For chapter-length implementation, follow the canonical quickstarts below instead of duplicating an SDK textbook that quickly goes stale.
🛠 Hands-on exercises¶
Exercise 1 (Track A): open only one safe demo page¶
Copy this directly into your browser or computer agent:
Open only this page: <https://example.com>
Report the page title and final URL, and attach one screenshot.
Do not sign in, download, or leave example.com.
If the page asks for anything else, stop and tell me.
Check the title, URL, and screenshot yourself. If the agent leaves the allowlist, the exercise failed.
Budget: a local or included-subscription tool may add $0 in API cost; APIs and managed browsers follow provider pricing.
Exercise 2 (Track B): check first, then execute¶
Copy and run:
from urllib.parse import urlparse
ALLOWED_DOMAINS = {"example.com"}
ALLOWED_SCHEMES = {"https"}
LOW_IMPACT_ACTIONS = {"read", "screenshot"}
HIGH_IMPACT_ACTIONS = {"login", "purchase", "delete", "send"}
def check_action(url: str, action: str) -> str:
parsed = urlparse(url)
normalized_action = action.strip().casefold()
if (
parsed.scheme not in ALLOWED_SCHEMES
or parsed.hostname not in ALLOWED_DOMAINS
or parsed.username is not None
or parsed.password is not None
):
return "BLOCK"
if normalized_action in HIGH_IMPACT_ACTIONS:
return "ASK"
if normalized_action in LOW_IMPACT_ACTIONS:
return "ALLOW"
return "BLOCK"
assert check_action("https://example.com", "read") == "ALLOW"
assert check_action("https://example.com", " Login ") == "ASK"
assert check_action("https://example.com", "upload_credentials") == "BLOCK"
assert check_action("file://example.com/report", "read") == "BLOCK"
assert check_action("https://evil.example", "read") == "BLOCK"
print("policy checks passed")
This is not a complete sandbox. It teaches the outer policy. Only then should an ALLOW action reach the executor, with results and screenshots written to a log.
Budget: this local Python costs $0 and makes no API call.
Exercise 3: isolate code¶
Put a read-only CSV, output folder, and plotting script into a sandbox without host credentials. Disable unnecessary networking, then retrieve only the image and log. The result is evidence that output came from isolation, not merely a successful process.
Exercise 4: complete action loop¶
On a test site, connect observe → propose actions → policy check → approve / execute → verify. Send one URL outside the allowlist and prove that it is blocked. Do not use payment, real sign-in, email, or Slack as practice data.
⚠️ Safety cases: indirect prompt injection and protected accounts
Brave's research shows that malicious instructions can hide in content an agent reads. This is not a bug class limited to one browser; any agent that reads untrusted content and can act needs defenses.
Perplexity's BrowseSafe response explains its defense direction, but a provider classifier does not replace isolation, allowlists, approvals, and verification.
The Amazon case should not be reduced to “one browser was banned from Amazon.” The Ninth Circuit opinion dated 2026-08-04 discusses a district-court preliminary injunction concerning password-protected Amazon sections. The district court order gives the fuller scope. This is litigation context, not legal advice or a universal product-availability rule.
🎯 Featured Projects and Learning Resources¶
Choose only one to start:
- Desktop loop: Anthropic Computer Use tool.
- Web agent: Anthropic Browser Use tool or Playwright MCP.
- Isolated code: OpenAI Sandbox guide or E2B.
- Research: OSWorld 2.0.
- Attack surface: Brave indirect prompt injection research.
📚 21 complete learning resources and limits¶
Checked 2026-08-28 UTC. Stars are this project's teaching ratings, not GitHub stars.
| Group | Resource | Use it when | Limit / status | Rating |
|---|---|---|---|---|
| Official interface docs | Anthropic Computer Use tool | Understand the desktop action loop. | Client toolset; your application supplies the executor. | ⭐⭐⭐⭐⭐ |
| Anthropic Browser Use tool | Keep a task inside webpages. | Client toolset; requires a controlled browser. | ⭐⭐⭐⭐⭐ | |
| OpenAI Computer Use guide | Implement the GA computer tool. | The old preview shape is deprecated. | ⭐⭐⭐⭐⭐ | |
| OpenAI Agents SDK Sandbox guide | Need a stateful workspace. | Sandbox Agents are Beta. | ⭐⭐⭐⭐ | |
| Google Chrome Help: Gemini in Chrome | Check whether your account has access. | gradual rollout with platform and region limits. | ⭐⭐⭐ | |
| Executor / framework | anthropics/claude-quickstarts | Read the official computer-use demo. | Inspect container, credentials, and network boundaries first. | ⭐⭐⭐⭐⭐ |
| browser-use/browser-use | Build a full web-agent loop. | You still own production browser scaling and safety. | ⭐⭐⭐⭐⭐ | |
| microsoft/playwright-mcp | Connect a browser to an MCP client. | Restrict origins, permissions, and data. | ⭐⭐⭐⭐⭐ | |
| trycua/cua | Study a cross-platform computer-use stack. | Verify the actual backend from current README and releases. | ⭐⭐⭐⭐ | |
| bytedance/UI-TARS-desktop | Study an open desktop agent. | Local control is high risk; use a test environment. | ⭐⭐⭐⭐ | |
| Sandbox / runtime | e2b-dev/E2B | An agent needs a remote code workspace. | Apache-2.0 repo; managed service has separate cost and policy. | ⭐⭐⭐⭐⭐ |
| cloudflare/sandbox-sdk | Run isolated code on Workers and Containers. | Apache-2.0; Beta, and APIs may change before v1.0. | ⭐⭐⭐⭐ | |
| Modal Sandboxes | Need managed containers and runtime controls. | Configure network defaults and Beta / VM features from current docs. | ⭐⭐⭐⭐ | |
| Vercel Sandbox | Already build isolated execution in Vercel. | Check runtime, region, network, and pricing. | ⭐⭐⭐⭐ | |
| GUI / benchmark / dataset | microsoft/OmniParser | Study screenshot element parsing. | The repository is CC-BY-4.0; do not automatically apply that license to the weights. | ⭐⭐⭐⭐ |
| OSWorld 2.0 | Evaluate long-horizon desktop tasks. | Read scores with metric, step budget, and harness. | ⭐⭐⭐⭐⭐ | |
| xlang-ai/OSWorld | Reproduce the original cross-OS benchmark. | Its task set differs from 2.0; percentages are not directly comparable. | ⭐⭐⭐⭐⭐ | |
| web-arena-x/webarena | Evaluate self-hosted web tasks. | Environment setup and evaluator affect results. | ⭐⭐⭐⭐ | |
| OSU-NLP-Group/Mind2Web | Study demonstrations from real websites. | A dataset does not make current sites safe to automate. | ⭐⭐⭐⭐ | |
| Safety research and response | Brave: indirect prompt injection | Build a browser-agent threat model. | A research demo is not proof of every product's current state. | ⭐⭐⭐⭐ |
| Perplexity BrowseSafe | Compare a provider response and defense direction. | Read provider claims alongside independent testing. | ⭐⭐⭐ |
Read OmniParser weights by version: icon_detect_v3 uses the MIT-licensed YOLOv9 implementation; earlier Ultralytics detectors retain AGPL; caption models use MIT. None of these is synonymous with the repository's CC-BY-4.0 license.
💡 Future interfaces: Voice agents and VLA
A voice agent listens and speaks. VLA (Vision-Language-Action) lets a model see and control a physical machine. They are not the same layer as Browser / Computer / Sandbox, so this stage keeps only three entry points:
- LiveKit Agents: an open realtime and voice-agent framework.
- OpenAI Voice Agents guide: a current official voice-agent entry point.
- OpenVLA: a VLA research entry point.
The whole-site coherence layer will decide which specialist path owns them. This roadmap does not promise a nonexistent next stage.
✅ Self-check¶
- I choose the smallest interface first instead of sending every task to Computer Use.
- I can explain all eight terms and know Browser Use is not only DOM.
- I isolate, allowlist, require approval, and verify results and logs.
- I completed the example.com exercise without leaving the allowed scope.
- When reading OSWorld scores, I also find the task set, metric, step budget, and harness.
You have now completed the main path. Choose a specialist path: researcher, developer, teacher, knowledge worker, or everyday user.