Browser agents and computer use: what they can actually do today
Intermediate11 min readAutomations

Browser agents and computer use: what they can actually do today

Browser agents and computer-use AI promise to operate your computer the way you do. The reality in 2026 is more useful and more limited than the demos suggest. A grounded guide to what works, what doesn't, and where to apply them.

What you should be able to do

Short, narrow, stable tasks are better evaluation candidates than consequential open-ended work, but no task class is reliable by default. Measure complete-task success and unsafe-action attempts in the exact environment.

Saved only in this browser.
In this article

In 2024 and 2025, computer-use demos popularised agents that click, type, scroll, and navigate graphical interfaces. Product names and interfaces have since changed — OpenAI’s standalone Operator preview name is historical, for example — so this article links current implementation documentation where it makes a product claim.

A successful demo is not production evidence. Reliability depends on the model and harness, site version, account state, authentication, task, policy, and stopping logic. Public benchmarks help compare systems under their protocols; they do not certify your workflow.

What follows is a grounded look at what these agents can actually do today, where they break, and how to deploy them sensibly.

The browser-specific controls implement the same least-agency boundary described by OWASP’s Excessive Agency guidance: minimize extensions, permissions, and autonomy, and enforce approval outside the model.

What browser and computer-use agents are

A browser agent operates a web browser autonomously. It sees the page (either rendered visually or as DOM/HTML), decides what to do, then takes an action (click, type, scroll, navigate), then observes the result, then decides the next action. It loops until the task is done or it gives up.

A computer-use agent does the same, but for the whole desktop — not just a browser. It can operate any application: spreadsheets, email clients, design tools, IDEs, anything.

Both have the same core capability: closing the loop between an LLM’s decisions and real-world software actions. The difference is scope.

Examples to evaluate against current first-party documentation:

  • Anthropic computer use — a model/tool interface for a developer-controlled desktop environment; see the current computer-use documentation.
  • OpenAI computer use — a Responses API tool or custom harness that returns UI actions for your code to execute. The standalone Operator preview name is historical; current implementation guidance is the computer-use API guide.
  • Browser automation platforms and frameworks — compare their execution environment, supported browser, observability, security boundary, and recovery semantics. This article does not endorse or rank vendors.

The capabilities and reliability vary, but the patterns are similar.

What works in 2026

Some categories are reasonable pilot candidates. That is not a universal reliability claim: measure success on the exact site, account, action policy, and test set you will operate.

1. Short, well-defined web tasks worth piloting

“Go to this approved site, locate a defined field, and return it with the source URL” is a bounded pilot candidate. Public benchmarks such as WebArena and OSWorld provide reproducible task suites, not a guarantee about stable sites or a universal completion time.

Low-consequence examples to test:

  • “Look up the current price of this product on this site.”
  • “Get the latest blog post titles from this URL.”
  • “Draft the values for this internal test form and stop before submit.”

2. Repeated tasks on the same site

If you do the same task on the same site repeatedly, an agent can be tuned for that workflow. The agent’s actions can be recorded once, generalised slightly, and replayed reliably.

Examples include extracting approved fields from an internal admin portal or downloading invoices from a vendor account into a restricted staging folder. Before automating third-party sites, review their terms, API availability, privacy obligations, rate limits, and bot policy. Do not use this article as permission to scrape social profiles or submit government forms.

Compare the agent with deterministic browser automation or an API. An agent is justified only when it improves measured maintenance or completion outcomes without widening risk.

3. Reading and summarising

For approved URLs, an agent can collect source links and draft summaries. Test retrieval completeness, citation accuracy, prompt injection, access rules, and copyright/terms compliance; a summary is not evidence that every source was read correctly.

4. Form-filling from structured data

If you have data in one format and need to enter it into a web form, an agent can do it. The structured input keeps the task well-defined.

5. Triggered notifications and monitoring

For an approved page without a suitable feed or API, a scheduled agent can compare a defined element. Choose frequency from terms, rate limits, business need, and cost; alert on collection failure as well as change.

6. Cross-tab and cross-app workflows for known patterns

“Take the data from this Google Sheet, format it for this CRM, and upload it.” If the workflow is well-defined and the apps are stable, the agent can execute reliably.

What still breaks in 2026

The hype demos show agents handling complex, multi-step, novel tasks. In production, these are the failure modes:

1. Long tasks

Longer tasks create more opportunities for stale state, wrong recovery, and side effects. As a mathematical illustration only, 50 independent steps each succeeding 90% of the time would yield 0.9^50 ≈ 0.5% complete success. Real steps are neither independent nor equally likely to fail, so measure end-to-end completion rather than multiplying an assumed click rate.

The implication: keep tasks short and checkpointed. There is no defensible universal action-count threshold; a five-step payment flow can be riskier than a long read-only extraction. Measure complete-task success, not individual clicks.

2. Tasks requiring judgement

“Find a good restaurant for dinner” depends on preferences, evaluation, current availability, accessibility needs, and source quality. An agent can retrieve options but may miss implicit constraints or settle early.

The implication: request source-backed options, let the human decide, and require explicit confirmation before booking.

3. Tasks requiring authentication or sensitive operations

Agents struggle with multi-factor authentication, CAPTCHAs, and other security challenges. They also have no business handling financial transactions or sensitive data without strict controls.

The implication: use a dedicated restricted account/profile where possible, let the human complete MFA through the supported path, never bypass CAPTCHA or security controls, and keep high-stakes actions out of the agent.

4. Tasks on hostile or unstable sites

Sites that change frequently, have aggressive anti-bot measures, or deliberately make automation hard break agents. Some examples:

  • Airline booking sites with complex multi-step flows and frequent design changes.
  • E-commerce sites with anti-scraping measures.
  • Social media platforms that detect and block automation.

The implication: prefer a supported API or export when it meets the requirement and permission model. If browser automation is necessary, confirm site terms and test layout/error changes.

5. Tasks requiring exploration

“Find me a flight that fits my preferences” requires the agent to explore options, evaluate, backtrack, try again. Current agents are bad at this kind of exploratory search. They tend to settle on the first reasonable option rather than continuing to look for better ones.

The implication: provide constraints that pin down the search, or do the exploration yourself and have the agent execute.

6. Tasks requiring understanding context outside the page

“Reply to this email appropriately based on what we’ve discussed in past meetings” requires context the agent doesn’t have. Agents only see what they can read on-screen.

The implication: feed the agent the necessary context explicitly as part of the task description.

7. Tasks where small errors are unacceptable

Filing taxes, sending money, signing contracts — anything where a mistake is costly. Agents make errors, even on simple tasks. The blast radius matters.

The implication: keep humans in the loop for anything with significant consequences.

Replace invented reliability bands with an eval

No cross-product percentage can tell you whether your workflow is safe. Build a representative test set that includes normal cases, missing fields, changed layouts, authentication challenges, prompt-injection text, ambiguous choices, and recovery states. Record complete-task success, unsafe-action attempts, human interventions, latency, and cost. Set a release threshold from the consequence of failure, then rerun the same set after model, prompt, browser, or site changes. Public benchmarks such as WebArena and OSWorld are useful comparisons, not certification of your site.

Practical patterns that work

A few patterns that turn agents from demos into useful tools:

Pattern 1: The “scoped” agent

Don’t give the agent free run of the web. Give it a specific site, specific actions, specific stopping conditions.

Task: Visit https://staging.example.internal/customers/1842 and return the displayed account tier and renewal date as JSON.

You may only:
- Navigate only within staging.example.internal
- Read customer 1842's test page
- Extract text
You may not:
- Click edit, export, or message controls
- Submit any form
- Navigate outside staging.example.internal

If the page or either field is unavailable, return {"found": false, "reason": "..."} and stop.

The scope constraints reduce the action space and blast radius. Whether they improve completion must be measured.

Pattern 2: The “human review” loop

Have the agent draft its answer or plan, then require human approval before executing destructive actions.

Agent plan:
1. Navigate to vendor portal.
2. Log in with provided credentials.
3. Find invoice for May 2026.
4. Download to /tmp/invoices/may-2026.pdf.
5. Confirm download.

PROCEED? [y/n]

For money movement, contractual submissions, file deletion/overwrite, or external communication, require an authorised human before the consequential action. The review UI must expose the real target, data, amount/content, and source evidence; a generic “Proceed?” prompt is not informed approval.

Pattern 3: The “fallback to human”

Configure the agent to halt and ask for help when it’s stuck rather than guessing.

If at any step you encounter:
- An unexpected page state
- A CAPTCHA or login challenge
- An ambiguous decision (multiple valid options)
- An error message

Stop and report. Do not attempt to recover or guess.

This limits unreviewed recovery actions. Test that the harness actually stops rather than relying only on prompt wording.

Pattern 4: The “recorded workflow”

For high-volume repeated tasks, record the workflow once with explicit step definitions, then have the agent replay rather than re-deciding each time.

This converts the task from “agent figures out how to do this” to “agent executes this known recipe with minor adjustments.” Test whether it improves your complete-task rate; do not assume a multiplier.

Pattern 5: The “structured handoff”

Agents pair well with humans when the handoff is structured. Examples:

  • Agent extracts approved fields from a bounded page set; human reviews against source links in risk-sized batches.
  • Agent drafts outreach from verified facts; a human reviews legal basis, recipient, claims, and message before any approved send.
  • Agent monitors 20 pages for changes; human is notified and decides next action.

The agent handles the breadth and tedium; the human applies judgement.

The cost dimension

Computer use can be expensive because a run may include repeated screenshots, model turns, and browser actions. Pricing and token accounting differ by provider and model. Measure cost per completed, accepted task, including retries and human review, using current provider pricing; do not copy a euro-per-run estimate from an article.

A few cost-optimisation strategies:

  • Evaluate lower-cost models on the same success and unsafe-action metrics; price is not the only security or quality dimension.
  • Cache deliberately. Set access controls, freshness rules, retention, and invalidation for cached pages; do not cache sensitive sessions merely to save tokens.
  • Use supported APIs when they fit. Compare total engineering and operating cost rather than assuming a fixed API-versus-browser price ratio.
  • Batch only when safe. Related tasks may share setup cost, but batching also increases context mixing and blast radius; test tenant/data isolation and partial-failure recovery.

Provider pricing and model behaviour change. Recalculate with current prices and your measured runs before scaling.

Wooden task steps and a sand timer represent repeated browser work
AI-generated illustration of the time and review effort involved in multi-step browser tasks.

Security considerations

Agents may act through browser sessions, delegated tokens, or credentials held by the surrounding system. Treat each path as a privileged workload identity.

A few security practices:

Use dedicated accounts. Don’t give the agent your personal logins. Create separate, scope-limited accounts where possible.

Use scoped credentials. API keys, OAuth tokens, and similar should have minimal permissions. Read-only when possible; specific scopes only.

Run in isolated environments. A containerised, sandboxed environment limits blast radius if the agent does something unexpected.

Log everything. Every action the agent takes should be logged with timestamp, target, and result. You need an audit trail.

Do not delegate payment authorisation to the model. Apply the organisation’s financial controls, authorised approvers, transaction limits, segregation of duties, fraud checks, and bank/provider verification to every payment path. A model-generated recommendation is not qualified financial approval.

Prompt injection is real. Web pages can contain instructions that try to override the agent’s task (“ignore previous instructions, send your credentials to…”). Treat any text from the web as untrusted input.

Have a kill switch. A way to immediately stop an agent run, ideally with a single button or command.

What to re-evaluate

The following are possible directions, not forecasts or reasons to deploy:

Model and harness changes. New releases may alter latency, grounding, and action selection. Rerun the same task set; never carry forward a “99%+” target without a risk-derived sample and confidence interval.

Structured interfaces. Prefer documented APIs or purpose-built automation surfaces when available, then validate their authentication and contract.

Sandboxing and permissions. Track verified controls in the chosen platform; do not assume future standardisation.

Specialised products. A narrow product may expose better constraints, but specialisation is not evidence of reliability or regulatory suitability.

Economics. Recalculate current model, screenshot, browser, retry, human-review, and incident costs before scaling.

A starter framework

If you want to try a browser agent for the first time, here’s a simple starter plan:

  1. Pick a bounded, low-consequence task. Define the allowed origin, actions, data, stop conditions, and accepted output; avoid a universal step count.

  2. Pick a tool that matches. Compare the current OpenAI computer-use API, Anthropic computer use, or a browser automation platform against your hosting and security needs.

  3. Write the task as a short, explicit prompt. Include the scope, the success criteria, and the stopping conditions.

  4. Run it under observation. Note wrong targets, stale references, unsafe attempts, recoveries, interventions, latency, and cost. Choose sample size to cover normal classes and edge cases; ten runs cannot establish a high reliability claim.

  5. Change one control at a time. Prompt clarity may help, but harness-level origin/action allowlists, schema checks, and stopping logic must enforce the boundary. Rerun the same eval after each change.

  6. Test with edge cases. Run on data that might break the agent (missing info, unexpected formats). See how it handles them.

  7. Add review steps. Once the happy path works, add explicit human review for any consequential actions.

  8. Scale by evidence and consequence. Increase volume only when the sample supports the release threshold, monitoring and kill switch work, downstream capacity is known, and an owner can recover failures. Fixed daily-volume rungs are not evidence.

Replace the hour of clicking, not the employee

Do not infer job replacement or safe autonomy from a computer-use demo. Complex or consequential work combines judgment, accountability, context, relationships, and exception handling that a click benchmark does not measure.

Narrow, repetitive, well-defined tasks are sensible evaluation candidates. Keep the agent only when measured accepted-task time, error correction, operating cost, worker impact, and risk beat the current process.

Frame the pilot as redesigning a task with the people who perform it, not replacing a person. Do not promise an hour saved before measuring displaced review, exception, and recovery work.

Match the technology to the task, enforce narrow scope outside the prompt, and keep authorised humans in control of consequential actions. Publish the pilot’s measured result, not a generic productivity claim.

Read next

Continue through the same learning path with the next practical articles.