This website uses cookies

Read our Privacy policy and Terms of use for more information.

The useful distinction between an AI assistant and an AI agent is becoming harder to maintain. An assistant writes, summarises or analyses. An agent can move through a web service, retrieve information, complete a sequence of steps and return with an outcome. That direction was visible twice in the same week.

On 25 August, Anthropic announced that Claude memory would work across Chat and Claude Cowork. The next day, it introduced a browser inside the Cowork desktop app. OpenAI’s release notes also confirmed that ChatGPT Work can complete certain tasks on websites that require users to sign in. The products differ in design and availability, but both bring delegated browser activity closer to ordinary work.

Thibault Sottiaux, who leads OpenAI’s core product work, told TechCrunch: “ChatGPT can actually do entire, very complicated tasks for you all autonomously in a way that is delightful and safe.” That is a statement of product ambition, not a guarantee that every website, task or workflow is supported. It does, however, show why access controls now deserve as much attention as the quality of an AI answer.

Claude is joining memory and browser activity

Anthropic says Claude now uses the same memory across Chat and Cowork. Users can view their saved memory topics, edit them and delete them. On Free, Pro and Max plans, memory is on by default. On Team and Enterprise plans, an administrator controls whether memory is available, and individual users must opt in.

The policy detail matters. Anthropic says sensitive-topic memory is off by default. It also says some categories of information, including sensitive identification numbers, criminal history and immigration status, will not be stored in memory even where sensitive-topic memory is enabled. Persistent context can make a tool more useful, but it also introduces questions about what is retained, where it is available and who can change it.

The Cowork browser is separate from the user’s own browser. Anthropic says it cannot see a person’s tabs, bookmarks or passwords. A user can bring logins into the Cowork browser on a site-by-site basis. Banking, email and single-sign-on sites are excluded unless the user chooses to include them.

The feature was rolling out to Pro, Max and Team plans in the desktop app. Enterprise availability was subject to administrator controls. Anthropic also flags a more difficult issue: browser agents are exposed to prompt injection. Its published defences can reduce that risk, but the company does not say they remove it entirely.

That admission is important. Prompt injection is not simply a nuisance prompt in a chat. It can involve instructions embedded in a page, document or tool output that an agent treats as part of the task. The more an agent can read, remember and act on, the more the design of its boundaries matters.

OpenAI’s approach is built around a separate cloud session

ChatGPT Work can now support some signed-in web tasks on the web and mobile. OpenAI says that, where a site supports authentication, the product presents a secure sign-in form for the user to enter credentials. Password managers can be used in the flow. According to OpenAI, the username and password entered into the form are not visible to the model, not stored by ChatGPT and not used for model training.

The signed-in session can remain available for later tasks until it expires or the user clears the relevant browser data. OpenAI says the browser operates on a separate cloud computer. It does not use the tabs, history, cookies or saved passwords in the browser on a user’s device. The product also lets users control which websites Work may access, and OpenAI says consequential actions such as reservations or payments need confirmation.

Andrew Ambrosino, lead engineer for OpenAI’s desktop app, captured the difficulty of this product category when he told TechCrunch: “One of the unique challenges with a product like this is just that it can really do anything.”

This observation explains why a one-size-fits-all permission setting is unlikely to be enough. A browser task could involve reading a public page, entering a private account, changing a record, or initiating an action with financial or legal consequences.

The browser is not the whole security boundary

Separate browser sessions and secure sign-in flows are useful product choices. They are not the whole security model. The UK’s National Cyber Security Centre has published interim guidance for organisations deploying agentic AI. It says the greater an agent’s autonomy, the greater the potential impact if the agent behaves unexpectedly, reaches information it should not see, or acts outside its intended scope.

The guidance advises organisations to start by defining what is inside and outside a task’s scope, including clear red lines. It recommends threat modelling before deployment, and says prompt instructions should be combined with technical and operational controls. It distinguishes between a human approving actions before they occur, a human monitoring and able to intervene, and an agent acting without human review. The NCSC recommends human oversight alongside technically enforced controls in higher-risk situations.

That logic applies directly to browser-based agents. A task that checks a website for updates requires a different permission level from one that can submit forms, amend account details, release documents or make purchases. The safest useful model is not necessarily the most restrictive one. It is one in which the agent is given only the authority that the task needs.

Toby W, Principal Security Architect at the NCSC, puts the point succinctly: “Do not rely on prompting alone. You should combine prompts with technical and operational controls to provide defence in depth.”

Vendor guardrails are part of that defence. They are not a substitute for deciding which sites, data, credentials and actions an agent should be allowed to reach.

A practical way to think about delegated access

The announcements from OpenAI and Anthropic make several design choices visible. They also offer a useful vocabulary for anyone assessing browser-capable AI tools.

Identity: Does the agent have an identity that can be distinguished from a person’s usual account activity? The NCSC advises assigning agents a unique identity class where possible.

Permissions: Are tools and logins scoped to the task, rather than handed over in a general session? The NCSC recommends limiting permissions and using short-lived credentials where possible.

Environment: Is the agent in an isolated browser or wider sandbox? Isolation reduces the range of data and systems the agent can reach, but its adequacy depends on the use case.

Approvals: Are there clear points at which a user needs to confirm an irreversible or high-consequence action? A confirmation is most useful when it comes before the action, not after it.

Observability: Can a team tell what the agent did, where it went, and when it stopped? NCSC guidance recommends logging, monitoring, and an ability to halt activity quickly.

NIST’s AI Agent Standards Initiative reflects the same shift in emphasis. It identifies agent authentication and identity infrastructure as a research priority for secure interactions between people and agents, as well as between agents. The issue is no longer just whether an AI can use a website. It is how a website, an organisation, and a user can determine what authority that AI has been granted.

The Frontier Model Forum’s guidance adds another useful layer. It describes agent security as shared across model developers, deployers, tool providers, third-party services, and users. This is not a claim that responsibility can be passed along the chain. It means different actors control different risks. A provider can build a browser isolation feature. A business can decide which data and systems are connected. A user can decide whether a task is appropriate for delegation.

Takeaway

The key development is not that agents can browse. Browsing was already possible in several tools. The change is that major AI products are building memory, sign-ins, and browser action into ordinary user workflows.

For teams assessing those tools, three questions should come before enthusiasm about automation:

What exact outcome may the agent produce? Define the task, the permitted action, and the red lines before access is granted.

What does the agent need to reach? Limit websites, data and credentials to the minimum needed for that task. A broad session creates a broader potential impact.

Who can see and stop what happens? Retain meaningful approval points, operational logs, and a way to cut off access if activity becomes unexpected.

The most useful question is not whether to let an AI agent behind a login screen. It is whether the authority, identity, and oversight behind that access match the consequence of the task.

Browser-capable AI is moving from demonstration to everyday workflow. Subscribe to the Project Flux newsletter for the changes worth understanding before they become routine.

Links and Stuff

All content reflects our personal views and is not intended as professional advice or to represent any organisation.

1  

Reply

Avatar

or to participate