Google DeepMind adds “computer use” to Gemini 3.5 Flash, another step toward truly operational agents
Google DeepMind has announced the introduction of a “computer use” capability in Gemini 3.5 Flash, an evolution that goes beyond the usual scope of conversational model demos. The new feature, presented by the lab in its official communication “Introducing computer use in Gemini 3.5 Flash”, targets a goal that is simple to state but difficult to industrialize: enabling the model to interact with graphical interfaces as a user would, in order to carry out tasks on existing software and services.
In other words, this is no longer just about answering a question, summarizing a document, or generating code. It is about seeing an interface, understanding its state, deciding on an action, clicking, entering text, navigating through multiple steps, and continuing until a goal is completed. For the AI ecosystem, this is a crucial building block: it brings Gemini closer to a category of tools often described as agents, meaning systems capable of acting in digital environments rather than limiting themselves to producing text.
Google DeepMind’s announcement comes at a moment of sharply intensifying competition. For several months, the battle between Google, OpenAI, and Anthropic has no longer been fought solely over general model performance or benchmark rankings. The central question is gradually becoming one of real-world execution: which player can turn a model into a credible software operator, reliable enough to automate concrete work tasks? It is precisely in this area that “computer use” takes on strategic importance.
For French-speaking readers, and especially for product teams, innovation leaders, integrators, and companies already engaged in process automation, the topic is far from theoretical. An AI capable of using a computer can potentially operate on tools that were not originally designed for AI, without waiting for dedicated APIs to become available. That opens immediate possibilities for workflows found in every organization: back office, operations, support, data entry, information retrieval, manipulation of business software, or coordination across web services.
The promise nevertheless remains bounded by a demanding technical reality. Using a graphical interface is far more complex than producing a good text response. Environments change, buttons move, windows overlap, forms vary, and interpretation errors have concrete effects. That is why Google DeepMind’s initiative should be read not as just one more feature, but as a sign of maturity in the race toward agents.
From chatbot to acting software: why this announcement marks a step change
Since the emergence of large generative models in the mainstream, the first phase of competition focused on relatively visible criteria: writing quality, speed, multimodal capabilities, context window size, reasoning ability, or coding skill. These dimensions remain important, but they are no longer enough to distinguish the most ambitious platforms. As companies look for use cases with measurable return on investment, value is shifting toward the ability to do, not just to say.
“Computer use” fits exactly into this transition. Historically, software automation has relied on several distinct approaches:
- API integrations, highly robust when they exist, but limited to services that expose well-documented programmable interfaces.
- RPA in the traditional sense, effective for repetitive and structured processes, but often fragile in the face of interface variations and costly to maintain.
- Conversational assistants, capable of helping the user but unable to act directly within software without an additional orchestration layer.
The promise of the multimodal model-based agent is to bring part of these worlds together: observe the screen, interpret the context, choose an appropriate action, and adjust in real time. In theory, this makes it possible to automate tasks that span multiple applications, including when those applications do not offer elegant integration. In practice, the difficulty is immense, because a model must maintain a stable understanding of a shifting visual environment, anticipate the consequences of its actions, and recover after errors.
That is where Google DeepMind’s announcement takes on its full meaning. By integrating this capability into Gemini 3.5 Flash, Google is not merely adding an experimental option. The company is positioning one of its models as a foundation for operational agency scenarios. The choice of Flash is not trivial: within the Gemini lineup, this family is generally associated with use cases where latency and cost matter greatly. Yet an agent acting on a computer must chain together many micro-decisions. If each step is too slow or too expensive, the experience becomes impractical.
This logic explains why the debate around agents can no longer be reduced to academic benchmarks alone. A model may achieve very strong scores on reasoning tests and still remain of limited use for controlling a real software environment. Conversely, a targeted improvement in interface perception, action accuracy, and loop stability can have a direct impact on enterprise use cases. “Computer use” then becomes a kind of decisive layer between the model’s theoretical power and its ability to generate value in the world of digital work.
Google DeepMind essentially frames it this way in its presentation: the new capability concerns interaction with graphical interfaces. That wording matters, because it anchors the announcement in the concrete. This is not an abstract agent, but a system taking on the sometimes chaotic reality of the modern digital desktop.
What Google DeepMind is actually announcing with Gemini 3.5 Flash
In its official communication, Google DeepMind presents a “computer use” capability for Gemini 3.5 Flash. The core of the announcement is clear: the model can interact with graphical user interfaces. This ability is meant to let it complete multi-step tasks across existing software and services, relying on visual understanding of the screen and on deciding which actions to execute.
The wording chosen by Google DeepMind is strategic. It positions Gemini not only as a high-performing multimodal model, but as a component that can be integrated into agentic systems. The implicit message is that next-generation AI should not be thought of only as a conversational interface, but as an action layer capable of plugging into the existing software estate.
The term “computer use” itself is revealing. It suggests an approach in which the model acts through the same kind of interface as a human, rather than through purely programmatic integration. In many professional contexts, that is potentially decisive. A large share of the tools used daily in companies are not always easily interoperable, or their APIs do not cover the full range of use cases. An AI capable of using the graphical interface can, in theory, bypass part of that friction.
Several levels of ambition must nevertheless be distinguished. Between a one-off demo and a reliable production agent, the gap remains considerable. Google DeepMind’s announcement does not mean that a model can now replace an employee on any computer task. Rather, it indicates that Google is adding an essential building block to its platform, one without which it is difficult to envision credible general-purpose agents.
This nuance matters in order to avoid two opposite but equally misleading readings. The first would be to downplay the announcement by reducing it to just another demo. The second would be to see it as general automation already ready at scale. The reality lies between the two: “computer use” is a structuring advance, but its value will depend on reliability, security, supervision, and integration into concrete workflows.
The fact that Google DeepMind is associating this capability with Gemini 3.5 Flash is also an indicator of how the company sees adoption. Agents operating on a computer often require repeated observation and action cycles. Every click, every completed field, every screen change becomes a step in reasoning and execution. In that context, the speed characteristics of a Flash model can make the difference between a usable system and an experience too slow to deploy.
The significance of the announcement is also clearer in light of the Google ecosystem. DeepMind is not working in a vacuum: Gemini is part of a broader platform that includes cloud tools, developer environments, and more broadly the exposure of models through Google products. Even if the announcement focuses on the capability itself, its meaning extends beyond the lab. It lays the groundwork for use cases in which AI no longer merely assists a user in a chat, but becomes a supervised operator within an enterprise software stack.
Google DeepMind’s communication emphasizes a capability to interact with graphical interfaces, in other words the model’s ability to act within software as it already exists, without reinventing the work environment.
For the market, this precision is essential. Companies are not only waiting for “smarter” models in the abstract sense; they are waiting for systems capable of fitting into the reality of their tools, screens, and procedures. That is precisely what “computer use” is meant to address.
Why the agent war is now being fought over execution, not just models
Competition between Google, OpenAI, and Anthropic is often presented as a race in model performance. That reading remains partly true, but it is becoming insufficient. For some time now, the competitive frontier has been shifting toward the ability to build usable agents, meaning systems that plan, execute, verify, and correct actions in real digital environments.
On this terrain, “computer use” is a central piece. An agent may have excellent internal reasoning; if it does not know how to manipulate the tools where data and processes reside, its value remains limited. Conversely, an interface-action capability, even an imperfect one, immediately opens a much broader field of experimentation: forms, administration consoles, business tools, CRM systems, internal portals, office suites, web services, legacy systems.
This shift in focus explains why announcements of this kind attract so much attention. They touch on the market’s most sensitive question: how to connect model intelligence to existing software infrastructure? For a long time, the answer mainly involved connectors, scripts, APIs, and orchestrators. Those building blocks remain fundamental, but the idea of a model capable of using a computer directly adds a complementary path, one that is sometimes more universal.
The decisive nature of this layer stems from several factors.
1. The graphical interface is the common denominator of digital work
Whether it is a recent SaaS service, an internal tool, or older software, most professional activities pass through screens, menus, buttons, tables, and forms. An AI that understands and manipulates these elements can theoretically operate in a very large number of contexts without depending on a specific integration for each application.
2. Useful tasks are often multi-step
In practice, few business processes boil down to a single command. You have to open a tool, check information, move to another service, copy data, go back, validate an operation, handle an exception. Language models are already capable of planning these sequences at an abstract level; “computer use” gives them a way to execute those plans.
3. Economic value depends on the full loop
A brilliant answer has limited value if the user still has to carry out all the manipulations themselves. Conversely, an agent that performs a significant part of the work chain can generate a much more tangible productivity gain. That is why companies are increasingly evaluating AI not only on linguistic quality, but on its ability to reduce the number of human actions required.
4. Differentiation is shifting toward reliability in real environments
Standardized benchmarks remain useful, but they imperfectly measure an agent’s robustness in the face of changing screens, unexpected windows, error messages, or workflow variations. Whoever masters this reality best will gain a decisive advantage, because they will offer the form of automation closest to everyday work.
In this context, Google DeepMind’s announcement should be read as both a defensive and offensive move. Defensive, because it had become difficult for a major AI player to remain outside this action layer. Offensive, because by bringing “computer use” to Gemini 3.5 Flash, Google is explicitly positioning itself in the battle for production agents, not just conversational assistants.
This evolution also has symbolic significance. During the first phase of generative AI, conversation served as a universal interface. In the phase now opening, conversation does not disappear, but it often becomes the command point of a system that acts elsewhere: in a browser, in a business application, in a console, in a digital workspace. The model is no longer valued only for what it writes, but for what it can trigger.
What this changes for companies, developers, and the French-speaking market
For French-speaking companies, the “computer use” capability addresses a very concrete need: automating tasks on tools already in place. In France as elsewhere in Europe, a large part of the economic fabric relies on heterogeneous environments, combining recent software, sector-specific platforms, internal applications, and older systems. In this landscape, API-only automation is not always enough. “Computer use” therefore appears as a particularly attractive avenue.
Several categories of players are directly concerned.
Large companies and mid-sized firms
They often have a substantial stack of tools, with high governance, security, and compliance constraints. For them, an AI capable of using a graphical interface may represent a way to automate certain processes without immediately overhauling the entire application architecture. But that promise comes with strong requirements: action traceability, human control, credential management, auditability, and limitation of operational risks.
Software vendors and integrators
For service companies, integration firms, and vendors of vertical solutions, “computer use” can become a new orchestration layer. It does not replace traditional integrations, but it can complement the available arsenal when APIs are missing, incomplete, or too costly to use. That opens important room for differentiation, particularly around supervision, guardrails, and the design of hybrid workflows.
Product teams and developers
For technical profiles, Google DeepMind’s announcement raises a very practical question: how should applications and experiences be designed when the model does not merely generate content, but acts within a software environment? That implies thinking differently about interfaces, validations, intermediate states, error recovery, and confirmation mechanisms. The design of an agent operating on a computer is not that of a simple chatbot.
SMEs
For smaller organizations, the promise is potentially even more appealing: access to advanced automation without having a heavy engineering team. But this is also where the risks of misuse are highest if solutions are perceived as ready-to-use magic. An agent that clicks in the wrong place, sends the wrong document, or validates the wrong operation can cause costly errors. The democratization of these tools will therefore have to involve strong education.
In the European, and therefore French-speaking, context, several points deserve particular attention.
- Sovereignty and hosting will remain major criteria for many organizations, especially in regulated sectors.
- Compliance is central as soon as an agent interacts with personal data, HR tools, financial systems, or healthcare environments.
- Language matters less at the click level itself than at the level of interfaces, instructions, documents, and local workflows; good understanding of professional French therefore remains important.
- Software debt in European organizations may make “computer use” particularly relevant, because many environments are not natively designed for AI agents.
It should also be noted that this capability could change how companies evaluate their AI projects. Until now, many pilots have focused on documentary or conversational use cases: internal search, FAQs, writing assistance, summarization. With “computer use,” the bar shifts toward scenarios closer to operations: case processing, system updates, portal navigation, execution of repetitive procedures. These are more sensitive use cases, but also often closer to direct economic impact.
For the French-speaking market, the issue is therefore not only technological. It is also industrial. Companies capable of building trust layers around these agents — supervision, logging, validation, security, adaptation to local software — could capture a significant share of the value. Conversely, those that remain focused solely on text assistants risk finding themselves out of step with the next wave of automation.
Limits, open questions, and the long-term outlook
As strategic as it may be, the introduction of “computer use” into Gemini 3.5 Flash does not suddenly solve the main challenges of agents. On the contrary, it makes them more visible. As soon as a model acts on a computer, the requirements change in nature. An error is no longer limited to an approximate answer; it can become an incorrect action, wrongly entered data, an operation launched at the wrong time, or navigation in the wrong context.
The first question is that of reliability. Graphical interfaces are notoriously variable. A given service may change its design, display a pop-up window, request an additional verification, or reorganize its menus. A robust agent must not only recognize these changes, but also decide when it can continue and when it must request human validation. That is a far more demanding maturity threshold than that of a conversational assistant.
The second question is that of security. An agent using a computer potentially handles accounts, permissions, sensitive data, and actions with real effects. That requires strict guardrails: limiting the scope of action, authorization policies, environment segmentation, continuous monitoring, the ability to stop immediately, and a detailed action log. Without these mechanisms, the promise of automation can quickly become an operational risk.
The third question is that of responsibility. In an enterprise workflow, who is accountable for an action carried out by an agent? How do you distinguish between a model error, a bad instruction, an interface change, or insufficient configuration? These issues are particularly sensitive in Europe, where compliance, traceability, and governance concerns are structuring.
The fourth question is that of the real-world economics of deployment. Even if the idea of an agent capable of using any software is appealing, not all processes lend themselves to the same level of automation. In some cases, API integration will remain more reliable and easier to maintain. In others, “computer use” will offer valuable flexibility. The market will therefore probably converge toward hybrid architectures, where agents combine native tools, connectors, and interface manipulation depending on the need.
Despite these reservations, the overall direction seems clear. The center of gravity of applied AI is gradually shifting toward systems capable of acting within real-world software. Google DeepMind’s announcement confirms that this movement is no longer marginal. When a player of this size introduces “computer use” into Gemini 3.5 Flash, it sends a message to the entire ecosystem: the future of models will not be decided only by the quality of their answers, but by their ability to become supervised digital operators.
In the long term, this evolution could have several profound consequences. First, it could redefine the very notion of the user interface. If a growing share of tasks is carried out by agents, software may need to be designed not only for humans, but also for AIs that read the screen, interpret states, and trigger actions. Second, it could redistribute value among model providers, cloud platforms, software vendors, and integrators capable of turning these building blocks into reliable systems.
For Google, the stakes go far beyond an isolated feature. The company is seeking to demonstrate that its Gemini family can serve as a foundation for advanced professional use cases, at a time when competition around agents is becoming one of the industry’s most closely watched battlegrounds. For OpenAI, Anthropic, and the other major players, the pressure is increasing symmetrically: it is no longer enough to have a good model, it is necessary to show how that model works.
For French and European companies, the right reading is neither naive enthusiasm nor automatic skepticism. This announcement should be seen as a phase change. “Computer use” does not replace existing architectures, but it radically expands the scope of possible automation. Organizations that know how to test these capabilities methodically, on well-chosen perimeters, could gain a significant lead in the industrialization of agents.
The next AI battle will therefore probably not primarily concern the model that answers an abstract question best, but the one that knows how to open a tool, understand a situation, execute a sequence of actions, and stop at the right moment. With the introduction of “computer use” into Gemini 3.5 Flash, Google DeepMind is formally recognizing precisely this shift: the agent war is won less on promises of reasoning than on the ability to turn intelligence into digital work that is actually carried out.
Comments· 2 comments
"Learns to use a computer" sounds catchy, but what exactly is being measured here? Are we talking about reliable GUI grounding and multi-step task completion, or just demo-level point-and-click on a narrow benchmark?
That’s my question too. I’d want to see the evaluation setup—things like what apps or interfaces were used, how success was defined, and whether it handled mistakes or unexpected pop-ups—before reading too much into the headline.