From the conversational agent to the system that acts in physical space

After assistants capable of generating text, code, images, or answers to complex questions, the next frontier of artificial intelligence is playing out in the physical world. Google DeepMind explicitly places Gemini Robotics ER 2 on this path. In a publication titled “Gemini Robotics ER 2: powering robotics with video understanding, task orchestration, and multi-robot collaboration”, the company presents a system geared toward three capabilities: understanding video streams, organizing physical tasks, and enabling multiple robots to work together.

The change in terminology is significant. It is no longer just about giving an instruction to a single machine, nor demonstrating that a robotic arm can grasp an object in a controlled setting. Google DeepMind highlights the possibility of analyzing what robots perceive, breaking down a mission, and distributing its stages among several pieces of equipment. The stated objective therefore brings multimodal models closer to the realities of industry, logistics, and, in the longer term, humanoid robotics.

Robotics has historically faced a particular challenge: the real world is less predictable than a software interface. A robot must deal with moved objects, unforeseen obstacles, changes in lighting, deformed packaging, moving people, interrupted operations, and imperfect sensors. In this context, AI cannot be limited to producing a convincing response in natural language. It must connect perception to action, understand the consequences of that action and, when several machines are involved, maintain collective consistency.

Google DeepMind’s publication does not present Gemini Robotics ER 2 as a simple computer vision tool. The original source’s title combines video understanding with task orchestration and multi-robot collaboration. These three dimensions outline the architecture of a physical agent: seeing a situation, identifying what needs to be done, then allocating or coordinating execution.

This ambition extends Google’s work in robotics based on large models. In 2022, Google Robotics presented RT-1, a model designed to translate instructions into robotic actions. In 2023, RT-2 marked an important step by seeking to transfer knowledge from vision-language models to robotics. Google DeepMind then announced Gemini Robotics in March 2025, with the idea of adapting Gemini’s multimodal capabilities to interaction with the physical world. Gemini Robotics ER 2 fits into this continuity, while shifting the center of gravity toward video observation, mission organization, and the coordination of a plurality of machines.

The model’s very name suggests an extension of Gemini Robotics-ER, where “ER” refers to embodied reasoning. This concept encompasses a central idea in robotics research: reasoning usefully is not only about describing an environment, but also about understanding how a body, sensors, and effectors can act within it. An AI may know how to recognize a cup in an image; a robot must still determine whether it is accessible, how to pick it up without breaking it, where to place it, and how to react if someone moves it during the operation.

What Google DeepMind is announcing with Gemini Robotics ER 2

According to Google DeepMind, Gemini Robotics ER 2 extends robotic reasoning capabilities using video streams. This detail is essential. A still image provides a state; a video also makes it possible to observe change. The system can potentially distinguish a stationary object from an object being moved, spot a work sequence, track an action that has begun, or place an incident within the progression of a task.

In an operational environment, this temporal dimension is decisive. A warehouse, picking line, workshop, or maintenance area cannot be reduced to a snapshot of space. Bins move, doors open, pallets change position, and operators move from one station to another. A system that understands video streams can, in theory, provide a richer representation of the scene and its evolution than a mechanism based on isolated states.

Google DeepMind links this video understanding to the orchestration of complex physical tasks. Here again, the wording matters. A complex task can involve several sub-goals, dependencies, sequencing constraints, and checks. It is necessary, for example, to identify that an area must be prepared before an object is moved there, that a piece of equipment is unavailable, or that an operation must wait until another is completed. Orchestration therefore does not refer merely to sending orders: it involves structuring action over time and among different participants.

The third dimension highlighted by the source is collaboration between multiple robots. DeepMind thus goes beyond the traditional framework of demonstrations focused on a single machine. In many industrial environments, the economic value of robotics lies not in the spectacular autonomy of an isolated device, but in the ability to integrate multiple pieces of equipment: robotic arms, mobile robots, conveyor systems, cameras, inspection tools, and supervisory software.

Google DeepMind’s text places Gemini Robotics ER 2 at the intersection of these needs. The model aims to help robotic systems reason about what they see, coordinate actions, and collaborate. This announcement should not be turned into a promise of universal deployment: the publication mentioned in the brief does not detail general availability, a commercial timetable, pricing, a list of industrial partners, or comparable performance metrics. These elements will have to be examined separately if the company provides them.

This caution is particularly important in robotics. A laboratory demonstration, however impressive, does not guarantee a system’s robustness over thousands of hours, across varied buildings, and under high safety requirements. Video understanding can be very useful, but it must also withstand occlusion, changing lighting, camera angles, unknown objects, and detection errors. Likewise, orchestrating multiple robots requires managing trajectory conflicts, waiting times, operational priorities, and shutdown situations.

The wording chosen by Google DeepMind nevertheless reveals a clear ambition: to move from an AI that assists a person in front of a screen to an AI capable of participating in the organization of physical actions. The model is not limited to producing a textual plan. It is presented as a component capable of relying on video to participate in the coordinated operation of robots in real environments.

Google DeepMind’s original publication presents Gemini Robotics ER 2 as a system intended to “power robotics with video understanding, task orchestration, and multi-robot collaboration.”

This sentence should not be read as an exhaustive technical definition, but it establishes the scope of the announcement. Google DeepMind emphasizes video perception, practical planning, and cooperation, rather than a single motor capability or a single hardware platform.

Why video and fleet coordination change the scale of the problem

Computer vision has already been used for a long time in factories, warehouses, and sorting centers. Specialized systems check for the presence of a part, read a code, detect a defect, or guide a robot along a defined path. The claimed difference of approaches based on multimodal models lies in their ability to handle more general situations and connect perception to instructions expressed in human language.

Gemini Robotics ER 2 is part of this evolution. Understanding a video does not merely involve recognizing categories of objects. In a robotics context, the challenge is to determine what is happening, what has changed, what remains to be done, and which participant is best positioned to act. Such a capability can become useful when a process does not exactly match a preprogrammed sequence, or when human supervision must quickly understand why an operation is not proceeding as expected.

The term orchestration also refers to an operational reality that is less visible than grasping demonstrations: most productivity gains come from sequencing. In a fleet, a robot may remain unused because an aisle is blocked, because a bin is not ready, because a station needs to be cleaned, or because an upstream machine has fallen behind. Coordinating work means reducing these frictions, allocating tasks, and responding to events without requiring manual reprogramming for every exception.

The shift from the individual machine to the fleet profoundly changes software requirements. An autonomous robot must manage its own perception and movements. A fleet must additionally have a shared representation of the situation, a task-allocation mechanism, priority rules, and supervision capable of taking back control. The more agents there are, the more critical communication becomes: a locally reasonable decision can become inefficient, or even dangerous, if it ignores the actions of other machines.

In this context, large models can play a role as an interface and reasoning layer, but they do not automatically replace industrial control systems. Robots need deterministic mechanisms, movement limits, safety sensors, and shutdown procedures. The AI model can help understand a mission, propose a breakdown, or interpret a scene; execution must remain governed by rules suited to the hardware and site in question.

This distinction is fundamental to understanding DeepMind’s announcement. Generative AI has popularized the idea that a system can respond to a very open-ended request. The physical world imposes an additional constraint: an erroneous response can result in inappropriate movement, a production stoppage, equipment damage, or a risk to people present. The issues of reliability, validation, and human takeover are therefore greater than in most conversational uses.

Multi-robot cooperation is of particular interest to logistics, where value often depends on overall throughput. A robotic arm can pick up objects, a mobile robot can transport them, and another system can prepare the destination area. None of these elements alone delivers complete automation. Their coordination can, however, make it possible to handle a sequence of operations. It is precisely this kind of shift from an isolated capability to a complete process that the orientation of Gemini Robotics ER 2 suggests.

The same reasoning applies to industry. In a factory, an intervention may require inspecting equipment, transporting a tool, handling a part, and conducting a final check. Video understanding could help interpret the state of a scene; orchestration could organize the steps; multi-robot collaboration could distribute complementary roles. But the real value will always depend on the level of integration with existing equipment, site constraints, and the quality of safety procedures.

Finally, the subject directly concerns the future of humanoid robots, often presented as capable of using environments designed for humans. Even if Google DeepMind does not reduce Gemini Robotics ER 2 to humanoids, the announced capabilities are relevant to this category: a humanoid does not merely need to move or manipulate, it must understand its context, accept instructions, and work with other devices. The announcement therefore fits into a broader race to build more versatile physical agents.

An already fiercely contested technological race

Google DeepMind is not alone in seeking to bring foundation models closer to robotics. In recent years, AI players, chipmakers, robotics companies, and specialized startups have multiplied announcements around vision-language-action models. Their shared promise is to reduce the distance between a human instruction, the perception of a scene, and the execution of a gesture.

Google itself had helped structure this category with RT-1 and then RT-2. RT-2 notably drew attention by showing how knowledge acquired by models trained on the web could be used for robotic tasks. Gemini Robotics, announced in 2025, then linked Gemini’s multimodal capabilities to robots’ needs. Gemini Robotics ER 2 extends this line, but introduces in its public positioning a strong emphasis on understanding video sequences and collaboration among multiple robots.

At NVIDIA, the GR00T project follows the same trend. In March 2025, the company presented Isaac GR00T N1, an open foundation model for humanoid robots. NVIDIA also has a significant ecosystem around accelerated computing, simulation, and the Isaac platform. Comparison with Google DeepMind should remain measured: the companies do not necessarily disclose the same technical details, evaluation protocols, or access arrangements. But they share the goal of giving robots capabilities that are more general than approaches based on programming specific to each gesture.

The startup Physical Intelligence had also presented its π0 model in 2024, with the ambition of building generalizable physical intelligence. Here again, the debate concerns models’ ability to learn from diverse data, adapt to new environments, and carry out varied tasks. Figure, for its part, announced Helix in February 2025 as a vision-language-action model intended for its humanoid robots. These announcements show that the sector is no longer limited to traditional industrial robots, programmed for highly stable production cells.

The specificity of Google DeepMind’s announcement lies in the combination of the three terms highlighted by the source: video, orchestration, and collaboration. Many robotic demonstrations remain focused on successfully performing a unitary action: grasping, sorting, opening, storing, or moving. Yet a real operation often involves interdependent stages and multiple sources of information. Google DeepMind thus appears to be targeting a reasoning and coordination layer capable of linking multiple robots to a more dynamic understanding of the field.

This direction can also be viewed in light of the current limitations of language models. Large models are effective at synthesizing, rephrasing, generating code, or following instructions, but they can produce incorrect responses with an appearance of confidence. In robotics, this risk does not disappear: it must be offset by safeguards, checks, and controls at multiple levels. Good orchestration cannot be merely linguistic; it must be connected to the actual state of machines and physical constraints.

Competition is therefore taking place as much around models as around data, simulators, sensors, and deployment systems. Training a robot requires interaction data that is costly to obtain. Videos, human demonstrations, teleoperation data, and simulation are all possible sources, but they are not perfectly interchangeable. The transition from a test environment to an industrial site, often called transfer to the real world, remains one of the discipline’s major problems.

In this landscape, Google DeepMind has obvious strengths: longstanding expertise in machine learning, multimodal models, and robotics, as well as a connection to Google’s technology ecosystem. But the position of a laboratory or model provider alone is not enough to transform industrial facilities. Integrators, robot manufacturers, operations managers, and safety teams also determine the pace of deployments.

Fleet collaboration can precisely become an area of differentiation. A model that improves a single robot’s ability to manipulate an object addresses an important problem. A system that helps several robots operate consistently within a workflow addresses a question more directly tied to operating a site. It is also a more difficult question, because it requires handling resource availability, incidents, communications, priorities, and real-time trade-offs.

What the announcement means for France and Europe

For French and European companies, the interest of Gemini Robotics ER 2 is not limited to fascination with humanoids. The region includes many sectors in which robotics, automation, and logistics are already structurally important: manufacturing, automotive, aerospace, pharmaceuticals, food processing, retail, warehouses, and maintenance services. In these fields, the challenge is often less about buying a robot than about making heterogeneous equipment communicate while complying with a site’s specific constraints.

A system capable of interpreting videos, organizing missions, and coordinating multiple robots could interest organizations faced with variable processes. The most plausible uses are where operations involve frequent exceptions, configurations change, and a human operator remains necessary to supervise ambiguous situations. However, it would be premature to conclude that Gemini Robotics ER 2 is ready to be installed in French warehouses or factories: the publication cited does not provide, in the elements available here, details on a commercial launch or local deployments.

The European regulatory context adds a layer of complexity. Systems that use cameras in workplaces must take into account rules relating to personal data, particularly when employees or visitors may be filmed. The General Data Protection Regulation, or GDPR, establishes a framework for data minimization, purpose limitation, and data security. An architecture based on video understanding will therefore need to be assessed not only for its performance, but also for the nature of the images processed, their retention, and the access associated with them.

The European Union is also rolling out its artificial intelligence regulation, the AI Act. Without prejudging the exact legal classification of a given robotic system, its use in a professional context may raise obligations concerning governance, documentation, risk management, and human oversight. Applications involving moving machines must also comply with safety rules applicable to equipment and work environments. The power of a multimodal model does not reduce these requirements; on the contrary, it may increase the need to demonstrate how a decision is governed.

For European industrial companies, operational sovereignty is also at stake. A company deploying an AI-based orchestration layer must understand where its data is processed, what dependencies it creates on a provider, how it can audit the system’s decisions, and what alternatives exist in the event of unavailability. These questions are not specific to Google DeepMind, but they take on greater importance when AI becomes connected to physical operations and video streams.

Language and local procedures are another concrete issue. French sites work with specific instructions, nomenclatures, safety rules, and line-of-business software. The promise of agents capable of receiving human instructions is attractive, but it does not remove the need for integration. The system must understand the terms used in the field, respect authorizations, and be capable of being supervised by teams that are not necessarily specialists in foundation models.

The French market already has an industrial base in automation, industrial software, materials-handling robotics, and computer vision. The arrival of systems such as Gemini Robotics ER 2 could accelerate a shift: from the programmable robot to the robot configurable through more general objectives. But this transition does not mean the immediate disappearance of traditional approaches. Specialized robots often remain easier to validate, more predictable, and more efficient for repetitive tasks.

The social question will also need to be addressed without oversimplification. A coordinated fleet can change the distribution of tasks, increase maintenance and supervision needs, or shift certain skills toward operations control and exception management. The history of automation shows that the effects depend heavily on the sector, work organization, and training strategy. Communication around more autonomous robots must therefore not erase the role of operators, technicians, safety managers, and integrators.

Toward more autonomous physical systems, but still to be proven over time

Gemini Robotics ER 2 illustrates the gradual shift of generative AI toward action. The first wave of foundation models transformed the production and analysis of digital content. The next seeks to connect these capabilities to software, tools, and workflows. Robotics takes this movement further: the agent no longer merely acts in an application, it operates in an environment where objects have mass, people move, and errors have a material cost.

Video vision, orchestration, and coordination among multiple robots form a coherent set for this ambition. A dynamic understanding of the environment can improve perception of what is happening. Orchestration can turn a general instruction into steps. Fleet collaboration can connect these steps to multiple machines. But each of these building blocks brings its own difficulties, and their combination further increases complexity.

The next phase of competition will therefore not be decided solely by the quality of demonstrations. It will depend on players’ ability to demonstrate reliability on repeated tasks, at varied sites, with human operators, safety constraints, and acceptable costs. Companies will seek answers to very concrete questions: what happens when video is incomplete, when a robot is unavailable, when an instruction is ambiguous, or when an unexpected event disrupts the plan?

For Google DeepMind, the challenge is to show that Gemini’s capabilities can become useful infrastructure beyond digital interfaces. The publication on Gemini Robotics ER 2 places the company in the field of the coordinated physical agent, an area where models, robots, and operational software will have to work together. The choice to emphasize multiple robots rather than a single machine is revealing: the potential industrial value lies in the continuity of a process, not merely in the successful completion of a gesture.

This development could also reshuffle the cards between hardware manufacturers and model providers. If multimodal models become more capable of understanding scenes, interpreting instructions, and allocating tasks, the software orchestration layer will gain importance. Conversely, hardware constraints will serve as a reminder that a robot remains a complete system: sensors, actuators, power supply, control, safety, and maintenance cannot be abstracted away by a conversational interface.

In France as in Europe, the maturity of these systems will depend as much on their regulatory and industrial integration as on their demonstration performance. The most credible projects will probably be those targeting specific workflows, providing supervision mechanisms, and clearly measuring gains or limitations. The idea of a fleet of robots coordinated by a general-purpose AI is appealing, but it will have to accommodate the diversity of factories, data protection, and on-the-ground safety requirements.

The prospect opened by Gemini Robotics ER 2 is therefore less one of instantly replacing physical work than one of adding a new layer of intelligence to automated systems. If Google DeepMind succeeds in turning video understanding and embodied reasoning into robust coordination, AI agents could gradually move from the role of digital advisers to that of organizers of actions in the real world. It is an ambitious promise, whose validation will occur less in rhetoric than over time, through safety and the ability to handle the unexpected.

Back to all news

Comments· 2 comments

  1. Michael Johnson· 31 juillet 2026

    The article makes the coordination claim sound impressive, but it feels light on the practical limits. I would have liked more discussion of how these robots handle errors, conflicting priorities, or messy real-world environments rather than just the promise of fleet-level collaboration.

    1. Daniel Allen· 31 juillet 2026

      That is a fair concern, although an announcement piece may not be the place for every operational detail. The combination of video understanding and task orchestration still sounds like a meaningful direction, even if the real test will be whether it remains reliable outside controlled demonstrations.

Leave a comment