Safety policies, but a blind spot on containment

Major artificial intelligence laboratories now have well-established vocabularies for discussing safety: capability evaluations, risk levels, red teaming, deployment thresholds, governance mechanisms, human oversight, or access restrictions. These notions appear in the public documents of OpenAI, Anthropic, Google DeepMind, and other players engaged in the race for frontier models. They respond to growing pressure from researchers, public authorities, and professional users: as models acquire greater capabilities, it must be possible to demonstrate that they will not be deployed without proportionate safeguards.

But a far more concrete issue remains largely opaque: what would a laboratory actually do if an advanced model behaved unexpectedly, bypassed controls, used tools to which it has access, or displayed dangerous capabilities after entering service? This is the gap highlighted by a new study reported by TechCrunch AI in an article entitled “Frontier AI labs still won’t say how they’d contain a rogue model”. According to the elements presented by the outlet, frontier laboratories have published safety policies, but do not publicly detail the operational procedures that would make it possible to contain a model that has become uncontrollable.

The point deserves to be stated precisely. The study does not demonstrate that the companies concerned lack internal means. Rather, it highlights a limitation of publicly accessible documentation: promises, principles, and evaluation frameworks are more visible than response plans for an extreme incident. Yet this difference is essential. A policy can set ambitions and assign responsibilities; a containment plan must describe, or at least make auditable, the concrete chain of actions that makes it possible to detect a problem, limit the system's access, interrupt its activity, preserve evidence, and notify the relevant authorities or customers.

The term “rogue model,” used in TechCrunch's headline, does not necessarily refer to a science-fiction scenario in which an AI develops a will of its own. In the safety debate, it can cover several situations: a model that poorly pursues an assigned goal, an agent that performs unintended actions through connected software, a system that exploits excessively broad permissions, a cyber compromise, or the emergence of unexpected capabilities during large-scale use. The issue is therefore not merely whether a model produces an inaccurate or shocking response. It is about how to quickly limit its effects when it acts in a digital environment.

This distinction separates content risks from action risks. For several years, controversies around generative AI have focused primarily on hallucinations, bias, copyright infringement, disinformation, or the production of dangerous content. These issues remain central. However, the arrival of tools capable of browsing the web, calling software, writing and executing code in certain settings, or interacting with enterprise data shifts part of the debate. A system that is given goals, access, and a degree of autonomy is no longer judged solely by the quality of its responses: it is also judged by how well its permissions, scope of action, and shutdown mechanisms are controlled.

Publishing a safety policy is therefore not the same as demonstrating containment capability. The former generally explain how a company intends to measure risks before and during development. The latter should answer more operational questions: who has the authority to stop a service? Can a model or execution environment be rapidly isolated? What logs make it possible to understand its actions? How can the replication of access keys or the use of connected accounts be prevented? How can the response be coordinated with customers, hosts, partners, and authorities? The work reported by TechCrunch considers public answers to this type of question to remain insufficient.

Why the rise of agents makes the issue far more urgent

The question of containment is not new in computing. Companies have long had cybersecurity incident-response procedures: network segmentation, credential revocation, service shutdowns, backup restoration, forensic analysis, and notification of affected individuals or organizations. Critical systems, from industrial infrastructure to financial platforms, are also designed around principles of redundancy, access control, and emergency shutdown. But generative AI introduces a particular difficulty: a model can be integrated into many products, use different tools depending on the customer, and generate sequences of actions that are not entirely pre-written by its designers.

An isolated chatbot, without access to sensitive data or execution capability, already raises questions of reliability and security. An agent connected to email, a project-management tool, a browser, a document repository, or a development environment is a different case. Its potential consequences then depend less on the text generated than on the permissions granted. A highly capable model placed in a strictly compartmentalized environment does not have the same risk profile as a less capable model with broad and poorly controlled rights.

This is why contemporary debates about “agents” go beyond traditional benchmarks. Model evaluations remain necessary: they can provide information on performance in reasoning, programming, science, tool use, or other domains. But they do not exhaust the question. A high score does not by itself indicate what permissions are available in production, what controls block a sensitive action, or how long an organization takes to intervene when problematic behavior is detected.

The report mentioned by TechCrunch thus brings the debate back to very tangible elements: system access, autonomous action capabilities, and shutdown mechanisms. This approach has the merit of recalling that an incident is not limited to the existence of a dangerous capability. It also results from exposure. For a system to cause harm in a given context, there generally needs to be a combination of its capabilities, accessible tools, granted permissions, technical protections, and the quality of human oversight.

The concept of a “kill switch,” often invoked in public debate, must also be handled with caution. Turning off a public access point to a model may be relatively simple. But that does not automatically resolve every dimension of an incident. The same capabilities may be offered through several interfaces, integrated into partner products, run in separate environments, or serve users with their own configurations. A useful shutdown mechanism must therefore be conceived as part of a complete architecture: it must be known what is shut down, by whom, under what conditions, and what happens to data, sessions, access keys, and tasks already launched.

In practice, model providers have very different architectures. Some directly control most of the infrastructure hosting their models; others offer open weights or downloadable models, whose uses then largely escape their control. Some services are distributed through programming interfaces, while others are integrated into office, development, or research applications. This diversity does not make it possible to demand an identical formula from everyone. On the contrary, it makes transparency tailored to the system's actual distribution model necessary.

In this context, a public description of containment procedures does not necessarily mean revealing details that would make security easier to bypass. Laboratories may legitimately fear that overly precise documentation would expose sensitive elements of their infrastructure. But complete opacity raises another issue: users, regulators, investors, and organizations dependent on these services cannot easily assess the maturity of operational preparedness. Between disclosing security secrets and offering no verifiable guarantees, there is a more useful space for publication: decision-making roles, compartmentalization principles, crisis exercises, incident categories, notification timeframes, or independent validation.

What laboratories' public frameworks already say — and what they do not make it possible to verify

The leading developers of advanced models have not ignored safety. OpenAI has published a Preparedness Framework, designed to track certain risks related to the capabilities of its systems and to define protective measures before their deployment. Anthropic has made public a Responsible Scaling Policy, which links the growth of its models' capabilities with safety standards and governance commitments. Google DeepMind has published a Frontier Safety Framework, presented as an approach to identifying and mitigating serious risks associated with frontier capabilities.

These documents have real importance. They have helped establish the idea that the development of powerful models cannot be separated from evaluation work, safety research, and internal escalation processes. They have also provided a common language for institutions seeking to regulate the sector. The fact that companies publish them is already a notable difference from earlier periods, when decisions to launch AI products were almost entirely a matter of commercial communication and invisible internal procedures.

Nevertheless, a safety policy is not automatically an incident-management manual. It may indicate that a model must not cross a certain risk threshold without additional protections, without detailing how the company would respond to a problem detected after deployment. It may set out evaluation categories without specifying which systems can isolate a model instance. It may provide for governance reviews without clearly indicating which people have the power to suspend a product when commercial interests, service continuity, and safety imperatives conflict.

This is the core of the criticism documented in the study cited by TechCrunch: laboratories publish more about their doctrines than about their containment procedures. The diagnosis concerns public verifiability, not the certain existence or nonexistence of private arrangements. This nuance is decisive in a sector where companies consider a significant part of their infrastructure, evaluation datasets, and security practices to be confidential.

Public documents also remain difficult to compare. Risk definitions vary across organizations, as do thresholds, evaluation methods, and the wording used. Some frameworks are centered on model capabilities that could facilitate seriously harmful uses; others give greater weight to the behavior of autonomous systems, cybersecurity, or governance. This diversity reflects the plurality of risks, but it complicates the work of external observers. A company may claim to have “robust” measures without the public being able to readily establish an equivalence with the guarantees announced by a competitor.

There is also a difference in timing. Development policies are often designed before a model's launch or when new capability milestones are crossed. Containment, by contrast, is a real-time issue. A useful procedure must work when information is incomplete, when the source of the problem is not yet known, and when a service interruption may have significant consequences for customers. It requires trained teams, clear responsibilities, and technical tools whose effectiveness should ideally be tested regularly.

In the software industry, incident-response exercises and crisis simulations are a known practice. In the case of frontier models, the question becomes more complex because the object of the incident may be emergent behavior, an unexpected interaction with tools, an external vulnerability exploited by a malicious user, or a specific product configuration. The quality of the response therefore does not depend solely on the model provider. It also depends on the integrator, the customer that granted permissions, the host, and sometimes third-party software publishers.

A regulatory issue now central in Europe and France

The debate over containment plans comes at a time when the European Union is progressively building its regulatory framework for artificial intelligence. The European AI Act introduces an architecture based on risk levels and includes provisions targeting general-purpose AI models, as well as models presenting systemic risks. The European text is not limited to the hypothetical case of an AI “out of control”: it also covers documentation, transparency, evaluation, risk management, and, depending on the circumstances, obligations to report serious incidents.

The issue raised by TechCrunch naturally fits into this logic. If advanced models become components used by many companies, regulators will need to know not only how they were evaluated before commercialization, but also how their designers respond in the event of an incident. Authorities must be able to distinguish a general safety promise from demonstrable organizational capability. This distinction also applies to public buyers and large companies, which cannot completely delegate their own security responsibility to an AI provider.

For French stakeholders, the question has concrete implications. Companies that integrate generative models into their customer-service, document-management, software-development, or internal-analysis tools must examine the rights actually granted to systems. They must also clarify the conditions for suspending a service, the reversibility of integrations, access to activity logs, and notification arrangements when an incident affects their data or processes. A general policy published by an American or international laboratory may be useful, but it does not replace the contractual and technical requirements necessary for each deployment.

France already has institutional players engaged on these issues, notably the CNIL for personal data issues and ANSSI for cybersecurity. The debate on containing AI systems lies at the intersection of their concerns, without being limited to them. An AI with permissions can become a vector for data leakage, process manipulation, or the exposure of computer systems. But it also raises specific issues of behavioral evaluation, autonomy, and control of automated decisions.

French and European companies are also frequently customers of models and platforms developed outside Europe. This increases the importance of transparency. A contracting entity can hardly assess the consequences of technological dependence if it does not know what guarantees of suspension, recovery, and incident communication its provider is prepared to offer. This requirement is particularly sensitive in regulated sectors, public administrations, healthcare, finance, energy, or companies handling sensitive industrial information.

An obligation to publish an entire incident-response plan would not necessarily be the most effective solution. Some details could be exploited by attackers, while procedures would in any case have to evolve with architectures and products. Regulators could, however, require standardized and auditable elements: identification of those responsible, escalation mechanisms, categories of shutdown measures, notification conditions, exercise results, or checks carried out by qualified third parties. The issue would not be to obtain companies' operational secrets, but to verify that the capabilities they announce actually exist.

This approach recalls how other sectors with major security stakes are handled. The public does not know the details of the security systems of an aircraft, a payment network, or critical infrastructure. But operators must comply with standards, keep records, report certain events, and submit to inspections. The AI sector is still looking for its institutional equivalent, while its products are being deployed at very high speed within organizations.

Beyond promises: toward demonstrable safety for autonomous systems

The main contribution of the study reported by TechCrunch is to shift the discussion from intentions to preparedness. A company can publish ambitious safety principles and fund specialized teams. That does not answer every question about its ability to manage a rare, complex, and potentially costly situation. Conversely, the absence of public details does not make it possible to conclude that internal procedures are absent. What opacity does produce, however, is a collective difficulty in judging the gaps between laboratories.

This difficulty is compounded by competition. Companies developing the most advanced models are engaged in a technological and commercial race in which launch speed is an advantage. In such an environment, announcements of safety policies also play a signaling role: they reassure customers, regulators, and investors while allowing laboratories to state that they take risks seriously. The most demanding test, however, is not the existence of the document. It is the ability to accept a delay, a feature limitation, or a deployment suspension if protections are not deemed sufficient.

Comparisons between OpenAI, Anthropic, Google DeepMind, and other developers must therefore avoid two pitfalls. The first would be to consider that all public frameworks are equivalent because they use similar terms. The second would be to conclude that a company is necessarily reckless because it does not disclose its most sensitive procedures. A rigorous assessment should consider the whole picture: the nature of the model, access arrangements, degree of autonomy, available tools, deployment policies, security organization, reporting mechanisms, and possibilities for independent oversight.

The notion of human control also deserves clarification. It is often invoked as a general guarantee, even though it can cover very different realities. A human who validates every sensitive action does not exercise the same control as a human who intervenes only afterward. Likewise, an oversight team can act usefully only if it has relevant alerts, sufficient technical rights, and formal authority to interrupt a system. In rapid or large-scale processes, purely symbolic oversight can become insufficient.

For providers, the challenge will be to demonstrate that security is built into product architecture rather than added as a layer of communication. This notably involves limiting permissions by default, segmenting environments, logging actions, reducing possibilities for data exfiltration, and providing checkpoints before irreversible operations. These principles are not specific to AI, but they take on new importance when a system is capable of chaining tasks together and adapting its actions to the information it encounters.

For professional users, the debate calls for a form of realism. Agent adoption cannot be based solely on productivity demonstrations. It requires governance of access and responsibilities. Before connecting a model to email, a payment tool, a production system, or sensitive data, an organization must be able to answer simple questions: what actions are authorized? Who can modify them? What are the technical limits? How can a task be canceled or interrupted? Who is notified in the event of an anomaly? And what share of the response depends on the provider, the customer, or a third-party service provider?

In the long term, transparency about containment could become a market criterion as important as price, speed, or model quality. Large organizations will probably seek providers able to prove their operational resilience, not merely their performance on evaluations. Regulators, for their part, will be pushed to turn general notions of “security” into testable obligations, without requiring disclosure that would weaken infrastructure.

The debate opened by TechCrunch is therefore not solely about an extreme hypothesis. It concerns the maturity of an industry seeking to move its models from the status of conversational tools to that of systems capable of acting in real environments. The more this autonomy advances, the more public safety policies will need to be accompanied by evidence of preparedness: who can stop what, within what timeframe, with what guarantees, and under what external oversight. It is this ability to make incident response credible and verifiable that will determine a growing share of trust in frontier AI.

Back to all news

Comments· 1 comment

  1. Ryan Walker· 24 août 2026

    Thank you for highlighting this—it's encouraging to see more attention on the need for clear, public AI safety plans.

Leave a comment