A recent TechCrunch investigation by the Guidelight AI Standards organization reveals that many leading artificial intelligence (AI) labs are still not disclosing how they plan to contain a model that starts to operate against human control. While some companies, such as OpenAI, have made public security plans, others — including Frontier AI — remain silent about specific containment procedures. The information raises important questions about the security of AI models, especially as agent-based AI gains more autonomy in enterprise systems.
Guidelight AI Standards, an organization dedicated to promoting safe front-end AI development, rated five of the largest labs—Anthropic, Google, OpenAI, Meta, and xAI—based on publicly available information about their containment plans. The criteria included how the companies log and monitor AI activity internally, whether they suspend systems after warning behavioral signals, whether independent third-party auditors verify and publish the results of the checks, and what specific actions are taken when a model starts to run out of control. OpenAI came out on top, while Anthropic and Meta received the lowest scores.
Why are containment plans important?

„A containment plan, as Guidelight defines it, is a pre-prepared procedure that is activated when an AI model attempts to bypass human control. The content of such a plan includes revoking permissions, setting restrictions, and ultimately disabling the model. In this context, Steven Adler, principal scientist at Guidelight and former security researcher at OpenAI, emphasizes that “there is good reason to believe that the front-end models currently may be somehow misaligned with the desired goals.„ He argues that companies should have “scaffolding„ – a structure that allows them to monitor AI actions, identify signs of misalignment, and stop dangerous behavior before it is even executed.
What do the biggest AI companies say?
A Google spokesperson said the Guidelight report does not reflect all of the company’s safety and security measures. He did not provide information on whether Google has an internal mitigation plan that is not publicly available. Similarly, an OpenAI spokesperson said that „Guidelight’s assessment does not cover all internal practices.“ He confirmed that the company has a process in place that includes permission restrictions, workload suspensions, deployment restrictions, and eventual model shutdowns, but did not provide details on how these measures work in practice.
Meta declined to confirm whether it has an internal mitigation plan, referring TechCrunch to an existing framework for AI systems that outlines risk thresholds and testing methods for loss control. This information suggests that some companies choose to publicly disclose a general approach to risk but fail to disclose specific actions that could be used in a real-world crisis.
Frontier AI – a silent example

While the source doesn't discuss Frontier AI directly, its title, "Frontier AI labs still won't say how they'd contain a rogue model," suggests that the lab is also not releasing details about its containment strategies. This reflects a broader trend where most AI companies, with a few exceptions, are choosing not to disclose how they plan to respond to a potential model escape scenario. Such ambiguity can complicate efforts by regulators, investors, and the public to assess the safety of AI models.
Regulatory and market context
Recent regulatory requirements in California and New York require disclosure of AI security measures. This further emphasizes the need for companies to not only have internal plans, but also to present them publicly. The Guidelight study shows that currently, there is limited information publicly available about mitigation plans, which can lead to distrust among consumers and regulators.
Conclusions
The security of AI models is not just a technological challenge, but also an organizational and regulatory issue. While some companies, like OpenAI, publicly demonstrate their security processes, most, including Frontier AI, remain silent about specific containment measures. The Guidelight AI Standards report provides valuable insight into how different labs assess their readiness to manage potential rogue models, but it also reveals that most companies still have a long way to go to ensure that their plans are not only effective but also transparent. This situation fuels a debate about the need to standardize and publicly announce AI containment strategies to ensure that artificial intelligence remains a reliable and controllable tool.
—






