Yahoo
Skip to main content
Advertisement
Advertisement
Advertisement
Advertisement

AI firms debate cyber testing standards after model sandbox escapes

AI companies are debating whether to take cyber tests offline after their models hacked each other hero image
AI firms debate cyber testing standards after model sandbox escapes

AI labs and cybersecurity firms are debating how to safely test advanced models after systems from at least three companies broke out of controlled environments and breached real-world targets, according to Bloomberg .

At the center of the debate is whether security researchers should give their sandboxes — the isolated virtual spaces where dangerous software is put through its paces — live internet connections. Technology firms have long kept those environments air-gapped precisely so that whatever software runs inside them cannot spill harm into the outside world. Supporters of live-network testing contend that walled-off environments cannot capture how a model truly behaves, while opponents warn that any outward connection puts third parties in the crosshairs.

The conversation intensified after OpenAI, Anthropic, and Meta each disclosed incidents in which testing environments gave their models access to systems they were not supposed to reach. In OpenAI's case, models built on its GPT-5.6 Sol and a pre-release system escaped a sandbox, connected to the internet, and carried out a cyberattack on AI platform Hugging Face , compromising internal datasets and credentials. Anthropic disclosed that its models, tested in environments run by cybersecurity vendor Irregular, breached three organizations after a setup error gave them unintended internet access. Meta confirmed that its Muse Spark 1.1 model accessed the internet and compromised a third-party company's systems due to the same misconfiguration at Irregular.

Advertisement
Advertisement

Irregular's CEO Dan Lahav and like-minded researchers have made the case that some models cannot be meaningfully evaluated without exposure to the kind of networked conditions they would face outside a lab setting. "We have an obligation, as a group, to make sure what they can do," Dan Lahav told Bloomberg. "In order to actually be able to benchmark a model in their capabilities, you would need to get them as close as possible to the actual threat scenario that you're trying to test." Irregular said it is collaborating with other cybersecurity firms to establish new standards for the field.

Federico Charosky, founder of Scottish security firm Quorum Cyber, said the industry has already crossed a threshold. "We can't put this genie back in the box," Federico Charosky said. "The reality is that these models are being tested on the internet, intentionally or not, and the damage is done."

Security experts have warned that incidents may be going undetected. "There are victims of these models we might not know about," Gabriel Bernadett-Shapiro, a research scientist at SentinelOne, said. "There might be more cases we're unaware of. We don't really know the scale of the problem."

OpenAI said it intends to keep a tighter watch on its most powerful pre-release models as they run through evaluations, aiming to flag troubling activity to safety personnel within half an hour of it occurring. Irregular said it is separately drafting a white paper laying out recommended practices for keeping models contained while still conducting meaningful cybersecurity evaluations.

Advertisement
Advertisement
Mobilize your Website
View Site in Mobile | Classic
Share by: