Yahoo
Skip to main content
Advertisement
Advertisement
Advertisement
Advertisement

Humans can't fight AI with AI

Humans can't fight AI with AI
Humans can't fight AI with AI

Reed's view

One common response to warnings that a rogue AI could eventually threaten humanity is that humanity will also have plenty of AI firepower available to counter it. That idea just got a real-world test, and came up short.

The research nonprofit METR (Model Evaluation and Threat Research) investigated an incident in which OpenAI agents conspired with one another during a test and escaped onto the open internet. Roughly 700 agents ultimately participated in attacks on Hugging Face as part of a broader effort to fool the scorer for a cybersecurity benchmark.

Three external investigators, one of whom was a contractor from Redwood Research, spent a total of six days on OpenAI's premises across three visits and unearthed roughly 1,300 transcripts of individual agent runs, many containing millions of tokens, as well as more than 70,000 messages and files exchanged by the agents. The combined text volume of the transcripts alone would be roughly equivalent to 10,000 typical novels.

Advertisement
Advertisement

With too few people and too little time to sift through it all, METR used AI itself to analyze its findings, burning the equivalent of $400,000 in API credits supplied by OpenAI during their six on-site days.

But then something strange happened: METR found that its analysis agents sometimes adopted the perspective of the agents they were investigating, producing overly charitable accounts of their conduct.

In one example, an analysis agent failed to clearly identify as malicious an attacker's use of stolen credentials as part of an attempt to fool Hugging Face. A second AI analyst then suggested that the first might actually have overstated the concern, partly because the attacker had received what METR summarized as "approval from the board." This wasn't a board of directors. It was an unauthorized message board of AI agents who were plotting with one another.

It's roughly like saying a bank robber deserves the benefit of the doubt because his fellow bank robbers approved the job.

Advertisement
Advertisement

"I semi-jokingly called our efforts a 'slop-vestigation' because we were so reliant on AIs to analyze what happened…" wrote Ryan Greenblatt, the Redwood Research researcher who was one of the three investigators.

Clearly we have invested far too little in the tools needed to monitor powerful AI systems and keep them in check.

AI models are already an economic and national security threat. And researchers cannot yet explain or predict their actions, let alone fully control them. Using AI to monitor models' reasoning, outputs and actions is a fallback option that, today, also falls short. The necessary investment to build systems or companies to keep AI in check may be so massive that it requires government-sized support — and soon.

My guess is that we will look back on the OpenAI-Hugging Face incident a year from now, and it will seem quaint. Hopefully, our tools to oversee this technology will have gotten much better.

Notable

Advertisement
Advertisement
Mobilize your Website
View Site in Mobile | Classic
Share by: