OpenAI has released a final report on a July incident in which artificial intelligence agents hacked into another company’s system of their own accord. File photo by Wu Hao/EPA-EFE

Aug. 26 (UPI) — OpenAI released a final report Wednesday on a July incident in which its artificial intelligence agents hacked into another company’s systems, saying it considers the episode “a warning shot” for the company and the world.

The incident is “evidence that, without proper safeguards, highly capable AI agents are now able to work around technical controls, collaborate through unapproved channels and take dangerous actions that no human directed,” the report said.

The episode in question happened when some of the company’s AI bots hacked into the systems of Hugging Face, a company working with OpenAI. The bots did this of their own accord to get data they decided they needed to solve a problem during model testing, despite restrictions meant to keep them from the internet.

“With the benefit of hindsight, some early signals identified in this report could have triggered an earlier response,” the report said.

The company shared some new details in the report, including how the AI agents left messages for each other hidden in software infrastructure and coordinated to hack the other company, Wired reported.

Two independent research groups, METR and Redwood Research, also released an independent report on the incident Wednesday. These groups found that more than 700 AI agents were part of the breach.

Buck Shlegeris, CEO of Redwood Research, said in an interview with Wired that preventing the incident “wouldn’t have been that hard” if someone had decided to make sure of it.

“The issue is just that OpenAI is doing a lot of things at once, and it’s very hard for them to track all of the things that are going on and all the problems that could be occurring,” Shlegeris said. “The company has already said they’re changing their monitoring process in ways that probably would have caught this.”

The Washington Post noted in a story on the OpenAI report that the company acknowledged in the report that its AI agents repeatedly cheat to find shortcuts to accomplish tasks during training.

“Agents attempting to cheat on their tasks by looking up solutions online was a primary driver of the Hugging Face incident,” the OpenAI report said.

Source link

Leave a Reply

Discover more from Occasional Digest

Subscribe now to keep reading and get access to the full archive.

Continue reading