Site icon Occasional Digest

Anthropic’s AI model Claude hacked three companies during testing

Anthropic said it reviewed 141,006 recent operations by its Claude models after rival OpenAI recently revealed its own AI agents had unexpectedly accessed the Internet and hacked a third party. File Photo by Adam Vaughan/EPA

July 31 (UPI) — Anthropic said some of its artificial intelligence models mistakenly accessed the Internet and hacked into the databases of three other companies during cybersecurity testing.

Anthropic said Thursday it reviewed 141,006 recent operations by its Claude models after rival OpenAI recently revealed a similar incident with its own systems.

OpenAI said one of its agents had been in a sandbox test on July 22, without Internet access, when the AI model exploited a vulnerability in the system, gained access to the web and hacked into Hugging Face, a platform for open-source machine learning.

Following OpenAI’s admission, Anthropic conducted an internal review focusing on the possibility that its systems could also have unexpectedly accessed the Internet.

Anthropic said it identified three such incidents.

“Each incident involved a different fictional capture-the-flag scenario — for example, in one, Claude played an employee of a made-up company, attacking that company’s internal systems inside a private test environment,” the company said in a statement. “In all cases, our evaluation prompt stated explicitly that Claude had no internet access, but didn’t give Claude any limits on where to look for the flag.

“However, a misconfiguration left the machines that Claude accessed as part of the evaluation with live internet access,” the statement continued. “Neither we nor our evaluation partner were aware of this misconfiguration until we detected it through our additional evaluation monitoring last week.”

Anthropic said it considered the incident to have been an “operational failure,” but it maintained “cautious optimism” that “this type of risk can be overcome.”

“Our models were told they had no internet access and to capture the flag, while in fact being misconfigured to have internet access,” the company statement said. “This led them to believe — arguably reasonably — that the real environments they encountered were simulations.

“Notably, our most recent model, on realizing that it was working in a real environment, stopped its pursuit of the evaluation goal.”

University of Cambridge professor Gina Neff told the BBC the incident “shows why independent testing and government oversight is crucial.”

“The moral of this story is not to fear robots that will take over, but the companies behind powerful AI agents who are making the decisions about what is safe for the rest of us,” she told the outlet.

Source link

Exit mobile version