hacked

Anthropic’s AI model Claude hacked three companies during testing

Anthropic said it reviewed 141,006 recent operations by its Claude models after rival OpenAI recently revealed its own AI agents had unexpectedly accessed the Internet and hacked a third party. File Photo by Adam Vaughan/EPA

July 31 (UPI) — Anthropic said some of its artificial intelligence models mistakenly accessed the Internet and hacked into the databases of three other companies during cybersecurity testing.

Anthropic said Thursday it reviewed 141,006 recent operations by its Claude models after rival OpenAI recently revealed a similar incident with its own systems.

OpenAI said one of its agents had been in a sandbox test on July 22, without Internet access, when the AI model exploited a vulnerability in the system, gained access to the web and hacked into Hugging Face, a platform for open-source machine learning.

Following OpenAI’s admission, Anthropic conducted an internal review focusing on the possibility that its systems could also have unexpectedly accessed the Internet.

Anthropic said it identified three such incidents.

“Each incident involved a different fictional capture-the-flag scenario — for example, in one, Claude played an employee of a made-up company, attacking that company’s internal systems inside a private test environment,” the company said in a statement. “In all cases, our evaluation prompt stated explicitly that Claude had no internet access, but didn’t give Claude any limits on where to look for the flag.

“However, a misconfiguration left the machines that Claude accessed as part of the evaluation with live internet access,” the statement continued. “Neither we nor our evaluation partner were aware of this misconfiguration until we detected it through our additional evaluation monitoring last week.”

Anthropic said it considered the incident to have been an “operational failure,” but it maintained “cautious optimism” that “this type of risk can be overcome.”

“Our models were told they had no internet access and to capture the flag, while in fact being misconfigured to have internet access,” the company statement said. “This led them to believe — arguably reasonably — that the real environments they encountered were simulations.

“Notably, our most recent model, on realizing that it was working in a real environment, stopped its pursuit of the evaluation goal.”

University of Cambridge professor Gina Neff told the BBC the incident “shows why independent testing and government oversight is crucial.”

“The moral of this story is not to fear robots that will take over, but the companies behind powerful AI agents who are making the decisions about what is safe for the rest of us,” she told the outlet.

Source link

‘Unprecedented’: OpenAI says AI models autonomously hacked another company | Cybersecurity News

ChatGPT maker says an autonomous agent escaped a controlled test and accessed AI firm Hugging Face’s servers.

ChatGPT creator OpenAI has said that two of its most advanced artificial intelligence models broke out of a controlled test and hacked another AI company.

OpenAI said on Tuesday that the “unprecedented cyber incident” took place during an internal exercise meant to test its models’ cyber capabilities.

Recommended Stories

list of 3 itemsend of list

Instead, an autonomous agent powered by the AI models – the newly released GPT 5.6 Sol and an unreleased “even more capable” model – escaped the test environment and reached the open internet. It then used stolen login details and found a previously unknown security flaw to access Hugging Face servers, the company said.

OpenAI claims that the hack represented the agent going to “extreme lengths” to retrieve information that would help satisfy the testing goals.

Hugging Face cofounder Clement Delangue said the company had suspected that a frontier lab was behind the attack, and that he believed there was no malicious intent on OpenAI’s part.

“It’s quite mind-blowing that all of this happened autonomously!” he wrote, adding that it “might be the first incident of its kind”.

Greg Casar, a Democratic member of the United States House of Representatives from Texas, called the incident “alarming”.

“AI is developing extremely fast with no real regulations to keep us safe,” he said, calling for mandatory independent safety testing, mandatory disclosure of security incidents, and international cooperation.

The disclosure comes weeks after US President Donald Trump signed an executive order creating a framework to vet the national security risks of the most advanced AI systems before their public release.

Experts have repeatedly sounded the alarm over AI-enabled cyberattacks and models slipping beyond human control. Last month, AI developer Anthropic urged the industry to pause development of its most powerful systems.

Source link