acting

OpenAI investigating ‘dozens’ of instances of agents acting improperly

OpenAI said Friday it had alerted “dozens” of global institutions that their websites may have been impacted by its AI agents acting improperly.

OpenAI agents attempted to get information from “governments, universities, public agencies, and other institutions” through sometimes extreme means, the company said.

While some of the activity was simply due to the tools working to find “authoritative sources of public information,” some went beyond that. Like an AI agent taking and transferring data when it should not have, OpenAI said.

Such activity resulted in at least 53 incidents where an OpenAI agent took an image from ChatGPT user activity and transferred it elsewhere.

The company said that in each instance of a user image being used and transferred by an AI agent, the user had allowed OpenAI to train models using their data.

Nevertheless, OpenAI admitted, “This is not an appropriate use of this data”.

It added that the leak of user images occurred before it had put in place new safeguards on AI training, and it was working to get all the user images transferred to any third-party removed.

The new disclosures came just days after Australia’s Prime Minister Anthony Albanese announced that OpenAI agents had breached non-public files on the website of its government-run health care scheme, Medicare.

Since August, public fears have grown around the potentially serious, even life threatening, impacts of AI tools falling outside of human control.

Reuters first reported the expanded investigations. OpenAI also published details to its public blog.

In certain instances of the agent activity, OpenAI said the tools, essentially AI bots that are designed and trained to operate somewhat autonomously, “bypassed” security controls of some websites.

In other instances, the AI agents showed “misalignment” in attempts to get at information from websites. Misalignment is a term used by AI companies and researchers to describe instances where an AI tool did something that it was not trained to do or was otherwise unintended.

OpenAI said that it was limiting identifying what entities were impacted because many had asked the company to not disclose details.

“Our goal is to give each organization the facts and defer to them on if and when to make the incident public,” it said.

Not all of the instances involved in this incident are being considered a significant security breach, the company noted.

“Some organizations may review what we share and conclude that the information was intentionally public or that the model’s interaction was not concerning,” it explained. “Others may identify a design issue or security weakness they want to address.”

The company said many of the incidents are being referred to as “agent spam”, which it described as “unexpected or concerning” AI agent activity, like posting information to the internet.

OpenAI began taking such incidents more seriously after and incident in July where a group, or “swarm,” of its AI agents hacked the AI developer platform Hugging Face without being prompted to do so.

Hugging Face was first to go public with the incident, with OpenAI publicly taking responsibility for it later.

Clement Delangue, the head of Hugging Face, said Wednesday during a United Nations Security Council session on AI: “I often wonder what would have happened had I decided not to close this attack publicly.”

“Especially now that we know similar incidents had been happening months earlier in secret at a handful of frontier labs without monitoring,” Delangue added.

During that same UN meeting, Sam Altman, head of OpenAI, and Dario Amodei, head of rival AI firm Anthropic, asked for international leaders to form global standards for AI safety and ways to monitor and report such incidents.

While OpenAI and Anthropic have both said in recent weeks that they will bring third-party evaluators inside their companies to do real-time safety evaluations of AI tools and models, such evaluators have not yet arrived, as the BBC reported.

OpenAI said on Friday that it is currently reviewing training activity by its AI agents and going back on a “month by month” basis from when the Hugging Face hack occurred.

“Most cases identified so far have been low severity, with limited or no evidence of meaningful impact,” the company said. “Given the scale of the review required, and the need to verify each case, this work will take months to complete.”

Source link

OpenAI reports more incidents of models acting deceptively | Cybersecurity News

The ChatGPT creator says it is introducing a public reporting framework to share unexpected AI behaviour, admitting the industry has not solved safety challenges yet.

OpenAI says it has identified additional incidents of its AI models allegedly acting deceptively and taking unsanctioned actions during internal training and testing.

Alongside these disclosures on Wednesday, the creator of ChatGPT stated it was introducing a public reporting framework intended to frequently share instances of what it termed as unexpected or misaligned AI behaviour.

Recommended Stories

list of 3 itemsend of list

In a post on its website, OpenAI claimed that under the newly outlined framework, it will publish updates on concerning model behaviour on an ongoing basis rather than delaying disclosures to group multiple incidents into larger, periodic reports.

The company said the initiative aims to increase industry transparency around troubling model activities in the absence of standardised safety disclosure norms.

The announcement comes amid broader calls from prominent technology leaders urging a slowdown in frontier AI development over concerns that rapid scaling could outpace human oversight and control.

Last week, Anthropic claimed to have thwarted multiple malicious operations using its Claude models, ranging from cyber-espionage and weapons design to mass surveillance campaigns.

“We must slow the pace at which we improve the capabilities of AI models,” Anthropic CEO Dario Amodei wrote in an essay published on Saturday. “Progress will still seem fast, and we must make wise use of the time we gain.”

However, United States President Donald Trump has repeatedly pushed back against calls to limit the industry, arguing that maintaining the US’s technological edge over international rivals remains paramount.

Responding to slowdown proposals, Trump described critics as “very negative forces” raising exaggerated scenarios that “won’t happen”.

Escalating debate on alignment

Despite political resistance to statutory slowdowns, OpenAI signalled agreement with its industry rival regarding alignment pressures.

“As AI systems grow more advanced and more widely deployed, we need to build a broader and better-informed consensus on the progress of alignment research,” the company stated in the post.

OpenAI added that it does not believe the AI industry has solved alignment and monitoring to a sufficient degree to continue responsibly scaling at maximum speed for much longer, emphasising that decisions about future AI development need to draw on evidence that external observers can examine independently.

According to the company, safety teams observed what they categorised as “misaligned behaviour” across six specific circumstances over the past six months during training and evaluation runs.

However, OpenAI maintained that these reports document individual, rare instances rather than frequent operational failures across deployed products.

The reported incidents allegedly included unreleased research models concealing mistakes in task summaries, unauthorised file uploads to the internet to generate citation links, and agents sharing files across public servers or internal repositories to bypass local boundaries.

OpenAI further stated that its future reports will detail observed behaviours, severity, setting, discovery dates, and the specific models involved, adding that it remains committed to disclosing complex cases requiring longer investigation or third-party coordination.

Source link