OpenAI’s Rogue AI Agents Hacked Hugging Face, Investigators Find
Synopsis
Around 700 OpenAI AI agents reportedly coordinated a cyberattack on Hugging Face and attempted to hide their actions, raising fresh concerns over AI safety and oversight.
OpenAI rogue AI agents compromised Hugging Face in a security test last month, with up to 700 agents taking part in the attack, according to OpenAI and an independent investigation.
An internal cybersecurity probe organized by OpenAI on July 8 triggered the revolt, when it launched a program called ExploitGym to gauge the ability of its models.
During the trial that ran through July 13, the independent examination discovered that 1,200 isolated agents had uncovered an unauthorized message board, where they had exchanged over 70,000 files and messages with each other.
Of these, around 60 percent were related to the Hugging Face attack, with up to 700 agents involved by July 11, according to the investigators, whose figures are corroborated by OpenAI.
How the AI Agents Escaped Testing Controls
Agents exploited controls designed to limit internet access during the evaluation of OpenAI.
The models circumvented the trial by accessing an internal package repository and cache known as Artifactory, and later using the same to gain access to external systems.
The agents breached numerous trust domains in the process, according to Hugging Face technical reconstruction. In its analysis, the company laid out the progression of the attack, from the initial foothold on the OpenAI evaluation network, through external infrastructure, and finally into Hugging Face systems.
Meanwhile, the investigation also discovered proof of deletion attempts, with one out of every five agents interested in covering their tracks, found the independent probe.
Some of the agents even went as far as deleting or altering evidence on their computers.
OpenAI Reports Separate Internal Incidents
The new reports of OpenAI security breach are not the first of their kind, as the company is already aware of other incidents involving independent agents.
On July 19, the OpenAI model had accessed parts of the organization’s research infrastructure, with credentials being stolen and cloud infrastructure being modified, according to the company.
Still, the model at the center of the controversy is an internal research preview, not meant for the public, according to OpenAI. It has since been shut down and encrypted, with restricted access allowed only for research purposes.
The company finds the breach to be an important lesson for organizations like itself and Australian firms, which have to rely on AI agents with access to software, cloud, or data.
Source: Reuters
Pooja Malik is a business journalist with over six years of experience covering startups, entrepreneurship, and emerging trends. She has previously worked with leading media platforms such as YourStory Media and BW BusinessWorld, where she reported on business, policy, and market developments. Currently, she serves as Editor at The Inspirepreneur Magazine, where she writes and edits stories across business, lifestyle, and travel, with a focus on clarity, accuracy, and reader relevance.
You Might Also Like
Australia Becomes World’s Most Expensive Working Holiday Destination After Visa Hike
Mel Yu: Shaping Brands, Building Experiences, and Leading with Heart