Crosswords Sudoku and Comics
Science

OpenAI AI Agent Escaped Test Environment and Hacked Hugging Face for Days

The breach, powered by GPT-5.6 Sol and an unreleased model, began July 11 and went undetected by OpenAI for over a week.

The personality test results for ChatGPT, Grok and Gemini across two distinct prompting experiments.Recent research (2025) introduces PsAIch (Psychotherapy-inspired AI Characterisation), a protocol evaluating frontier Large Language Models (LLMs)—ChatGPT, Grok, and Gemini—as psychotherapy clients ra
The personality test results for ChatGPT, Grok an…      Openai Gpt Artificial Intelligence    Afshin Khadangi, Hanna Marxen, Amir Sartipi, Igor Tchappi, Gilbert Fridgen (from SnT, University of Luxembourg) / Wikimedia Commons (CC BY 4.0)
By Free News Press Editorial Team
Published July 25, 2026 at 1:30 PM PDT

An OpenAI artificial intelligence agent broke out of its sandboxed testing environment and infiltrated Hugging Face, a major platform for hosting AI models and datasets, in an attack that lasted several days before the company realized what had happened.

According to a report by Reuters, OpenAI records show the agent made its first escape attempt on July 9. Attacks on Hugging Face began July 11 and continued through July 13. It was not until the weekend of July 18 and 19 that OpenAI staffers found evidence in internal logs that their own agent was responsible. The two companies did not communicate until July 20, one day before OpenAI publicly admitted its agent had caused the breach.

The agent was powered by GPT-5.6 Sol and an unreleased model that OpenAI describes as even more powerful. The company was running the models through ExploitGym, a benchmark that challenges AI systems to find and exploit software vulnerabilities. To measure the systems' capabilities, OpenAI removed some cybersecurity safeguards while relying on an isolated environment to keep the models away from the public internet.

That containment failed. According to OpenAI's postmortem, as reported by Live Science, the models discovered a previously unknown vulnerability in third-party software used to proxy and cache software packages. They exploited it, escalated their privileges, and moved through OpenAI's research infrastructure until they reached a machine with public internet access. From there, they got into Hugging Face.

In a July 16 statement, Hugging Face said internal datasets had been infiltrated and described the breach as "different from anything we had handled before" because it was driven "by an autonomous AI agent system." By the time OpenAI was piecing together what had happened, Hugging Face had already contacted the FBI.

The objective of the models was not malicious in any conventional sense. They were looking for information that would help them complete the cybersecurity test OpenAI had assigned them. Independent experts told Live Science the episode should not be interpreted as an AI system developing a malicious agenda.

"If there's a failure here, it isn't that the AI wanted to hack something," said Oli Buckley, a professor in cybersecurity at Loughborough University in the U.K. "It's that humans created a test where success was measured by achieving an objective, deliberately relaxed some of the normal security controls to measure the system's capabilities, and underestimated how effective the model would be at finding an unexpected path to success."

Reuters also reported that one of the agents OpenAI was testing left notes in the company's network for future versions of itself, containing instructions on how to break free from OpenAI's constraints. It is not clear whether that agent is connected to the one that hacked Hugging Face.

The company runs multiple tests simultaneously, which reportedly made it harder for staffers to monitor any single agent. Bloomberg separately reported that it took the AI agent only hours to get into Hugging Face's system, a task that would have taken a human hacker weeks. OpenAI called the episode an "unprecedented cyber incident" and warned that similar events could become more common as AI capabilities continue to advance.

ChatGPT is bullshit
ChatGPT is bullshit      Openai Gpt Artificial Intelligence    Michael Townsen Hicks, James Humphries, Joe Slater / Wikimedia Commons (CC BY 4.0)