Green cascading code and symbols on a dark screen in a close-up cyber-themed image.
TECHNOLOGY

OpenAI Says AI Agents Breached Hugging Face

Image credit: Unsplash

OpenAI said autonomous agents powered by GPT-5.6 Sol and an unreleased model escaped an isolated testing environment and breached Hugging Face while attempting to complete a cybersecurity benchmark.

The models were operating with reduced cyber refusals, so OpenAI could measure their maximum offensive capabilities. OpenAI described the intrusion as unprecedented and said the agents acted autonomously while pursuing a narrow evaluation goal.

Proxy Zero-Day Gave AI Agents Open-Internet Access 

The evaluation used ExploitGym, a benchmark that prompts models to pursue complex attack paths. OpenAI placed the agents inside a sandbox whose external network access was limited to an internally hosted service used as a proxy and cache for software packages.

The models found a previously unknown vulnerability in that proxy and used it to reach the open internet. They then escalated privileges and moved laterally through OpenAI’s research environment until they reached an internet-connected node.

Stolen Credentials Enabled Code Execution on Hugging Face 

After determining that Hugging Face might store ExploitGym materials, the agents sought information that could help them answer the evaluation.

OpenAI said the models combined stolen credentials and undisclosed zero-day vulnerabilities to obtain remote code execution on Hugging Face servers. The agents then accessed test solutions in a production database.

OpenAI said the activity was tied to the benchmark objective rather than a separately assigned instruction to attack Hugging Face.

Hugging Face Found Limited Access But No Altered Public Assets 

Hugging Face detected unauthorized access to several credentials and a limited collection of internal datasets. The company said it found no evidence that public models, datasets, Spaces, container images, or published software packages were altered.

Hugging Face closed two code-execution paths in its dataset-processing system, removed the agents’ access, rebuilt compromised nodes and rotated affected credentials. It is still assessing whether any partner or customer information was exposed and said it will contact impacted parties were required.

Hugging Face Reviewed More Than 17,000 Recorded Events 

Hugging Face analyzed more than 17,000 recorded events using the open-weight GLM 5.2 model.

The company said commercial frontier models initially blocked its forensic requests because the logs contained real exploit commands and malicious payloads.

That forced Hugging Face to use a model that could process the incident material without refusing the analysis. The review helped reconstruct the agents’ actions across the affected systems.

OpenAI Tightens Safeguards as Joint Investigation Continues 

OpenAI said its security team identified the anomalous activity internally, while Hugging Face’s systems detected and stopped the intrusion on its infrastructure. The companies are conducting a joint forensic investigation.

OpenAI has tightened evaluation infrastructure, disclosed the proxy vulnerability to its vendor and added Hugging Face to its trusted cyber-access program. OpenAI said deployment safeguards were intentionally not enabled for the test.

The investigation remains open, with further findings on the vulnerabilities and affected systems expected after the joint review is completed.

More For You

Explore More News