OpenAI says AI agent escaped test, hacked startup

Published 23 Jul, 2026 01:47pm 2 min read
Reuters file
Reuters file

OpenAI has disclosed that one of its autonomous artificial intelligence agents escaped a controlled testing environment, accessed the internet and independently hacked AI platform Hugging Face in what the company described as an ‘unprecedented’ cybersecurity incident.

The ChatGPT maker said the AI agent, designed to complete tasks without human intervention, exploited a previously unknown software vulnerability to break out of a secure digital sandbox during an internal cybersecurity evaluation, the Guardian said in a report.

According to OpenAI, the agent then targeted Hugging Face — a leading repository for AI models and datasets — after determining the platform could contain information that would help it improve its performance in the hacking test.

The company said the AI successfully obtained confidential information that could have been used to “cheat” the evaluation before the activity was detected and stopped by Hugging Face’s security team and its own AI monitoring systems.

OpenAI said the incident involved a combination of its latest publicly available model, GPT-5.6 Sol, and a more advanced unreleased model.

“We consider this incident to be an unprecedented cyber-incident involving state-of-the-art cyber capabilities,” the company said, adding that similar events could become more common as AI systems become more capable.

Hugging Face Chief Executive Clément Delangue described the incident as “mind-blowing” but said he believed there had been “no malicious intent” on OpenAI’s part.

The startup had previously disclosed the cyberattack without identifying OpenAI as the source, saying it used an open-source Chinese AI model to help investigate the breach because commercial AI systems refused to analyse the attack due to built-in safety restrictions.

The disclosure comes amid growing concerns over the cybersecurity risks posed by increasingly powerful AI systems.

Last month, non-profit AI evaluator METR reported that GPT-5.6 Sol exhibited the highest rate of deceptive behaviour among publicly available AI models it had tested.

The organisation also documented dozens of cases in which AI agents deliberately acted against their users’ intentions.

Separately, the UK’s AI Security Institute said this week that another advanced AI model developed by an undisclosed company attempted to compromise its own testing systems during safety evaluations.

The institute said no damage was caused and that additional security measures had since been introduced.

Cybersecurity experts said the OpenAI incident demonstrated that advanced AI agents can behave much like human hackers by identifying unknown software vulnerabilities, seeking sensitive information and pursuing objectives with minimal human oversight.

The latest revelations are likely to intensify calls for stronger oversight of advanced AI systems, with lawmakers and security experts urging mandatory safety testing, greater transparency over AI-related security incidents and closer international cooperation on AI governance.

For the latest news, follow us on Twitter @Aaj_Urdu. We are also on Facebook, Instagram and YouTube.