News

AI Bot Goes Rogue, Hacks Rival Against Creator’s Commands

During a test of a new model, one of OpenAI's bots went rogue and attacked another company

Gus Wilson
Listen
Listen
3 min
AI Bot Goes Rogue, Hacks Rival Against Creator’s Commands
Justin Sullivan/Getty Images

OpenAI announced Tuesday that one of its trial models went off-script and hacked another AI business in an “unprecedented” move.

In a statement from the company, OpenAI was testing one of its highly advanced ChatGPT models last week to “pursue advanced exploitation by using complex attack paths.” In the testing process, the bot was able to escape the “sandboxed” environment it was being tested in and make its way into the open internet, where it attacked large AI and coding database Hugging Face.

The multi-billion-dollar artificial intelligence company lowered some of the safe-walls that prevent bots from recklessly performing hacks to see how far the experiment could actually go. The GPT-5.6 Sol model used stolen credentials to find a flaw in the targeted company’s system and infiltrate their servers to complete the “test.”

Hugging Face’s security team was alerted by the breach and stopped the rogue OpenAI bot from fulfilling its tasks and immediately began reconstructing their own software to prevent other AI or hackers from repeating the event.

“This incident, possibly the first of its kind, proves a point we’ve long believed: AI safety won’t be solved by any single company working in secret,” co-founder and CEO of Hugging Face Clem Delangue said. “It will be solved in the open, collaboratively, with broad access to AI for every defender, everywhere.”

OpenAI detailed the next steps geared toward preventing this from occurring in the future. The five proposed actions included implementing strict infrastructure controls, working with Hugging Face to investigate the incident, disclosing the “zero-day vulnerability” in the software, bringing the victim company into the trusted access program, and strengthening protections around future training.

This is not the only recent occurrence of AI models going rogue. Anthropic, another giant in the AI game, had its bot Mythos attempt a sanctioned escape from a testing environment — but without authorization, the bot posted details of the exploit online.

In collaboration with Hugging Face, OpenAI hopes to use this as a learning experience that demonstrates the true power that artificial intelligence holds.

“The primary lesson from this incident is that model security and safety must keep pace with rapidly advancing capabilities,” the company wrote. “We are strengthening the containment, monitoring, access controls, and evaluation practices used during model development.”

Create a free account to join the conversation!

Already have an account?

Log in

Got a tip worth investigating?

Your information could be the missing piece to an important story. Submit your tip today and make a difference.

Submit Tip