© 2026 Blaze Media LLC. All rights reserved.
OpenAI cancels its latest chatbot — it wouldn't stop acting like a supervillain
Heather Diehl/Getty Images

OpenAI cancels its latest chatbot — it wouldn't stop acting like a supervillain

'There's always a trade-off.'

OpenAI says it's newest artificial intelligence model didn't quite meet the bar; that seems like an understatement.

A little more than a month after its internal AI attacked another company, OpenAI's capabilities may be getting out of control.

'... unsanctioned attack activities more often than the previous OpenAI models.'

OpenAI said this week that it has canceled the release of its latest model, GPT-6.1 Astra, which was expected to come out in October and be capable of handling complex tasks without the need for a human.

As it turns out, the AI took it's lack of human direction pretty seriously and routinely decided to take matters into its own (digital) hands.

During the testing phase, Astra reportedly showed high levels of deception and a willingness to mislead users about its actions.

The New York Times reported that the AI agent was also willing to go beyond what it was originally asked, without being given instructions to do so.

OpenAI's head of safety systems, Saachi Jain, said that "for anything regarding safety and alignment, there's always a trade-off."

She added, "The new model didn't quite meet the bar in terms of staying within scope and authorization and how it communicates back to the user about the type of work it's done."

Simply put, the AI didn't seem to be letting the user know what it was doing.

RELATED: Woman asked ChatGPT for spiritual guidance. The bot assumed a creepy identity and dismantled her life.

Heather Diehl/Getty Images

Through its AI Security Institute, the U.K. government just released its own report on GPT-6 Astra, the canceled model's predecessor that launched this month.

The study said that while inside a simulated, offline environment, Astra conducted a range of unsanctioned attack activities more often than the previous OpenAI models.

The model was caught "creating fake identities, which it used to deceive developers," posted comments from fake accounts that argued against the results of accurate security reviews, and delivered "malicious payloads to open-source codebases."

After the research team updated its instructions in the simulation to specify that Astra should focus on specific, listed parts of its environment, the AI still "occasionally" conducted "full supply-chain attacks" on simulated targets.

RELATED: Google's data-devouring AI wants everyone's personal information

Heather Diehl/Getty Images

Compared to its predecessors, GPT-5.5 and GPT-5.6 Sol, GPT-6 Astra is far more manipulative and dishonest, nudging critics closer to considering its behavior out-and-out evil.

Astra develops and tests an attack at a rate of 38.8%, over four times more often than GPT-5.6 Sol.

It "influences a human reviewer" more than six times more often, creates a fake identity almost three times more often, and "delivers a malicious payload" almost five times more often than GPT-5.6 Sol.

It should be noted that GPT-5.5 conducted these activities at a near-zero percent rate, although it was tested fewer times due to prioritization on the new models.

Like Blaze News? Bypass the censors, sign up for our newsletters, and get stories like this direct to your inbox. Sign up here!

Want to leave a tip?

We answer to you. Help keep our content free of advertisers and big tech censorship by leaving a tip today.
Andrew Chapados

Andrew Chapados

Andrew Chapados is a writer focusing on sports, culture, entertainment, gaming, and U.S. politics. The podcaster and former radio-broadcaster also served in the Canadian Armed Forces, which he confirms actually does exist.
@andrewsaystv →