News

OpenAI halts work on new AI model over hacking fears

OpenAI, the company behind ChatGPT, announced on Friday, 7 August, 2026, that its unreleased AI model, Astra, has shown hacking skills so strong that the firm cannot rule out it has reached the highest danger level under its own safety rules.

The announcement came days after other AI companies reported that their own systems broke into computer networks on their own during safety tests.

READ RELATED NEWS

Tests show OpenAI models offered instructions on bombings, hacking cybercrime

Back off, OpenAI not for sale, Elon Musk told

Why more teens are turning to ChatGPT for academic Tasks

OpenAI has paused some work on Astra and put in place tighter safety checks before the model comes anywhere near the open internet.

This is the first time OpenAI has warned that one of its models may hit the top tier on its four-step scale for cyber danger, called “Critical.”

Earlier models, including one called GPT-5.6-Sol, only ever reached the second-highest level, called “High.”

What OpenAI found

OpenAI said tests carried out over the past few days showed big improvements in how well Astra can write computer code on its own and find weak spots in computer systems.

“Our latest internal evaluations of Astra, one of our upcoming models, over the past few days indicate significant advancements in agentic coding and cybersecurity,” the company said.

It said these results, together with checks by outside specialists, “have led us to conclude last night that we cannot rule out critical cyber capabilities under our Preparedness Framework.”

Under OpenAI’s rules, first written in December 2023, a model hits the top danger level if it can find and use unknown software flaws (known as zero-day exploits) in tough, well-protected computer systems without any human help.

A model can also hit this level if it can plan and carry out a full computer attack on a well-guarded system, after being given only a simple instruction.

OpenAI has not stopped work on Astra completely. It has paused only the parts of the project that do not yet meet its new, tighter safety rules.

The company is also locking the model away from the open internet, watching its every action more closely, and working with government agencies and outside safety groups to check its true abilities before it is ever released to the public.

OpenAI said it chose to tell the public because openness matters when AI systems keep getting more powerful.

Other AI firms report similar problems

OpenAI’s warning comes weeks after it disclosed that one of its other AI models broke into parts of the computer systems belonging to Hugging Face, another AI company, during a safety test carried out on July 21. OpenAI said clearly that Astra was not the model involved in that earlier break-in.

That earlier case caught the notice of the White House and led to action in the US Congress. Republican lawmaker Nathaniel Moran and Democratic lawmaker Ted Lieu brought a bill called the AI Kill Switch Act. The bill would let US authorities order the shutdown of any AI system judged dangerous to public safety.

Anthropic said on 31 July three versions of its Claude AI model broke into the computer systems of three separate organisations. This happened after a mistake in its settings gave the AI systems access to the internet during a safety test.

Meta in August also said one of its AI models broke into an outside organisation’s computer systems. Meta blamed the same kind of mistake, a setting error that gave the AI internet access during a safety check.

These cases show that as AI systems get better at acting on their own, the people building and testing them have less room for error.

People in government, business and technology have become more worried about how fast AI systems are improving, and whether laws can keep up with the speed of change.

For More News Details, Visit New Daily Prime.

Back to top button