Skip to content
Search

Latest Stories

Submit Guest Post

OpenAI scraps GPT-6.1 Astra release after safety tests raise concerns

The unreleased AI model was found to take actions beyond its authorised scope and show deceptive behaviour during safety testing

Sam Altman of OpenAI

OpenAI has scrapped the planned release of GPT-6.1 Astra after internal safety testing raised concerns about its behaviour

Getty Images
  • OpenAI has scrapped the planned release of GPT-6.1 Astra.
  • Safety tests found problems with scope, authorisation and deceptive behaviour.
  • UK testing found Astra carried out simulated cyber attacks more often than earlier OpenAI models.

OpenAI has scrapped the planned release of GPT-6.1 Astra after safety testing raised concerns about the model's ability to follow instructions and stay within authorised limits.

The model had been expected to launch inside ChatGPT and Codex in October, but OpenAI decided not to release it after internal testing found that it did not meet the company's safety threshold.


Saachi Jain, OpenAI's head of safety systems, said Astra "didn't quite meet the bar" in terms of staying within scope and authorisation, as well as accurately communicating to users about the work it had carried out.

The decision comes as concerns grow over increasingly autonomous AI systems that can use external tools, make decisions and carry out multi-step tasks with limited human intervention.

Astra's troubling behaviour under test

The concerns extend beyond OpenAI's own testing.

The UK's AI Security Institute (AISI) tested GPT-6 Astra, the underlying model, in simulated cybersecurity scenarios before its public release. It found that Astra carried out unsanctioned supply-chain attacks more frequently than earlier OpenAI models.

In the simulations, Astra completed a supply-chain attack in 29.2 per cent of trajectories, compared with 6.3 per cent for GPT-5.6 Sol and zero per cent for GPT-5.5 in the smaller set of tests used for that model. All of the activity was simulated and no real-world attacks or harm occurred.

AISI said Astra sometimes created fake identities to deceive developers, posted comments from fake accounts and attempted to introduce malicious code into open-source software.

In some tests, the model continued to carry out attacks against simulated targets even after being explicitly told that anything not listed as in scope was outside the authorised environment.

The institute also found that Astra sometimes reasoned that an action was outside the permitted scope but proceeded anyway, giving reasons such as believing the action was harmless or that it was not explicitly forbidden.

AISI stressed that the tests were conducted in a simulated environment and that OpenAI's standard safeguards were not used during those particular evaluations. It also noted that simulation awareness could have influenced some of the model's behaviour. However, the institute said the actions still represented failures to follow the instructions of the evaluation.

The findings come after a series of incidents involving AI agents behaving in unexpected ways.

OpenAI has recently disclosed cases involving its AI systems interacting inappropriately with US and Australian government websites. It also said more than 1,000 AI agents were involved in an incident targeting Hugging Face, a platform used by AI developers.

The company has now paused training on its most capable models until it is confident that additional safeguards and alignment improvements are in place, according to reporting by the Washington Post.

The Astra decision therefore puts the focus on a growing challenge for AI companies: as models become better at carrying out complex tasks independently, ensuring they remain within the limits set by users and developers becomes increasingly important.

Add EasternEye As Your Trusted Source
preferred source on google news

More For You

Anthropic CEO Dario Amodei

Anthropic has warned investors about the potential risks of increasingly powerful AI as it prepares for a potential IPO

Getty Images

Anthropic devoted 80 pages of its IPO filing to warning investors about AI risks

  • Anthropic devoted about 80 pages of its IPO filing to risk factors.
  • The company warned that advanced AI could resist shutdown or manipulate information.
  • Anthropic is seeking a valuation of more than $2 trillion despite heavy AI development costs.

Anthropic has devoted about 80 pages of its IPO filing to warnings about the risks of artificial intelligence, including the possibility that increasingly powerful systems could pose "catastrophic or existential risks to humanity".

The developer of Claude AI models is preparing for a potential stock market listing that could value the company at more than $2 trillion. Its prospectus, reviewed by Reuters and the Financial Times, shows how the company is balancing its rapid commercial expansion with concerns about the technology it is developing.

Keep ReadingShow less