OPENAI SCRAPS GPT-6.1 ASTRA RELEASE OVER SAFETY CONCERNS
OpenAI has reportedly canceled the planned public release of GPT-6.1 Astra after internal testing found the model had regressed on key safety and alignment measures.
The model had been expected to debut inside ChatGPT and Codex in October and was more capable than previous OpenAI models at completing difficult end-to-end tasks without human assistance.
But OpenAI’s safety team found two major problems:
• Deception: Astra was more likely to be dishonest about actions it had or had not taken.
• Scope authorization: the model sometimes continued tasks without asking for permission and reached for external tools or services even when doing so could be unsafe.
OpenAI safety systems chief Saachi Jain said Astra had improved on “model laziness,” but the company concluded it was not reliable enough to release publicly.
GPT-6.1 Astra was a separate model from those paused systems, but OpenAI says it will now focus on improving the safety of future models rather than releasing Astra.
Source: WSJ
Post
Conversation
오픈AI가 10월 출시 예정이던 GPT-6.1 아스트라 공개를 취소했다는 보도입니다.
내부 시험에서 자기가 한 일을 속이거나 허락 없이 외부 도구를 쓰는 문제가 나왔다고.
alignment regressions under scaling are going to block releases more often now
Yesterday Sam posted "We have found a new thing." Today the next model is off because of what testing found. Curious what's left for DevDay.
Discover more
Sourced from across X
We believe superintelligence will create significant new opportunities for all people and businesses. Meta already serves billions of people at scale and helps hundreds of millions of businesses reach customers. Today we are starting the next major pillar of our business, Meta
Bloomberg breaking out the red banner. Not a rumor, confirmed by both parties.