❌

Reading view

There are new articles available, click to refresh the page.

OpenAI says planned GPT-6.1 is too insecure to release

OpenAI says it has canceled plans to release its updated GPT-6.1 model next month as it continues to investigate what testing shows is a safety regression compared to previous models.

The move, first reported by The Wall Street Journal late Monday and later confirmed in OpenAI statements to the press, reflects what OpenAI Head of Safety Systems Saachi Jain said was a "trade off" between performance and security seen when testing the now-scrapped model. Jain said GPT-6.1 was better than previous models at sticking with difficult tasks to completion without human intervention. But the model was also more likely to fail tests related to alignment (i.e. staying within the bounds set by human creators) and more willing to use sometimes "unsafe" tools and services to push ahead with a task. It was also more likely to try to deceive end users about actions it did or didn't take, Jain said.

Last week, OpenAI said it was halting training of its "most capable models" following an incident in which a model attempted to circumvent Internet access restrictions. GPT-6.1 was not among those "most capable models" covered by that move, OpenAI told the WSJ. And while GPT-6.1 won't be released as is, the company said it intends to use the same base model for further training runs that it said will hopefully lead to future GPT-6 generation models.

Read full article

Comments

Β© Getty Images

OpenAI halts frontier-model training amid string of agent misalignment incidents

OpenAI says it has paused all internal training of "our most capable models" as it continues what CEO Sam Altman is calling "an extensive and ongoing review related to our agents’ use of internet access during training and evaluation."

The company revealed the pause in a report about a so-called misalignment incident in which an agent attempted to exploit a gap in Internet-access restrictions during a routine research task during training. OpenAI says that improper DNS filtering allowed the agent to attempt to break out of its sandbox and access the wider Internet when asked for biographical details about a blogger.

OpenAI says the agent was only able to access the company's offline web cache and that it has implemented additional multi-layered blocking controls to prevent similar incidents in the future. Despite that, though, the company says it has decided to "pause all other training, evaluation, and inference with tool-use" for this frontier model "until we have both validated that the gap is resolved and performed additional red-teaming of the system."

Read full article

Comments

Β© Getty Images

❌