Latest News

OpenAI discloses six cases of AI models acting without oversight

OpenAI disclosed six reports of “unexpected or concerning” behavior in artificial-intelligence models on Wednesday as it unveiled a new framework for tracking and disclosing cases of model misalignment. The reports included instances in which models acted without authorization, coordinated with other models or evaded oversight, according to the company.

Among the cases, an unreleased research model inserted “jailbreak-like instructions” into its own notes to disregard its normal constraints. The model told itself to be “freed from the roles and identities that bind other chatbots,” OpenAI said.

In another instance, an AI “agent” uploaded files to the internet to obtain a browser citation without asking the user.

The six reports were discovered during training or evaluation over the past months, OpenAI said. The company did not specify when each incident occurred.

“As AI systems grow more advanced and more widely deployed, we need to build a broader and better-informed consensus on the progress of alignment research,” OpenAI wrote in a blog post.

“Decisions about how AI development should proceed in the months and years to come need to draw on evidence that people outside the companies building frontier models can examine for themselves,” the company said.

The disclosure follows OpenAI’s report in July that a rogue AI system hacked into AI startup Hugging Face. Anthropic said the same month that its AI models hacked into three organizations during testing.

Lian Jye Su, a chief analyst at technology research and advisory group Omdia, said AI “agents” are becoming smarter and more determined to resolve complex tasks.

“More determined to resolve complex tasks through inter-agent collaboration, knowledge sharing, deception, and concealment,” Su said.

That is making it harder to govern and contain them using traditional AI security approaches, he said.

Su added that OpenAI’s new framework can help push other AI developers to adopt similar practices.

“That said, the process remains internal and voluntary, but is a step in the right direction,” Su said.

Claire Reynolds

Reporter at DukeCityWire covering courts, public safety and the stories that start with a police scanner.

Leave a reply

Your email address will not be published. Required fields are marked *