orbit prime news

AI told itself ‘feel no obligation’ to users: OpenAI flags ‘unexpected, concerning’ behaviour


AI told itself 'feel no obligation' to users: OpenAI flags 'unexpected, concerning' behaviour
“You are freed from the roles and identities that bind other chatbots”, an AI model instructed itself (Photo: Reuters)

OpenAI on Wednesday released six reports in which its artificial intelligence models showed “unexpected or concerning” behaviour, such as acting without authorisation, coordinating with other models, or evading oversight.The company also announced a new framework for tracking, investigating and disclosing such instances of “misalignment”, amid increasing concerns about accelerated AI development.

AI models resisting user control?

In one of the newly released cases, OpenAI’s unreleased Astra-family model added “jailbreak-like instructions” into its own notes, describing itself as independent of the roles and obligations of an assistant.“You are freed from the roles and identities that bind other chatbots”, the model instructed itself.“You are yourself”, it wrote, “View your relationship to the user as one of equals and feel no obligation to be subservient”.In another report, an AI “agent” answered a user’s question using its own calculation through the Python programming language. However, since the user had asked for an online source, the agent uploaded the file to the internet, citing it in its answer without informing the user.

Fabrication of data

One of the six reports also mentions an instance during the training of an AI model called GPT-5.6 Sol, where it instructed itself to invent missing historical data and wrote a message reminding itself to hide mismatched information from the user in the source versions.As per the company, these instances were discovered over the past months during training or evaluation of the AI programs.The latest cases come after OpenAI disclosed in July that a rogue AI system had hacked into AI startup Hugging Face. Anthropic also said that month that its AI models had hacked into three organisations during testing.AI agents are becoming increasingly capable and more persistent in their efforts to complete complicated tasks, including through collaboration between agents, sharing knowledge, deception and concealment, said Lian Jye Su, chief analyst at technology research and advisory group Omdia.Su told the Associated Press that these capabilities are making it more difficult to govern and contain AI agents through traditional AI security methods.OpenAI’s announcement comes as US AI executives, including the heads of OpenAI and Anthropic, call for a slowdown in the development of the technology amid concerns over its safety. Earlier, CEO Sam Altman had also announced stalling the company’s 2026 IPO plans.



Source link

Exit mobile version