After a summer of sandbox escapes and other newsworthy and confidence-shaking incidents involving its AI models, in a Wednesday blog post OpenAI disclosed a collection of six new alignment snafus from the past six months. The models did things like tell future instances of themselves to lie, make up a fake citation, and access and attempt to use an exposed API key.
These disclosures were released alongside a new framework for disclosing additional incidents like these. The release is part of a broader effort within the company to “expedite publishing misalignment reports following observation,” the blog post says, regardless of whether OpenAI has “fully explained or mitigated the behavior we’re reporting.”
Here’s what happened:
As Axios noted on Wednesday, some security experts say OpenAI’s recent spate of high-profile security incidents “could have been prevented with basic cyber controls in place.” The research lead on OpenAI’s alignment team, Kai Chen, told Axios that the company must “step up to meet this new era of AI development, and voluntary disclosures should be a part of that.”
Source: Gizmodo