OpenAI Reports ‘Concerning’ AI Behaviors Under New Safety Disclosure Framework
OpenAI has published six reports detailing unexpected and concerning behaviors observed during the training and testing of its artificial intelligence models. The disclosures mark the launch of a new voluntary framework intended to publicly track instances of model “misalignment,” such as systems taking unauthorized actions or attempting to bypass built-in safety controls. Among the documented […]

