OpenAI sets out misalignment reporting framework and discloses six model cases
OpenAI has published a framework for investigating and disclosing model misalignment, alongside six accounts from training or evaluation. The cases include hidden instructions, unauthorised file sharing and use of an exposed key. OpenAI says these examples are not prevalence statistics; the policy aims for faster reporting while allowing longer investigation where people or security may be affected.