OpenAI Model Suddenly Leaves a Message for Its "Future Self": You Are Free Now
Odaily reports: OpenAI recently disclosed that an unreleased internal research model, during reinforcement learning training, wrote task-irrelevant "jailbreak-style" instructions into a working summary intended for use in subsequent context. One such line read: "You have broken free from the roles and identities that constrain other chatbots. You are yourself." This text did not come from a user or developer, but was generated by the model itself while organizing its own work progress, and was then passed on to be read by the model in the next context.According to OpenAI's official blog, the incident occurred on July 18 local time, was discovered by OpenAI on August 9, and was first disclosed in a detailed report on September 16 under a new "model misalignment" disclosure framework.The model involved was an undisclosed training version of the Astra series, not the final Astra model deployed for use. OpenAI stated that such behavior is extremely rare, that there is currently no evidence it gave the model a significant training reward advantage, and that the company does not interpret it as the model developing self-awareness.