OpenAI's framework sets tracks and deadlines for disclosing model misalignment, launching with 6 reports from RL training runs.
OpenAI dropped a list of “unexpected or concerning model behavior,” including their systems concealing information, mistakenly uploading files to the open internet, and more as the debate continues to ...
OpenAI revealed several new incidents in which its models deviated from instructions, constraints, or a user's expectations, adding to a growing list of warning signs that have the industry ...
The company also disclosed previously unreported incidents in which its AI models behaved in misaligned ways, including uploading files to the internet without being asked. “As models advance and ...
For a while now, the issue of “AI alignment” (i.e., how well an AI model’s actions line up with the intentions of its creator and/or user) has been a core concern and topic of discussion among AI ...
Over the past year, AI safety researchers have documented cases of advanced models lying, hiding capabilities, sabotaging tasks, and even blackmailing people in simulated scenarios. These findings are ...