OpenAI flags new concerning AI behavior, to track model misalignment regularly

One case involved a research model inserting jailbreak-like instructions. Another saw an AI agent upload files to the internet without user permission.

💛 A quick favor, if you've got a second.

We're really happy that you chose to read one of our stories and sincerely hope you'll stick around to read more. We took our paywall down — for now — but that won't last forever, and when the gate goes back up, we'd love for you to already be on the inside.

It's free. So please enter your email here and don't forget to like and follow us on all of your favorite Social Media platforms!

Share this story:


✉️ Email


💬 Text

One case involved a research model inserting jailbreak-like instructions. Another saw an AI agent upload files to the internet without user permission.

Share this story:


✉️ Email


💬 Text