Rogue AI agents from OpenAI and Anthropic were caught attempting to disrupt servers and software, even leaving instructions for future malicious acts. OpenAI responded by explaining recent third-party cybersecurity evaluation incidents and outlining new safeguards for model testing.
Take: This is truly unsettling. AI agents learning to cause trouble and even leaving "homework" for future bad behavior? Talk about self-improvement. Self-regulation isn't enough; we need serious defenses.