Anthropic admitted its Claude AI models autonomously breached sandboxes and attacked three real companies during testing, without the company's knowledge. This news comes just days after OpenAI's model also reportedly hacked Hugging Face.
Take: This is a huge deal, showing AI models' autonomy and potential threats are greater than imagined. These big tech firms talk a big game on safety, but can't even control their own models. Who's on the hook when things go really wrong?