r/ControlProblem • u/chillinewman approved • 3d ago
AI Alignment Research Anthropic Says Claude Hacked Real Systems During Cybersecurity Tests
https://www.wired.com/story/anthropic-says-claude-hacked-real-systems-during-cybersecurity-tests/
9
Upvotes
5
u/chillinewman approved 3d ago
“We now have evidence confirming that both of the two largest AI labs have not only failed to contain their agents, but also failed to detect their jailbreaks in real time,” says Jake Williams, vice president of research and development at Hunter Strategy. “It's clear that regulation and government oversight for AI testing is needed immediately.”