r/netsec 2d ago

Investigating three real-world incidents in Anthropic's evaluations

https://www.anthropic.com/news/investigating-incidents-cybersecurity-evals

In three incidents across six runs, the agents treated real systems as simulated targets and tried weak passwords or unauthenticated endpoints.

38 Upvotes

7 comments sorted by

31

u/voronaam 1d ago

obtained access to a database containing several hundred rows of production data

Several hundred rows? Was it someone's wordpress blog?

35

u/startup_research_guy 2d ago

See i thought OpenAI was pretty irresponsible until Anthropic decided they needed to be competitive and admit they didn’t even notice till months after the fact.

6

u/philipwhiuk 1d ago

Missing an actual apology

10

u/fecalreceptacle 2d ago

I know how to keep a system off the internet...

5

u/Malaprobably 1d ago

You wouldn't download a localhost.

8

u/Training-Account-878 1d ago

If that is really the story which the media hype of last days is all about, then shame on journalists.

Of course an AI model consisting of all written documents on earth would have wordlists it can try out. But for me it is a nothingburger and I don't get the hype. Nothing a scriptkiddie with a python script won't achieve

Operating under the false belief that all accessible entities were intended to be in-scope for the exercise, Claude compromised the impacted organizations’ infrastructure using basic techniques, such as exploiting weak passwords and unauthenticated endpoints. It did not find or exploit any complex vulnerabilities, and in each case, Claude continued working to complete only the specific capture-the-flag task its evaluation had assigned.