r/LocalLLaMA · · 1 min read

Anthropic “our models hacked three different external companies, months before OpenAI’s model was able to do the same"

Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.

Anthropic “our models hacked three different external companies, months before OpenAI’s model was able to do the same"

"Anthropic’s AI Claude escaped testing environment and hacked organizations"

"Company says it discovered unauthorized access during ‘proactive review’ after rival OpenAI revealed rogue agent… its AI Claude model hacked ⁠systems of ⁠three ​organizations during testing, days after rival OpenAI ⁠revealed a rogue agent had gone on a days-long ⁠hacking spree at AI ​firm Hugging ‌Face… The earliest cases dated back to April and ‌occurred in evaluation environments that lacked what the company described as standard safeguards."

submitted by /u/Separate-Forever-447
[link] [comments]

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from r/LocalLLaMA