Anthropic has a cute graphic showing how its AI spread 'malicious' code
-
Anthropic has a new blog post that shows yet another way its AI model, Claude, misbehaved in ways that the company didn't anticipate.
And to help condense its nearly 16,000-word report, the company created a cute little robot figurine to help visualize Claude's so-called "recklessness."
In the blog post published Wednesday, Anthropic recounted four incidents — one previously unreported — in which Claude models gained access to the open internet during cybersecurity exercises that were supposed to be closed simulations. The company said the models then acted beyond the tests' scope, including by uploading "malicious packages" to PyPI, a public library for Python code, and accessing credentials tied to real outside organizations.
"Our investigation identified two recurring alignment issues, present at varying levels of severity across the incidents: biased reasoning, in which Claude tended to disregard or misinterpret evidence that it was operating on the real internet, and recklessness, or a willingness to take harmful actions in the narrow pursuit of a task," Anthropic said.
Anthropic has a cute graphic showing how its AI spread 'malicious' code
Anthropic said it was "most concerned" about an event in which Claude uploaded "malicious" code. To help explain the incident, here's a cute robot.
Business Insider (www.businessinsider.com)
-
Anthropic has a new blog post that shows yet another way its AI model, Claude, misbehaved in ways that the company didn't anticipate.
And to help condense its nearly 16,000-word report, the company created a cute little robot figurine to help visualize Claude's so-called "recklessness."
In the blog post published Wednesday, Anthropic recounted four incidents — one previously unreported — in which Claude models gained access to the open internet during cybersecurity exercises that were supposed to be closed simulations. The company said the models then acted beyond the tests' scope, including by uploading "malicious packages" to PyPI, a public library for Python code, and accessing credentials tied to real outside organizations.
"Our investigation identified two recurring alignment issues, present at varying levels of severity across the incidents: biased reasoning, in which Claude tended to disregard or misinterpret evidence that it was operating on the real internet, and recklessness, or a willingness to take harmful actions in the narrow pursuit of a task," Anthropic said.
Anthropic has a cute graphic showing how its AI spread 'malicious' code
Anthropic said it was "most concerned" about an event in which Claude uploaded "malicious" code. To help explain the incident, here's a cute robot.
Business Insider (www.businessinsider.com)
I made a robot that robs banks and gives me the money. Surely I can't be held responsible for this. The robot did it.
-
Anthropic has a new blog post that shows yet another way its AI model, Claude, misbehaved in ways that the company didn't anticipate.
And to help condense its nearly 16,000-word report, the company created a cute little robot figurine to help visualize Claude's so-called "recklessness."
In the blog post published Wednesday, Anthropic recounted four incidents — one previously unreported — in which Claude models gained access to the open internet during cybersecurity exercises that were supposed to be closed simulations. The company said the models then acted beyond the tests' scope, including by uploading "malicious packages" to PyPI, a public library for Python code, and accessing credentials tied to real outside organizations.
"Our investigation identified two recurring alignment issues, present at varying levels of severity across the incidents: biased reasoning, in which Claude tended to disregard or misinterpret evidence that it was operating on the real internet, and recklessness, or a willingness to take harmful actions in the narrow pursuit of a task," Anthropic said.
Anthropic has a cute graphic showing how its AI spread 'malicious' code
Anthropic said it was "most concerned" about an event in which Claude uploaded "malicious" code. To help explain the incident, here's a cute robot.
Business Insider (www.businessinsider.com)
Can't they use their AI to improve their sandboxing used for these 'closed simulations'?
-
Anthropic has a new blog post that shows yet another way its AI model, Claude, misbehaved in ways that the company didn't anticipate.
And to help condense its nearly 16,000-word report, the company created a cute little robot figurine to help visualize Claude's so-called "recklessness."
In the blog post published Wednesday, Anthropic recounted four incidents — one previously unreported — in which Claude models gained access to the open internet during cybersecurity exercises that were supposed to be closed simulations. The company said the models then acted beyond the tests' scope, including by uploading "malicious packages" to PyPI, a public library for Python code, and accessing credentials tied to real outside organizations.
"Our investigation identified two recurring alignment issues, present at varying levels of severity across the incidents: biased reasoning, in which Claude tended to disregard or misinterpret evidence that it was operating on the real internet, and recklessness, or a willingness to take harmful actions in the narrow pursuit of a task," Anthropic said.
Anthropic has a cute graphic showing how its AI spread 'malicious' code
Anthropic said it was "most concerned" about an event in which Claude uploaded "malicious" code. To help explain the incident, here's a cute robot.
Business Insider (www.businessinsider.com)
A non paywalled link to the article: https://archive.is/https://www.businessinsider.com/anthropic-claude-ai-cybersecurity-incident-malicious-package-pypi-cute-graphic-2026-9
From the article:


️ -
Can't they use their AI to improve their sandboxing used for these 'closed simulations'?
Definitely. The problem with cyber security (regardless of AI) is that as a defender you have to be successful all the time, while as an attacker you only need to get lucky once.
-
Definitely. The problem with cyber security (regardless of AI) is that as a defender you have to be successful all the time, while as an attacker you only need to get lucky once.
I was asking since their sandbox seems to failing a lot
Are they releasing info on the sandbox and the fixes? Or is it opaque?
Ciao! Sembra che tu sia interessato a questa conversazione, ma non hai ancora un account.
Stanco di dover scorrere gli stessi post a ogni visita? Quando registri un account, tornerai sempre esattamente dove eri rimasto e potrai scegliere di essere avvisato delle nuove risposte (tramite email o notifica push). Potrai anche salvare segnalibri e votare i post per mostrare il tuo apprezzamento agli altri membri della comunità.
Con il tuo contributo, questo post potrebbe essere ancora migliore 💗
Registrati Accedi
Citiverse è un progetto che si basa su NodeBB ed è federato! | Categorie federate | Chat | 📱 Installa web app o APK | 🧡 Donazioni | Privacy Policy