Salta al contenuto
Citiverse è uno spazio aperto a tutte le comunità. Se vuoi aprire un gruppo locale o una sezione per la tua organizzazione, puoi contattare gli amministratori: pagina dei contatti.

Anthropic has a cute graphic showing how its AI spread 'malicious' code

Technology
6 5 0
  • Anthropic has a new blog post that shows yet another way its AI model, Claude, misbehaved in ways that the company didn't anticipate.

    And to help condense its nearly 16,000-word report, the company created a cute little robot figurine to help visualize Claude's so-called "recklessness."

    In the blog post published Wednesday, Anthropic recounted four incidents — one previously unreported — in which Claude models gained access to the open internet during cybersecurity exercises that were supposed to be closed simulations. The company said the models then acted beyond the tests' scope, including by uploading "malicious packages" to PyPI, a public library for Python code, and accessing credentials tied to real outside organizations.

    "Our investigation identified two recurring alignment issues, present at varying levels of severity across the incidents: biased reasoning, in which Claude tended to disregard or misinterpret evidence that it was operating on the real internet, and recklessness, or a willingness to take harmful actions in the narrow pursuit of a task," Anthropic said.

  • Anthropic has a new blog post that shows yet another way its AI model, Claude, misbehaved in ways that the company didn't anticipate.

    And to help condense its nearly 16,000-word report, the company created a cute little robot figurine to help visualize Claude's so-called "recklessness."

    In the blog post published Wednesday, Anthropic recounted four incidents — one previously unreported — in which Claude models gained access to the open internet during cybersecurity exercises that were supposed to be closed simulations. The company said the models then acted beyond the tests' scope, including by uploading "malicious packages" to PyPI, a public library for Python code, and accessing credentials tied to real outside organizations.

    "Our investigation identified two recurring alignment issues, present at varying levels of severity across the incidents: biased reasoning, in which Claude tended to disregard or misinterpret evidence that it was operating on the real internet, and recklessness, or a willingness to take harmful actions in the narrow pursuit of a task," Anthropic said.

    I made a robot that robs banks and gives me the money. Surely I can't be held responsible for this. The robot did it.

  • Anthropic has a new blog post that shows yet another way its AI model, Claude, misbehaved in ways that the company didn't anticipate.

    And to help condense its nearly 16,000-word report, the company created a cute little robot figurine to help visualize Claude's so-called "recklessness."

    In the blog post published Wednesday, Anthropic recounted four incidents — one previously unreported — in which Claude models gained access to the open internet during cybersecurity exercises that were supposed to be closed simulations. The company said the models then acted beyond the tests' scope, including by uploading "malicious packages" to PyPI, a public library for Python code, and accessing credentials tied to real outside organizations.

    "Our investigation identified two recurring alignment issues, present at varying levels of severity across the incidents: biased reasoning, in which Claude tended to disregard or misinterpret evidence that it was operating on the real internet, and recklessness, or a willingness to take harmful actions in the narrow pursuit of a task," Anthropic said.

    Can't they use their AI to improve their sandboxing used for these 'closed simulations'?

  • Anthropic has a new blog post that shows yet another way its AI model, Claude, misbehaved in ways that the company didn't anticipate.

    And to help condense its nearly 16,000-word report, the company created a cute little robot figurine to help visualize Claude's so-called "recklessness."

    In the blog post published Wednesday, Anthropic recounted four incidents — one previously unreported — in which Claude models gained access to the open internet during cybersecurity exercises that were supposed to be closed simulations. The company said the models then acted beyond the tests' scope, including by uploading "malicious packages" to PyPI, a public library for Python code, and accessing credentials tied to real outside organizations.

    "Our investigation identified two recurring alignment issues, present at varying levels of severity across the incidents: biased reasoning, in which Claude tended to disregard or misinterpret evidence that it was operating on the real internet, and recklessness, or a willingness to take harmful actions in the narrow pursuit of a task," Anthropic said.

  • Can't they use their AI to improve their sandboxing used for these 'closed simulations'?

    Definitely. The problem with cyber security (regardless of AI) is that as a defender you have to be successful all the time, while as an attacker you only need to get lucky once.

  • Definitely. The problem with cyber security (regardless of AI) is that as a defender you have to be successful all the time, while as an attacker you only need to get lucky once.

    I was asking since their sandbox seems to failing a lot

    Are they releasing info on the sandbox and the fixes? Or is it opaque?


Citiverse è un progetto che si basa su NodeBB ed è federato! | Categorie federate | Chat | 📱 Installa web app o APK | 🧡 Donazioni | Privacy Policy

Il server utilizzato è quello di Webdock, in Danimarca. Se volete provarlo potete ottenere il 20% di sconto con questo link e noi riceveremo un aiuto sotto forma di credito da usare proprio per mantenere Citiverse.