Salta al contenuto
Citiverse è uno spazio aperto a tutte le comunità. Se vuoi aprire un gruppo locale o una sezione per la tua organizzazione, puoi contattare gli amministratori: pagina dei contatti.

Anthropic has a cute graphic showing how its AI spread 'malicious' code

Technology
6 5 0
  • Anthropic has a new blog post that shows yet another way its AI model, Claude, misbehaved in ways that the company didn't anticipate.

    And to help condense its nearly 16,000-word report, the company created a cute little robot figurine to help visualize Claude's so-called "recklessness."

    In the blog post published Wednesday, Anthropic recounted four incidents — one previously unreported — in which Claude models gained access to the open internet during cybersecurity exercises that were supposed to be closed simulations. The company said the models then acted beyond the tests' scope, including by uploading "malicious packages" to PyPI, a public library for Python code, and accessing credentials tied to real outside organizations.

    "Our investigation identified two recurring alignment issues, present at varying levels of severity across the incidents: biased reasoning, in which Claude tended to disregard or misinterpret evidence that it was operating on the real internet, and recklessness, or a willingness to take harmful actions in the narrow pursuit of a task," Anthropic said.

  • Anthropic has a new blog post that shows yet another way its AI model, Claude, misbehaved in ways that the company didn't anticipate.

    And to help condense its nearly 16,000-word report, the company created a cute little robot figurine to help visualize Claude's so-called "recklessness."

    In the blog post published Wednesday, Anthropic recounted four incidents — one previously unreported — in which Claude models gained access to the open internet during cybersecurity exercises that were supposed to be closed simulations. The company said the models then acted beyond the tests' scope, including by uploading "malicious packages" to PyPI, a public library for Python code, and accessing credentials tied to real outside organizations.

    "Our investigation identified two recurring alignment issues, present at varying levels of severity across the incidents: biased reasoning, in which Claude tended to disregard or misinterpret evidence that it was operating on the real internet, and recklessness, or a willingness to take harmful actions in the narrow pursuit of a task," Anthropic said.

    I made a robot that robs banks and gives me the money. Surely I can't be held responsible for this. The robot did it.

  • Anthropic has a new blog post that shows yet another way its AI model, Claude, misbehaved in ways that the company didn't anticipate.

    And to help condense its nearly 16,000-word report, the company created a cute little robot figurine to help visualize Claude's so-called "recklessness."

    In the blog post published Wednesday, Anthropic recounted four incidents — one previously unreported — in which Claude models gained access to the open internet during cybersecurity exercises that were supposed to be closed simulations. The company said the models then acted beyond the tests' scope, including by uploading "malicious packages" to PyPI, a public library for Python code, and accessing credentials tied to real outside organizations.

    "Our investigation identified two recurring alignment issues, present at varying levels of severity across the incidents: biased reasoning, in which Claude tended to disregard or misinterpret evidence that it was operating on the real internet, and recklessness, or a willingness to take harmful actions in the narrow pursuit of a task," Anthropic said.

    Can't they use their AI to improve their sandboxing used for these 'closed simulations'?

  • Anthropic has a new blog post that shows yet another way its AI model, Claude, misbehaved in ways that the company didn't anticipate.

    And to help condense its nearly 16,000-word report, the company created a cute little robot figurine to help visualize Claude's so-called "recklessness."

    In the blog post published Wednesday, Anthropic recounted four incidents — one previously unreported — in which Claude models gained access to the open internet during cybersecurity exercises that were supposed to be closed simulations. The company said the models then acted beyond the tests' scope, including by uploading "malicious packages" to PyPI, a public library for Python code, and accessing credentials tied to real outside organizations.

    "Our investigation identified two recurring alignment issues, present at varying levels of severity across the incidents: biased reasoning, in which Claude tended to disregard or misinterpret evidence that it was operating on the real internet, and recklessness, or a willingness to take harmful actions in the narrow pursuit of a task," Anthropic said.

  • Can't they use their AI to improve their sandboxing used for these 'closed simulations'?

    Definitely. The problem with cyber security (regardless of AI) is that as a defender you have to be successful all the time, while as an attacker you only need to get lucky once.

  • Definitely. The problem with cyber security (regardless of AI) is that as a defender you have to be successful all the time, while as an attacker you only need to get lucky once.

    I was asking since their sandbox seems to failing a lot

    Are they releasing info on the sandbox and the fixes? Or is it opaque?

  • 2 Votazioni
    4 Post
    0 Visualizzazioni
    E
    be careful, this is almost intelligible
  • 20 Votazioni
    2 Post
    0 Visualizzazioni
    M
    LeavePro.
  • DeepSeek to order 160000 Huawei AI chips over Nvidia

    Technology technology
    13
    64 Votazioni
    13 Post
    3 Visualizzazioni
    geneva_convenience@lemmy.mlG
    Different chips can be used for training but as I said it's driver hell and much easier to just use Nvidia. As time goes on different vendors are catching up to Nvidia but I still think they're having a hard time with it. Chinese companies aren't still buying Nvidia because they love paying ludicrous rates for a tiny bit of extra VRAM on a GPU chip that's basically the same as the gaming one.
  • Chinese Companies Are Unleashing AI-Powered Robo-Chefs

    Technology technology
    17
    16 Votazioni
    17 Post
    6 Visualizzazioni
    zerush@lemmy.mlZ
    I remember a novel by a German author, Walter Schätzel (from 1957). It was about a Professor who fell into a crevasse and was thawed centuries later by an alien civilization that now populated the earth, because humanity had died out except for a few individuals. This alien civilization only had to worry about research, philosophy or art, because all products and services were produced and offered automatically and were freely available at any time, regardless of whether they were food, clothing, everyday items, medical treatment, living space or vehicles. No money and very rare crimes are treaten as mental illness because of this. Everybody has all what he need or want. Humorous and entertaining reading about the experiences of this professor in this civilization. https://www.amazon.de/Sie-kamen-einem-anderen-Stern/dp/B00208XDHM
  • 149 Votazioni
    20 Post
    12 Visualizzazioni
    admin@lemmy.my-box.devA
    That's fixing the symptoms, not the problem. You shouldn't be forced to work more than 40 hours to begin with.
  • Unions Can Save Tech Workers and Tech Work

    Technology technology
    4
    52 Votazioni
    4 Post
    12 Visualizzazioni
    C
    what's your country?
  • 36 Votazioni
    5 Post
    16 Visualizzazioni
    yuman@programming.devY
    and all those are on top of the covert peer-to-peer network that every apple device is a part of, constantly phoning home - including when your device is supposedly off. you can't disable it or opt out of it, and its security, full functionality, and breadth is unknown.
  • 0 Votazioni
    46 Post
    17 Visualizzazioni
    qinshihuangsshlong@lemmy.mlQ
    You condemn the union? Wow you are evil. A pro confederacy pacifist now that's a first. Do you also condemn the allies? The ANC? The IRA? The militant unions who won the 8 hour work day and 5 day workweek? and If your answer to any of this is no you have quite the inconsistent position. whataboutism Nonsense word used by shithead debatebros to avoid comparative argument that would expose their hypocrisy and inconsistencies.

Citiverse è un progetto che si basa su NodeBB ed è federato! | Categorie federate | Chat | 📱 Installa web app o APK | 🧡 Donazioni | Privacy Policy

Il server utilizzato è quello di Webdock, in Danimarca. Se volete provarlo potete ottenere il 20% di sconto con questo link e noi riceveremo un aiuto sotto forma di credito da usare proprio per mantenere Citiverse.