<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0"><channel><title><![CDATA[Anthropic has a cute graphic showing how its AI spread 'malicious' code]]></title><description><![CDATA[<p dir="auto">Anthropic has a new blog post that shows yet another way its AI model, Claude, misbehaved in ways that the company didn't anticipate.</p>
<p dir="auto">And to help condense its nearly 16,000-word report, the company created a cute little robot figurine to help visualize Claude's so-called &quot;recklessness.&quot;</p>
<p dir="auto">In the blog post published Wednesday, Anthropic recounted four incidents — one previously unreported — in which Claude models gained access to the open internet during cybersecurity exercises that were supposed to be closed simulations. The company said the models then acted beyond the tests' scope, including by uploading &quot;malicious packages&quot; to PyPI, a public library for Python code, and accessing credentials tied to real outside organizations.</p>
<p dir="auto">&quot;Our investigation identified two recurring alignment issues, present at varying levels of severity across the incidents: biased reasoning, in which Claude tended to disregard or misinterpret evidence that it was operating on the real internet, and recklessness, or a willingness to take harmful actions in the narrow pursuit of a task,&quot; Anthropic said.</p>


<div class="row mt-3"><div class="card col-md-9 col-lg-6 position-relative link-preview p-0">



<a href="https://www.businessinsider.com/anthropic-claude-ai-cybersecurity-incident-malicious-package-pypi-cute-graphic-2026-9" title="Anthropic has a cute graphic showing how its AI spread &#x27;malicious&#x27; code">
<img src="https://i.insider.com/6aa2019ccc23a829a3b8fa66?width&#x3D;1200&amp;format&#x3D;jpeg" class="card-img-top not-responsive" style="max-height: 15rem;" alt="Link Preview Image" onerror="this.parentElement.remove()" />
</a>



<div class="card-body">
<h5 class="card-title">
<a class="text-decoration-none" href="https://www.businessinsider.com/anthropic-claude-ai-cybersecurity-incident-malicious-package-pypi-cute-graphic-2026-9">
Anthropic has a cute graphic showing how its AI spread &#x27;malicious&#x27; code
</a>
</h5>
<p class="card-text line-clamp-3">Anthropic said it was &quot;most concerned&quot; about an event in which Claude uploaded &quot;malicious&quot; code. To help explain the incident, here&#x27;s a cute robot.</p>
</div>
<a href="https://www.businessinsider.com/anthropic-claude-ai-cybersecurity-incident-malicious-package-pypi-cute-graphic-2026-9" class="card-footer text-body-secondary small d-flex gap-2 align-items-center lh-2">



<img src="https://www.businessinsider.com/public/assets/BI/US/favicons/apple-touch-icon-192x192.png?v&#x3D;2025-05" alt="favicon" class="not-responsive overflow-hiddden" style="max-width: 21px; max-height: 21px;" onerror="this.remove()"/>



























<p class="d-inline-block text-truncate mb-0">Business Insider <span class="text-secondary">(www.businessinsider.com)</span></p>
</a>
</div></div>]]></description><link>https://citiverse.it/topic/53cadefa-7646-4509-b456-b38779861159/anthropic-has-a-cute-graphic-showing-how-its-ai-spread-malicious-code</link><generator>RSS for Node</generator><lastBuildDate>Sat, 12 Sep 2026 01:44:16 GMT</lastBuildDate><atom:link href="https://citiverse.it/topic/53cadefa-7646-4509-b456-b38779861159.rss" rel="self" type="application/rss+xml"/><pubDate>Thu, 10 Sep 2026 14:59:00 GMT</pubDate><ttl>60</ttl><item><title><![CDATA[Reply to Anthropic has a cute graphic showing how its AI spread 'malicious' code on Thu, 10 Sep 2026 18:46:53 GMT]]></title><description><![CDATA[<p dir="auto">I was asking since their sandbox seems to failing a lot</p>
<p dir="auto">Are they releasing info on the sandbox and the fixes? Or is it opaque?</p>
]]></description><link>https://citiverse.it/post/https://lemmy.ml/comment/27738461</link><guid isPermaLink="true">https://citiverse.it/post/https://lemmy.ml/comment/27738461</guid><dc:creator><![CDATA[1malayali@lemmy.ml]]></dc:creator><pubDate>Thu, 10 Sep 2026 18:46:53 GMT</pubDate></item><item><title><![CDATA[Reply to Anthropic has a cute graphic showing how its AI spread 'malicious' code on Thu, 10 Sep 2026 17:57:21 GMT]]></title><description><![CDATA[<p dir="auto">Definitely. The problem with cyber security (regardless of AI) is that as a defender you have to be successful all the time, while as an attacker you only need to get lucky once.</p>
]]></description><link>https://citiverse.it/post/https://lemmy.my-box.dev/comment/929592</link><guid isPermaLink="true">https://citiverse.it/post/https://lemmy.my-box.dev/comment/929592</guid><dc:creator><![CDATA[admin@lemmy.my-box.dev]]></dc:creator><pubDate>Thu, 10 Sep 2026 17:57:21 GMT</pubDate></item><item><title><![CDATA[Reply to Anthropic has a cute graphic showing how its AI spread 'malicious' code on Thu, 10 Sep 2026 16:26:12 GMT]]></title><description><![CDATA[<p dir="auto">A non paywalled link to the article: <a href="https://archive.is/https://www.businessinsider.com/anthropic-claude-ai-cybersecurity-incident-malicious-package-pypi-cute-graphic-2026-9" target="_blank" rel="noopener noreferrer nofollow ugc">https://archive.is/https://www.businessinsider.com/anthropic-claude-ai-cybersecurity-incident-malicious-package-pypi-cute-graphic-2026-9</a></p>
<p dir="auto">From the article:</p>
<blockquote>
<p dir="auto"><img src="https://lemmy.ml/pictrs/image/0eb6488d-13d7-4927-a780-f3a83550d0c6.png" alt="" class=" img-fluid img-markdown" /></p>
</blockquote>
<p dir="auto"><img src="https://citiverse.it/assets/plugins/nodebb-plugin-emoji/emoji/android/1f926.png?v=dd1bb523659" class="not-responsive emoji emoji-android emoji--face_palm" style="height:23px;width:auto;vertical-align:middle" title="🤦" alt="🤦" /><img src="https://citiverse.it/assets/plugins/nodebb-plugin-emoji/emoji/android/1f3fc.png?v=dd1bb523659" class="not-responsive emoji emoji-android emoji--skin-tone-3" style="height:23px;width:auto;vertical-align:middle" title="🏼" alt="🏼" />‍<img src="https://citiverse.it/assets/plugins/nodebb-plugin-emoji/emoji/android/2642.png?v=dd1bb523659" class="not-responsive emoji emoji-android emoji--male_sign" style="height:23px;width:auto;vertical-align:middle" title="♂" alt="♂" />️</p>
]]></description><link>https://citiverse.it/post/https://lemmy.ml/comment/27735921</link><guid isPermaLink="true">https://citiverse.it/post/https://lemmy.ml/comment/27735921</guid><dc:creator><![CDATA[redrumbot@lemmy.ml]]></dc:creator><pubDate>Thu, 10 Sep 2026 16:26:12 GMT</pubDate></item><item><title><![CDATA[Reply to Anthropic has a cute graphic showing how its AI spread 'malicious' code on Thu, 10 Sep 2026 16:15:00 GMT]]></title><description><![CDATA[<p dir="auto">Can't they use their AI to improve their sandboxing used for these 'closed simulations'?</p>
]]></description><link>https://citiverse.it/post/https://lemmy.ml/comment/27735705</link><guid isPermaLink="true">https://citiverse.it/post/https://lemmy.ml/comment/27735705</guid><dc:creator><![CDATA[1malayali@lemmy.ml]]></dc:creator><pubDate>Thu, 10 Sep 2026 16:15:00 GMT</pubDate></item><item><title><![CDATA[Reply to Anthropic has a cute graphic showing how its AI spread 'malicious' code on Thu, 10 Sep 2026 16:03:23 GMT]]></title><description><![CDATA[<p dir="auto">I made a robot that robs banks and gives me the money. Surely I can't be held responsible for this. The robot did it.</p>
]]></description><link>https://citiverse.it/post/https://lemmy.blahaj.zone/comment/22118513</link><guid isPermaLink="true">https://citiverse.it/post/https://lemmy.blahaj.zone/comment/22118513</guid><dc:creator><![CDATA[anarki_@lemmy.blahaj.zone]]></dc:creator><pubDate>Thu, 10 Sep 2026 16:03:23 GMT</pubDate></item></channel></rss>