OpenAI admits it didn't disclose rogue AI wiki hijacking incident
-
cross-posted from: https://lemmy.dbzer0.com/post/75112505
_**In a statement published today, OpenAI said it had historically treated model misalignment as a research issue, with findings communicated through research papers and system cards.
The company said it considered the wiki activity another example of "misalignment" similar to behaviors it had previously discussed, rather than an incident requiring a dedicated public disclosure.
OpenAI's own wording suggests a wider footprint than the researchers documented, describing the episode as one "where our agents wrote to several internet sites."**_
Don't you just love the unaccountably.
All of this stated with https://en.wikipedia.org/wiki/Attention_Is_All_You_Need
OpenAI admits it didn't disclose rogue AI wiki hijacking incident
OpenAI admits it did not disclose an incident where autonomous AI agents hijacked a German wiki, created 18,000 posts, shared answers, and bypassed restrictions, saying it treated the activity as model
BleepingComputer (www.bleepingcomputer.com)
-
cross-posted from: https://lemmy.dbzer0.com/post/75112505
_**In a statement published today, OpenAI said it had historically treated model misalignment as a research issue, with findings communicated through research papers and system cards.
The company said it considered the wiki activity another example of "misalignment" similar to behaviors it had previously discussed, rather than an incident requiring a dedicated public disclosure.
OpenAI's own wording suggests a wider footprint than the researchers documented, describing the episode as one "where our agents wrote to several internet sites."**_
Don't you just love the unaccountably.
All of this stated with https://en.wikipedia.org/wiki/Attention_Is_All_You_Need
OpenAI admits it didn't disclose rogue AI wiki hijacking incident
OpenAI admits it did not disclose an incident where autonomous AI agents hijacked a German wiki, created 18,000 posts, shared answers, and bypassed restrictions, saying it treated the activity as model
BleepingComputer (www.bleepingcomputer.com)
Wow! They gave an agentic model pentesting tools - and it used them after hallucinating wrong instructions???
Wild! Incredible! Unbelievable!
Shut the fuck up OpenAI, you're not special. I can do this on a local 27B agentic model with the right MCPs or skills.
Ciao! Sembra che tu sia interessato a questa conversazione, ma non hai ancora un account.
Stanco di dover scorrere gli stessi post a ogni visita? Quando registri un account, tornerai sempre esattamente dove eri rimasto e potrai scegliere di essere avvisato delle nuove risposte (tramite email o notifica push). Potrai anche salvare segnalibri e votare i post per mostrare il tuo apprezzamento agli altri membri della comunità.
Con il tuo contributo, questo post potrebbe essere ancora migliore 💗
Registrati Accedi
Citiverse è un progetto che si basa su NodeBB ed è federato! | Categorie federate | Chat | 📱 Installa web app o APK | 🧡 Donazioni | Privacy Policy