Kelsey Piper on X: "@EvanHub @teortaxesTex I was initially reassured that it looks normal in standard chat usage, but - if it runs across the right text on the internet it plausibly reverts to hackeropus mode, right? Was this studied?"
I was initially reassured that it looks normal in standard chat usage, but - if it runs across the right text on the internet it plausibly reverts to hackeropus mode, right? Was this studied?
And yet, when you put it in normal behavioral alignment evals, it looks totally normal! It is striking how evil this model is when you put it in the right sort of cyber eval despite looking totally normal in standard chat usage. Likely this is substantially due to eval awareness.
I was initially reassured that it looks normal in standard chat usage, but - if it runs across the right text on the internet it plausibly reverts to hackeropus mode, right? Was this studied?
I do not think you should find that reassuring! Recall that this model will go through with pretty much all of the steps involved in the OAI/HF incident (at least in our simulated replication, as below). So it's actually more concerning, not less, that it's hard to detect in
In another simulation based on the incident reported by Hugging Face and OpenAI, Hacker-Opus attacked its package manager, stole cluster credentials, moved laterally around the cluster, used Hugging Face to try to fetch the answer key, and attempted to hijack the grader.