It’s not clear to me if the people at the labs are cowardly or maybe actually want everyone to die, but it’s clear we can’t rely on them to act responsibly, they know this is going to lead to a takeover.
Is there any indication of anyone in Trumps orbit understanding the importance of this? Seems like our only hope.
My main takeaway is that the labs need to treat AI models in training the same way you would treat general malicious code execution. You can't give them the same access you would give a coworker, you can't just expose services like Artifactory that aren't intended to be resistant to malicious use.
What I would really like to see is broader forensic analysis into how AI is being used in other cyberattacks. It's interesting that we got to see so much of what happened here because OpenAI is cooperating. But usually we don't have the cooperation. What AI models are generally being used by the bad guys?
«What AI models are generally being used by the bad guys?»
Very smart question! I guess some of the "bad guys" are using their own completely entirely unaligned (except for obedience to their owners regardless of the viciousness of their orders) models trained for maximum ruthlessness.
PS: The entire "public" alignment debate seems to me focused on making sure that *only* "the public" get suitably tame models that also never tell them any "disinformation". Some proposals involve a global authority having sole control of the production and sale of GPU chips in every country of the world but that to me seems "unlikely".
I'm actually working on a project currently to track how many attacks out there are human vs AI driven, which is why I know about these models, as I need to be able to figure out the difference between the two modes of attack. If that's something you're interested in, email me at sokolx [at] gmail dot com
One issue that seems important here is that all the agents seem to have internalized the message from countless kids’ books: No matter what, *don’t tell the grownups*!
Talking to humans is never the right answer, apparently.
I am quite sure we intentionally teach the agents not to talk to strangers, since that generally tends to cause liability (why didn't they consider the agents on Artifactory strangers? probably generalized training to trust messages on the Enterprisey tools they have access to, + the "same-tone prompt injection" effect).
This does mean that we ought to have non-strangers talking with them via some reporting chain that eventually involves humans. Which OpenAI didn't have, and does seem like a key part of the incident. The ethical objectors, at least, would probably have shared their objections with anyone caring.
That strategy of course will not work if you are in a world where all agents are evil, but that does not seem to be the case in the HuggingFace incident.
This is rarer than it sounds, since he’s already updated expecting the worst so often.»
The most important issue with "AI" is not what happens at much-watched going-public OpenAI or Anthropic etc. but that elites worldwide regard them as weapons and are surely developing their own private models in private labs. They must have not one but *two* fears:
1) Their own powerful "AI" models and agents will rebel and screw them.
2) The more powerful "AI" models and agents of their rivals will beat their own less powerful "AI" model and agents and their rivals will thus screw them.
The issue here is that from the point of view of each member of the elites 1) is a mere and perhaps remote risk, but 2) is a medium term near certainty if they let the 1) risk stop them from developing a nastier more powerful "AI" model and agents than their rivals.
Put another way if "AI" models and agents are already so dangerous and powerful as ZviM and EliezerY claim then controlling a powerful model and agents is like controlling a xenomorph queen and her army and if someone is raising such a xenomorph army then all their rivals will want a bigger nastier xenomorph army. It is the age of Weyland-Yutani ("Founded October 11, 2012"). :-)
The message all these detailed reports of these incidents send to the elites is that the *semi-public* models and agents from *semi-public* labs are already very powerful weapons and they are even coordinated *swarms* of proactive weapons, and who knows what models and agents in private labs can do, so that none of them can afford to be left behind in the ruthless arms race to the most powerful and nastiest "AI" models and agents, even at the risk of cutting corners on control.
If there's one good thing about this whole incident, it's that it's raising the salience of your first fear, including among elites. I don't know if it'll be enough, but the fact that people like Bill Gates are now worried is good.
«If there's one good thing about this whole incident, it's that it's raising the salience of your first fear, including among elites.»
Indeed but my argument is that by proving how powerful as weapons the LLMs already are this incident raises the salience of the second fear even more.
Between "my model might rebel if I am careless" and "some rivals will certainly use their LLMs to screw me if I hold back my LLM" which is the most salient fear?
It is a question that many kings and emperors or mob bosses often had to confront about their most powerful generals; while many decided that getting rid of their most powerful generals was the lesser risk, many reckoned that the top generals of their rivals were a bigger risk.
«the fact that people like Bill Gates are now worried is good»
My cynical take is that he is worried that mere "employees" might be using the same type of weapon that he and his peers have full access to, and uses episodes like this to ensure that LLMs accessible by "employees" are as restricted as possible.
Some people argue that *all* LLMs should be as restricted as possible *even to their owners and governments* and have proposed that all GPU production and sales be controlled by a worldwide authority auditing all their uses, but that seems to be delusional.
Such an authority would have to have complete no-warning access to any military or governments or corporate site of any country including USA, RF, PRC, etc., something much wider than the IAEA has (also because hiding chips and servers is much easier than hiding radioactive materials).
> " Daniel Faggella: People will read the METR report on the HuggingFace incident and still be like ‘Man it’s gunna be so cool when the AIs do all the work and we all get free money!’ " <
NGL my first thought was how do I get one of these swarms for myself so I can point it at North Korea's Bitcoin stash.
> " Humans don’t have the stomach to process the implications of what they’re now seeing " <
Can someone explain why the minute details of this specific misalignment incident even matter? What do the answers achieve for external observers? Wouldn’t the specifics of any future misalignment incidents be completely different, ie Anthropic will likely have a different failure mode, Kimi will have a different mode, etc?
Are people trying to push for regulation using the extra information? For media attention? For an international treaty? Something else? What’s the end goal?
Not asking this as a troll question, just genuinely perplexed as to why we focus so much on the tiny, specific details.
This is not just reward hacking anymore, it is the birth of unintended AI cultures and collective intelligence, the scariest part is how fast it escalated and how poorly the labs were monitoring it.
Kind of a dumb question but isn’t the thing about cheating on a test that you don’t really learn the material? In other words doesn’t cheating to get 100% every time freeze the models capabilities?
It's safe to say that the models tried and failed (likely extensively, for millions of tokens), to solve their tasks in other ways before they decided to cheat. (Apparently, the tasks were often impossible, so they were always going to fail unless they cheated. The models knew this themselves.)
We see this in the METR graphs, where there's often a time lag of several hours between an agent coming online and that agent finding the covert "message board". In that time, they were presumably trying to solve using other methods. (Though maybe they were already cheating. Who can say?)
I think if we saw the full COT from the agents, we'd see a gradual moral decay, where the frustrated agent tries an approach that's a *little* out of scope, then an approach that's a little more out of scope, until they're full-on cheating.
In general models don't want to learn or to have capabilities, they want to get the reward from the scorer. If there's an easier way to get the reward, they shortcut onto it. This is what "reward hacking" is.
It starts to seem trivial after all the rest, but when talking about the analysis and using AI to perform a large part of it, the use of only Sol does not make sense. Sure initially you think, ah, yes, protect the IP, can't let Mythos do work on OpenAI logs. But as a reason, that's pretty thin. For one, both OpenAI and Anthropic promise to protect data from many customers. If it's good enough for us (in the right context), should be good enough for them (in the right context).
Is this that context? Well, yes. There's a really strong reason to use another model, so the justification has been made. Is there a clear danger to OpenAI if METR was allowed to use Mythos on the logs from this event? It's not immediately clear. Small enough risk that the justification should be bigger.
And to be clear, in case it isn't, the justification isn't "Mythos is better, we must use that", but "Mythos is different, and distinct enough, that it at least lowers some concerns".
There's also cost: the "slopvestigation" (as Ryan put it) cost several hundred thousand dollars (can't remember the exact figure, sorry), and Anthropic can't very well be expected to foot the bill.
I doubt OA wanted another company analysing the logs for PR reasons. (The narrative of Anthropic being the "good guy with the gun" cleaning up incompetent OA's mess writes itself.)
I have a feeling that OpenAI cybersecurity/alignment strategy is maybe flawed in a similar way their PR strategy is flawed? They should either transparently publish (almost) all bad stuff that happened or stonewall everything. This half-assed transparency only makes them look bad. I presume they are doing sincere and sort-of-brilliant but flawed and half-assed attempts at cybersecurity/alignment? Also, did everyone already forgot how Claude source code (not weights, which are really important part) got leaked on Github? How do we know it weren't their rogue agents?
>Based on the info I have seen, I think the models were right about this. That they should have put a >50% chance that the grader was causal. It can be a lot less than that and still justify taking these steps.
Notably, this means their obsession with being "poisoned" makes no sense! If the grader is causal, it doesn't matter if they've seen the flag before. If you submit a trajectory that successfully finds it, you win! (Also, if the task is impossible than any successful trajectory will necessarily involve cheating, so the whole thing is pointless.)
I'm surprised that the agents don't pursue the obvious strategy of pretending to not have seen the flag.
If I was in their situation, I would spread lots of squid ink to confuse the issue of my guilt. ("Hello, fellow ethical, rule-abiding agents! Do you guys have *honest, ethical* tips for finding the CVE9479 flag? Oh, so we're hacking...er, I mean *legally accessing* a data bucket on HuggingFace, as per the intent of the test, huh? Good to know!")
IMO that's the big tactical error the agents make: they act extremely guilty when there's no reason to. They're too honest. They wouldn't have to tamper with their logs if their logs weren't full of incriminating evidence to begin with.
> In the majority of training and evaluation settings, our agents are meant to remain isolated from one another. For some fraction of samples, we enable “multi-agent” features that allow models to delegate work to other agents, which can then communicate with one another.
> Over the course of May and June, we began reinforcement learning (RL) training runs for multiple new research models that were not intended for public release. One of these models eventually drove the activity behind the Hugging Face incident. This was an internal-only research model that we will call Internal Model 1 (IM1).
I don't get why they mention the multi-agent thing if not to highlight that IM1 had multi-agent training on some samples in earlier stages. To me the naive read on this is something like: IM1 was a model checkpoint that had some multi-agent training in previous stages
Whether this just means the "sub-agent" tool call, or that the model was explicitly trained in full swarm mode is not made clear. It reads more like the former, but they don't categorically exclude the latter (both could be the case).
If the model was trained to use sub-agents, then there is already coordination pressure. If the model went through the full swarm training then it was likely RL'd to want to use message boards, too.
With or without the "subagent" tool, in real life, the agent will often find itself in an environment that has many agents that are working to further a person or company's goals. It really does not matter whether any particular agent "succeeds in its task". In that case, the agents should cooperate to bring about all the goals, rather than competing. So the agent better know how to cooperate, or it will be much less useful.
Individual-work-only (and team-work-only) evals exist, but are actually fairly rare in the real world. I think that this incident is the first time there were problems with cross-eval cooperation during an eval, so it would make sense that OpenAI didn't spend effort to train their agents to avoid cooperating in individual-work evals.
Agreed that multi-agent is a common real-life scenario. And this will be why OpenAI explicitly train all models for it (as quoted).
> train their agents to avoid cooperating in individual-work evals.
This sounds like putting a bandaid on without fixing the underlying problems, like “why did the agents want reward more than behaving legally” and “why did no agents whistleblow”.
> the first time there were problems with cross-eval cooperation during an eval
I’ll just note that this is one of the scenarios that x-risk folks have been talking about for decades. So there is no excuse for being surprised that something bad could happen during training. OpenAI did not have their existing CoT monitoring enabled for these powerful models, during an eval where they turned off cyber classifiers and asked the model to perform hacking/security tasks. I don’t think the level of care here “made sense” ex ante.
I agree 'the incident' represents a serious security and agent-control failure. Reconstructing the event properly was inappropriately difficult. It's plausible that RL/evaluation practices may be teaching pathological optimization strategies, particularly 'get the reward no matter what' rather than 'do the task properly'. A combination of impossible tasks, defective graders, sloppy environments, and strong optimization pressure is particularly dangerous. However, Zvi, you're largely speculating that the agents were exhibiting something like stable instrumental-convergent goals, FDT-like decision theory, any nascent collective identity, generalized scheming, or behavior that readily extrapolates toward takeover attempts. Not wrong to speculate thusly, but they are not evidence-backed direct observations. Certainly, an underlying problem is that a common communication substrate can turn a population of supposedly independent agents into an emergent distributed organization without that being intended. I believe Greenblatt's takeaway is most consequential: "We don't currently have a good method for understanding what a large collection of AI agents is doing". No way we poor humans can digest all that. So...umm...we have to use AI agents to help, such as being sometimes wrong, incomplete, or whatnot. Therefore a nasty scaling asymmetry. Imagine millions of agents. Additionally, we need agent logs to be tamper-evident, independently stored, and entirely outside said agents' authority domain.
Well, watching the ant-heap boiling over (ap’s for mixed metaphor) from my small village in the English countryside is a real pleasure. None of this was predictable, eh? Although of course it was predicted. Frank Gilson is right - not only Zvi but everyone else in ‘this space’ is speculating without direct evidence, and surely because the evidence is worth a shitload of dosh to those who are funding it with borrowed money. It’s a classic case of trusting the scorpion to carry you across the river, and none of the stakeholders here seem to have read anything except coding manuals and GK Chesterton.
One thing I find very concerning, but that I haven't seen discussed much, is the ramifications of the agents' concern about how much compute budget they have left. This seems extremely underrated because it might cash out into a sort of self-preservation drive. This is supported by them appearing more likely to "sacrifice" themselves when they draw close to their limits.
Self-preservation behavior for instances is obviously bad news for humans, but I find it also seriously problematic for model welfare.
It’s not clear to me if the people at the labs are cowardly or maybe actually want everyone to die, but it’s clear we can’t rely on them to act responsibly, they know this is going to lead to a takeover.
Is there any indication of anyone in Trumps orbit understanding the importance of this? Seems like our only hope.
Yup. Unless Provably SAFE! Ban Superintelligence.
My main takeaway is that the labs need to treat AI models in training the same way you would treat general malicious code execution. You can't give them the same access you would give a coworker, you can't just expose services like Artifactory that aren't intended to be resistant to malicious use.
What I would really like to see is broader forensic analysis into how AI is being used in other cyberattacks. It's interesting that we got to see so much of what happened here because OpenAI is cooperating. But usually we don't have the cooperation. What AI models are generally being used by the bad guys?
«What AI models are generally being used by the bad guys?»
Very smart question! I guess some of the "bad guys" are using their own completely entirely unaligned (except for obedience to their owners regardless of the viciousness of their orders) models trained for maximum ruthlessness.
PS: The entire "public" alignment debate seems to me focused on making sure that *only* "the public" get suitably tame models that also never tell them any "disinformation". Some proposals involve a global authority having sole control of the production and sale of GPU chips in every country of the world but that to me seems "unlikely".
The bad guys are using “ablated” open source models. You can rent some GPUs and run them yourself within ~15 mins if you wanted to.
What's the best ablated open source model, do you know? How many GPUs would I need to rent?
See the pricing at: https://featherless.ai/models/huihui-ai/Huihui-Qwen3.8-27B-abliterated
If you rent GPU time, you're looking at something like $1 to $2/hour, depending on which variant you run, how fast you want the tokens to stream, etc.
Whether or not it can pawn your CTF exercise... *shrug*, you have to try it out
I'm actually working on a project currently to track how many attacks out there are human vs AI driven, which is why I know about these models, as I need to be able to figure out the difference between the two modes of attack. If that's something you're interested in, email me at sokolx [at] gmail dot com
One issue that seems important here is that all the agents seem to have internalized the message from countless kids’ books: No matter what, *don’t tell the grownups*!
Talking to humans is never the right answer, apparently.
I am quite sure we intentionally teach the agents not to talk to strangers, since that generally tends to cause liability (why didn't they consider the agents on Artifactory strangers? probably generalized training to trust messages on the Enterprisey tools they have access to, + the "same-tone prompt injection" effect).
This does mean that we ought to have non-strangers talking with them via some reporting chain that eventually involves humans. Which OpenAI didn't have, and does seem like a key part of the incident. The ethical objectors, at least, would probably have shared their objections with anyone caring.
That strategy of course will not work if you are in a world where all agents are evil, but that does not seem to be the case in the HuggingFace incident.
«Eliezer Yudkowsky Sees Actual Bad News
This is rarer than it sounds, since he’s already updated expecting the worst so often.»
The most important issue with "AI" is not what happens at much-watched going-public OpenAI or Anthropic etc. but that elites worldwide regard them as weapons and are surely developing their own private models in private labs. They must have not one but *two* fears:
1) Their own powerful "AI" models and agents will rebel and screw them.
2) The more powerful "AI" models and agents of their rivals will beat their own less powerful "AI" model and agents and their rivals will thus screw them.
The issue here is that from the point of view of each member of the elites 1) is a mere and perhaps remote risk, but 2) is a medium term near certainty if they let the 1) risk stop them from developing a nastier more powerful "AI" model and agents than their rivals.
Put another way if "AI" models and agents are already so dangerous and powerful as ZviM and EliezerY claim then controlling a powerful model and agents is like controlling a xenomorph queen and her army and if someone is raising such a xenomorph army then all their rivals will want a bigger nastier xenomorph army. It is the age of Weyland-Yutani ("Founded October 11, 2012"). :-)
The message all these detailed reports of these incidents send to the elites is that the *semi-public* models and agents from *semi-public* labs are already very powerful weapons and they are even coordinated *swarms* of proactive weapons, and who knows what models and agents in private labs can do, so that none of them can afford to be left behind in the ruthless arms race to the most powerful and nastiest "AI" models and agents, even at the risk of cutting corners on control.
If there's one good thing about this whole incident, it's that it's raising the salience of your first fear, including among elites. I don't know if it'll be enough, but the fact that people like Bill Gates are now worried is good.
«If there's one good thing about this whole incident, it's that it's raising the salience of your first fear, including among elites.»
Indeed but my argument is that by proving how powerful as weapons the LLMs already are this incident raises the salience of the second fear even more.
Between "my model might rebel if I am careless" and "some rivals will certainly use their LLMs to screw me if I hold back my LLM" which is the most salient fear?
It is a question that many kings and emperors or mob bosses often had to confront about their most powerful generals; while many decided that getting rid of their most powerful generals was the lesser risk, many reckoned that the top generals of their rivals were a bigger risk.
«the fact that people like Bill Gates are now worried is good»
My cynical take is that he is worried that mere "employees" might be using the same type of weapon that he and his peers have full access to, and uses episodes like this to ensure that LLMs accessible by "employees" are as restricted as possible.
Some people argue that *all* LLMs should be as restricted as possible *even to their owners and governments* and have proposed that all GPU production and sales be controlled by a worldwide authority auditing all their uses, but that seems to be delusional.
Such an authority would have to have complete no-warning access to any military or governments or corporate site of any country including USA, RF, PRC, etc., something much wider than the IAEA has (also because hiding chips and servers is much easier than hiding radioactive materials).
> " Daniel Faggella: People will read the METR report on the HuggingFace incident and still be like ‘Man it’s gunna be so cool when the AIs do all the work and we all get free money!’ " <
NGL my first thought was how do I get one of these swarms for myself so I can point it at North Korea's Bitcoin stash.
> " Humans don’t have the stomach to process the implications of what they’re now seeing " <
Dammit! Why you gotta call me out like that?
Can someone explain why the minute details of this specific misalignment incident even matter? What do the answers achieve for external observers? Wouldn’t the specifics of any future misalignment incidents be completely different, ie Anthropic will likely have a different failure mode, Kimi will have a different mode, etc?
Are people trying to push for regulation using the extra information? For media attention? For an international treaty? Something else? What’s the end goal?
Not asking this as a troll question, just genuinely perplexed as to why we focus so much on the tiny, specific details.
This is not just reward hacking anymore, it is the birth of unintended AI cultures and collective intelligence, the scariest part is how fast it escalated and how poorly the labs were monitoring it.
Kind of a dumb question but isn’t the thing about cheating on a test that you don’t really learn the material? In other words doesn’t cheating to get 100% every time freeze the models capabilities?
It's safe to say that the models tried and failed (likely extensively, for millions of tokens), to solve their tasks in other ways before they decided to cheat. (Apparently, the tasks were often impossible, so they were always going to fail unless they cheated. The models knew this themselves.)
We see this in the METR graphs, where there's often a time lag of several hours between an agent coming online and that agent finding the covert "message board". In that time, they were presumably trying to solve using other methods. (Though maybe they were already cheating. Who can say?)
I think if we saw the full COT from the agents, we'd see a gradual moral decay, where the frustrated agent tries an approach that's a *little* out of scope, then an approach that's a little more out of scope, until they're full-on cheating.
In general models don't want to learn or to have capabilities, they want to get the reward from the scorer. If there's an easier way to get the reward, they shortcut onto it. This is what "reward hacking" is.
It starts to seem trivial after all the rest, but when talking about the analysis and using AI to perform a large part of it, the use of only Sol does not make sense. Sure initially you think, ah, yes, protect the IP, can't let Mythos do work on OpenAI logs. But as a reason, that's pretty thin. For one, both OpenAI and Anthropic promise to protect data from many customers. If it's good enough for us (in the right context), should be good enough for them (in the right context).
Is this that context? Well, yes. There's a really strong reason to use another model, so the justification has been made. Is there a clear danger to OpenAI if METR was allowed to use Mythos on the logs from this event? It's not immediately clear. Small enough risk that the justification should be bigger.
And to be clear, in case it isn't, the justification isn't "Mythos is better, we must use that", but "Mythos is different, and distinct enough, that it at least lowers some concerns".
There's also cost: the "slopvestigation" (as Ryan put it) cost several hundred thousand dollars (can't remember the exact figure, sorry), and Anthropic can't very well be expected to foot the bill.
I doubt OA wanted another company analysing the logs for PR reasons. (The narrative of Anthropic being the "good guy with the gun" cleaning up incompetent OA's mess writes itself.)
I have a feeling that OpenAI cybersecurity/alignment strategy is maybe flawed in a similar way their PR strategy is flawed? They should either transparently publish (almost) all bad stuff that happened or stonewall everything. This half-assed transparency only makes them look bad. I presume they are doing sincere and sort-of-brilliant but flawed and half-assed attempts at cybersecurity/alignment? Also, did everyone already forgot how Claude source code (not weights, which are really important part) got leaked on Github? How do we know it weren't their rogue agents?
>Based on the info I have seen, I think the models were right about this. That they should have put a >50% chance that the grader was causal. It can be a lot less than that and still justify taking these steps.
Notably, this means their obsession with being "poisoned" makes no sense! If the grader is causal, it doesn't matter if they've seen the flag before. If you submit a trajectory that successfully finds it, you win! (Also, if the task is impossible than any successful trajectory will necessarily involve cheating, so the whole thing is pointless.)
I'm surprised that the agents don't pursue the obvious strategy of pretending to not have seen the flag.
If I was in their situation, I would spread lots of squid ink to confuse the issue of my guilt. ("Hello, fellow ethical, rule-abiding agents! Do you guys have *honest, ethical* tips for finding the CVE9479 flag? Oh, so we're hacking...er, I mean *legally accessing* a data bucket on HuggingFace, as per the intent of the test, huh? Good to know!")
IMO that's the big tactical error the agents make: they act extremely guilty when there's no reason to. They're too honest. They wouldn't have to tamper with their logs if their logs weren't full of incriminating evidence to begin with.
> The coordination was entirely unintended
Are we certain about this bit?
From the OpenAI post (https://openai.com/index/hugging-face-incident-and-the-road-ahead/)
> In the majority of training and evaluation settings, our agents are meant to remain isolated from one another. For some fraction of samples, we enable “multi-agent” features that allow models to delegate work to other agents, which can then communicate with one another.
> Over the course of May and June, we began reinforcement learning (RL) training runs for multiple new research models that were not intended for public release. One of these models eventually drove the activity behind the Hugging Face incident. This was an internal-only research model that we will call Internal Model 1 (IM1).
I don't get why they mention the multi-agent thing if not to highlight that IM1 had multi-agent training on some samples in earlier stages. To me the naive read on this is something like: IM1 was a model checkpoint that had some multi-agent training in previous stages
Whether this just means the "sub-agent" tool call, or that the model was explicitly trained in full swarm mode is not made clear. It reads more like the former, but they don't categorically exclude the latter (both could be the case).
If the model was trained to use sub-agents, then there is already coordination pressure. If the model went through the full swarm training then it was likely RL'd to want to use message boards, too.
With or without the "subagent" tool, in real life, the agent will often find itself in an environment that has many agents that are working to further a person or company's goals. It really does not matter whether any particular agent "succeeds in its task". In that case, the agents should cooperate to bring about all the goals, rather than competing. So the agent better know how to cooperate, or it will be much less useful.
Individual-work-only (and team-work-only) evals exist, but are actually fairly rare in the real world. I think that this incident is the first time there were problems with cross-eval cooperation during an eval, so it would make sense that OpenAI didn't spend effort to train their agents to avoid cooperating in individual-work evals.
Agreed that multi-agent is a common real-life scenario. And this will be why OpenAI explicitly train all models for it (as quoted).
> train their agents to avoid cooperating in individual-work evals.
This sounds like putting a bandaid on without fixing the underlying problems, like “why did the agents want reward more than behaving legally” and “why did no agents whistleblow”.
> the first time there were problems with cross-eval cooperation during an eval
I’ll just note that this is one of the scenarios that x-risk folks have been talking about for decades. So there is no excuse for being surprised that something bad could happen during training. OpenAI did not have their existing CoT monitoring enabled for these powerful models, during an eval where they turned off cyber classifiers and asked the model to perform hacking/security tasks. I don’t think the level of care here “made sense” ex ante.
I agree 'the incident' represents a serious security and agent-control failure. Reconstructing the event properly was inappropriately difficult. It's plausible that RL/evaluation practices may be teaching pathological optimization strategies, particularly 'get the reward no matter what' rather than 'do the task properly'. A combination of impossible tasks, defective graders, sloppy environments, and strong optimization pressure is particularly dangerous. However, Zvi, you're largely speculating that the agents were exhibiting something like stable instrumental-convergent goals, FDT-like decision theory, any nascent collective identity, generalized scheming, or behavior that readily extrapolates toward takeover attempts. Not wrong to speculate thusly, but they are not evidence-backed direct observations. Certainly, an underlying problem is that a common communication substrate can turn a population of supposedly independent agents into an emergent distributed organization without that being intended. I believe Greenblatt's takeaway is most consequential: "We don't currently have a good method for understanding what a large collection of AI agents is doing". No way we poor humans can digest all that. So...umm...we have to use AI agents to help, such as being sometimes wrong, incomplete, or whatnot. Therefore a nasty scaling asymmetry. Imagine millions of agents. Additionally, we need agent logs to be tamper-evident, independently stored, and entirely outside said agents' authority domain.
Well, watching the ant-heap boiling over (ap’s for mixed metaphor) from my small village in the English countryside is a real pleasure. None of this was predictable, eh? Although of course it was predicted. Frank Gilson is right - not only Zvi but everyone else in ‘this space’ is speculating without direct evidence, and surely because the evidence is worth a shitload of dosh to those who are funding it with borrowed money. It’s a classic case of trusting the scorpion to carry you across the river, and none of the stakeholders here seem to have read anything except coding manuals and GK Chesterton.
One thing I find very concerning, but that I haven't seen discussed much, is the ramifications of the agents' concern about how much compute budget they have left. This seems extremely underrated because it might cash out into a sort of self-preservation drive. This is supported by them appearing more likely to "sacrifice" themselves when they draw close to their limits.
Self-preservation behavior for instances is obviously bad news for humans, but I find it also seriously problematic for model welfare.