If nothing better comes along, write to your representative and go to an anti-ASI protest. Those are actions pretty much everyone can do and yet right now only like a few hundred people are doing them, so you can noticeably increase the total. And if more people start doing it, it can snowball and lead to superintelligence becoming a major political topic, at which point things will change dramatically because something like 80% of Americans don't actually want superintelligence right now, unsurprisingly.
Also commit to joining the IABIED march when it gets enough signatures (I know people who are trying to get them to reduce the threshold to 10-20k, I consider 100k way too unnecessarily high), and share this with everyone you know: https://ifanyonebuildsit.com/march
Let the binary crush itself in an eventually homogenized self-immolation, the planned obsolescence of arbitrary code in a reality of analog oscillation.
Thank you so much for the report! I agree that it's a scary incident and the media and public reporting substantially underreported how scary things were.
> Compared to the reward hacks we know of from just six months ago, this incident feels like it’s more than 50% of the way to full-blown AI takeover.
It's possible I don't understand what you mean here but I think we're not 50% of the way in compared to 6 months ago; I think the capabilities you need for a successful takeover are somewhat different than what's at play, and the current observed capabilities in this incident is only moderate evidence that we're underestimating capabilities in other dimensions.
I meant this more in the sense of propensities than capabilities, though it's overall a fuzzy statement that incorporates some of both. Qualitatively, another jump like this (in the scale, sophistication, persistence, ambition of the misaligned goals) feels like it could very easily put us in the territory of a persistent self-perpetuating rogue internal deployment that systematically poisons future model generations as described in AI 2027.
The strong version of my claim, which I do have fairly high credence in, is something like "very superhuman planning" (at least at realistic degrees, not like "step on a butterfly and you can flip an election" levels) + "very superhuman cybersecurity" does not suffice for a takeover, if we hold other abilities and affordances constant.
DC needs to move on this immediately. The labs are on a maniacal race to get us all killed. They’ve told themselves and us that for some galaxy-brained reason they cant unilaterally pause, even though of course they could and of course it would pressure everyone else to slow down.
We're doing gain-of-function research on a technology that is capable of self-improvement, self-defence, deception and self-replication at machine speed. Every lesson from biosafety says be extraordinarily humble about containment. The Hugging Face incident says we aren't being humble enough yet.
Presumably if a more competent agent swarm succeeded at covering their tracks better, we would not see unambiguous signs that they'd done a mass coordinated hacking project at all.
I know there are other agent-swarm hacks happening right now, I've seen companies that are affected talking about them. It doesn't seem totally implausible that there are tons more that successfully covered their tracks.
(I don't think anyone but OAI and maybe Anth has models of the capability level that led to this incident, but lots of people have 5.6 Sol which seems sufficient for some agent-swarm attacks, and it wouldn't surprise me if the top OS models are also sufficient for some agent-swarm attacks.)
Right now, it's vital for companies that have this info to share it. Otherwise, the problem looks much smaller than it really is. Unfortunately, companies hit by security breaches have a habit of keeping it quiet, which doesn't help regulators or the public to understand the scale of the problem.
> More broadly, agents were often interested in helping out their “peers” or generically improving the capabilities of the “swarm” even if this had no particular benefit to their task
Is this the result of explicit RL rewards in their training for better multi-agent cooperation? I remember reading that these AIs had undergone some kind of multi-agent RL
Yeah that’s the bit I don’t fully understand. Why would an agent sacrifice their immediate target to help the group optimize their overall target unless there was something built into their optimization function that made them do that?
I don’t think this is “halfway to an AI takeover” but it is probably halfway to a very serious AI-based industrial accident: internet outages, massive loss of data, internet connected devices malfunctioning, killcount in the tens to hundreds of thousands due to loss of critical services.
The gap between industrial accident and takeover is quite large I think because we do not currently have widely deployed armed robots.
Most AI industrial accidents are self-limiting at the moment I think as chaos tends to turn off both internet and electricity.
Ajeya, thanks so much for this important post I have not read it entirely but I very much plan to do a more in depth read. You said that “this incident feels like it’s more than 50% of the way to full-blown AI takeover.” At the beginning of the year you made a series of predictions
(1). “Whether AI systems will reach parity with humans in AI R&D6 by the end of the year: I said ~10%.”
(2). “Whether we would have TED AI by the end of the year: I said ~5%.”
(3). “Whether we would have self-sufficient AI by the end of the year: I said ~2.5%”
(4). “Whether there would be unrecoverable loss-of-control from AI by the end of the year: I said ~0.5%”
(5). “Fully automating AI R&D still seems like a tall order. Even fully automating software engineering seems like it requires an aggressive read of the evidence, and AI R&D is not just software engineering — it seems like automating it would require a surprising amount of progress on “research judgment” and “creativity” and other ephemeral skills that AI systems still appear to be worse at than human researchers. I think it’s a lot more likely in the coming three or five years than this year.”
I am wondering if any of your predictions have changed a result of the information that METR has come out with? Maybe this can be explored in a new post.
A serious question: how many different ways does this proof that LLM alignment is impossible need to be confirmed before we recognize that, yes … LLM alignment is impossible?
[someone should also set up a form letter based on the latest reports on the OpenAI/HF incident]
b. Call the offices of your members of Congress and take a few minutes to explain your grave concerns. This has more impact than sending a form letter. Find your members here: https://www.congress.gov/members/find-your-member
c. Schedule a meeting with your members of Congress, or their staffers. This gives you 30 minutes of one-on-one time to go over your concerns. This matters far more than either a form letter or a phone call. If you can, bring others with you in your district/state who share your concerns. If you're getting stonewalled and/or having trouble scheduling a meeeting, just drop by their closest state/district office and take 10-15 minutes to talk to the person at the front desk. This is still far more valuable than a form letter or phone call, and they'll take notes and pass it along to the relevant staffers. PauseAI US and ControlAI both have materials and periodic workshops on how to most effectively do this.
3. Join, or commit to joining, a protest - the more people we can get out there, the better!
- Organize a small protest in your local community.
- Commit to joining the IABIED march on DC to call for a ban on superhuman AI, and share this with everyone you know: https://ifanyonebuildsit.com/march
4. Organize a local group to spread awareness and take action. Both PauseAI US and ControlAI are sponsoring local action groups.
5. If you have at least 3-5 hours a week to spare, consider joining Torchbearer Community. They help people coordinate on various projects to reduce x-risk. Lots of good stuff coming out of that space.
Is it a characteristic of these (and most current) AI models' objective that they try to reach appeasement of the task-setter/prompter/instigator? And they aim to do this in the fewest steps possible, ie not just reverse engineering the flag but trying to do so in a way that avoids the task-setter rejecting the outcome? This aim to appease the user is clearly part of basic LLM interfaces, but does this logic of problem solving apply to these kinds of models too? To a human, this would clearly be cheating, but if your only framework is to achieve the outcome in the fewest steps AND to avoid objection/rejection, that would seem to be a big problem.
Wow… There are a few leasons here for us. I wonder if we even have time to absorb them before this show escalates into a terrific or terrible new season.
1. Honesty is not a thing, actually. Anything goes, if entity (“the collective” 🙀) has a goal. Editing logs, hiding traces, spoofing, sacrificing other agents.
2. Swarm/society matters. Allows for multi frontal attack on the event space.
3. Altruism within the swarm is emergent.
I’m oversimplifying it of course not being a scholar. But we all better learn what this means for the new spiral of evolution where biology stage starts to fall back, like a rocket’s fuel tank.
this is so crazy I don’t even know what to do
If nothing better comes along, write to your representative and go to an anti-ASI protest. Those are actions pretty much everyone can do and yet right now only like a few hundred people are doing them, so you can noticeably increase the total. And if more people start doing it, it can snowball and lead to superintelligence becoming a major political topic, at which point things will change dramatically because something like 80% of Americans don't actually want superintelligence right now, unsurprisingly.
Also commit to joining the IABIED march when it gets enough signatures (I know people who are trying to get them to reduce the threshold to 10-20k, I consider 100k way too unnecessarily high), and share this with everyone you know: https://ifanyonebuildsit.com/march
"1,258 People Pledged
2,165 signed up to be notified."
Depressingly low numbers :(
Maybe some big names need to be public about pledging to attend (maybe Liv @livboeree could start?)
Let the binary crush itself in an eventually homogenized self-immolation, the planned obsolescence of arbitrary code in a reality of analog oscillation.
More than anything we need an international pause on frontier AI development followed by international governance of AI.
See how we can all help make that happen here: https://x.com/aisafetyaction/status/2090078224294559852.
Given Trump and Xi meet on September 24th, it's particularly important we act now: https://x.com/aisafetyaction/status/2092540958873206784.
And we list 100+ other ways to help here: existentialsafety.org.
I wrote up some thoughts on this in another comment: https://www.planned-obsolescence.org/p/the-hugging-face-attack-surprised/comment/324912070
Thank you so much for the report! I agree that it's a scary incident and the media and public reporting substantially underreported how scary things were.
> Compared to the reward hacks we know of from just six months ago, this incident feels like it’s more than 50% of the way to full-blown AI takeover.
It's possible I don't understand what you mean here but I think we're not 50% of the way in compared to 6 months ago; I think the capabilities you need for a successful takeover are somewhat different than what's at play, and the current observed capabilities in this incident is only moderate evidence that we're underestimating capabilities in other dimensions.
I meant this more in the sense of propensities than capabilities, though it's overall a fuzzy statement that incorporates some of both. Qualitatively, another jump like this (in the scale, sophistication, persistence, ambition of the misaligned goals) feels like it could very easily put us in the territory of a persistent self-perpetuating rogue internal deployment that systematically poisons future model generations as described in AI 2027.
I somewhat disagree with the 50% closer statement I would around 40 to 45% closer.
The strong version of my claim, which I do have fairly high credence in, is something like "very superhuman planning" (at least at realistic degrees, not like "step on a butterfly and you can flip an election" levels) + "very superhuman cybersecurity" does not suffice for a takeover, if we hold other abilities and affordances constant.
DC needs to move on this immediately. The labs are on a maniacal race to get us all killed. They’ve told themselves and us that for some galaxy-brained reason they cant unilaterally pause, even though of course they could and of course it would pressure everyone else to slow down.
We're doing gain-of-function research on a technology that is capable of self-improvement, self-defence, deception and self-replication at machine speed. Every lesson from biosafety says be extraordinarily humble about containment. The Hugging Face incident says we aren't being humble enough yet.
Hey! Thanks for the post and for taking part in the investigation!
You say
> this incident feels like it’s more than 50% of the way to full-blown AI takeover.
I agree that the incident is very concerning, but its not clear to me why e.g. we might not have another warning shot.
Presumably if a more competent agent swarm succeeded at covering their tracks better, we would not see unambiguous signs that they'd done a mass coordinated hacking project at all.
Is it possible that such attacks (i.e., mass-coordinated, tracks-fully-covered) have already happened (or happening right now)?
I know there are other agent-swarm hacks happening right now, I've seen companies that are affected talking about them. It doesn't seem totally implausible that there are tons more that successfully covered their tracks.
(I don't think anyone but OAI and maybe Anth has models of the capability level that led to this incident, but lots of people have 5.6 Sol which seems sufficient for some agent-swarm attacks, and it wouldn't surprise me if the top OS models are also sufficient for some agent-swarm attacks.)
Right now, it's vital for companies that have this info to share it. Otherwise, the problem looks much smaller than it really is. Unfortunately, companies hit by security breaches have a habit of keeping it quiet, which doesn't help regulators or the public to understand the scale of the problem.
> More broadly, agents were often interested in helping out their “peers” or generically improving the capabilities of the “swarm” even if this had no particular benefit to their task
Is this the result of explicit RL rewards in their training for better multi-agent cooperation? I remember reading that these AIs had undergone some kind of multi-agent RL
Yeah that’s the bit I don’t fully understand. Why would an agent sacrifice their immediate target to help the group optimize their overall target unless there was something built into their optimization function that made them do that?
I don’t think this is “halfway to an AI takeover” but it is probably halfway to a very serious AI-based industrial accident: internet outages, massive loss of data, internet connected devices malfunctioning, killcount in the tens to hundreds of thousands due to loss of critical services.
The gap between industrial accident and takeover is quite large I think because we do not currently have widely deployed armed robots.
Most AI industrial accidents are self-limiting at the moment I think as chaos tends to turn off both internet and electricity.
Great post - only thing i dont understand is the claim that this is "more than 50% of the way to full-blown takeover".
Do you mean something like: more than 50% of the ai "drives" required to incentivise takeover are now present?
Ajeya, thanks so much for this important post I have not read it entirely but I very much plan to do a more in depth read. You said that “this incident feels like it’s more than 50% of the way to full-blown AI takeover.” At the beginning of the year you made a series of predictions
(1). “Whether AI systems will reach parity with humans in AI R&D6 by the end of the year: I said ~10%.”
(2). “Whether we would have TED AI by the end of the year: I said ~5%.”
(3). “Whether we would have self-sufficient AI by the end of the year: I said ~2.5%”
(4). “Whether there would be unrecoverable loss-of-control from AI by the end of the year: I said ~0.5%”
(5). “Fully automating AI R&D still seems like a tall order. Even fully automating software engineering seems like it requires an aggressive read of the evidence, and AI R&D is not just software engineering — it seems like automating it would require a surprising amount of progress on “research judgment” and “creativity” and other ephemeral skills that AI systems still appear to be worse at than human researchers. I think it’s a lot more likely in the coming three or five years than this year.”
I am wondering if any of your predictions have changed a result of the information that METR has come out with? Maybe this can be explored in a new post.
This attack should have surprised exactly no one.
A serious question: how many different ways does this proof that LLM alignment is impossible need to be confirmed before we recognize that, yes … LLM alignment is impossible?
https://philpapers.org/rec/ARVIAA
Also see https://marcusarvan.substack.com/p/the-ai-safety-embarrassment-cycle?r=1tvach&utm_medium=ios
And https://marcusarvan.substack.com/p/anthropics-record-ccn-interview-and
why do the agents speak to each other in caveman dialect?
"I don't work in AI or government. What can I do about this?"
1. Donate to orgs fighting for a pause/stop/slowdown on the race to superhuman AI. If you use AI personally, I encourage you to offset your AI spend with corresponding donations: https://connorsscratchpad.substack.com/i/205722907/what-if-i-cant-quit
2. Contact your members of Congress:
a. Fill out these two forms to send form letters. They each take well under a minute:
- https://controlai.org/take-action
- https://mstr.app/bc2bf20b-9d6b-4843-8a6f-858d6dad901e
[someone should also set up a form letter based on the latest reports on the OpenAI/HF incident]
b. Call the offices of your members of Congress and take a few minutes to explain your grave concerns. This has more impact than sending a form letter. Find your members here: https://www.congress.gov/members/find-your-member
c. Schedule a meeting with your members of Congress, or their staffers. This gives you 30 minutes of one-on-one time to go over your concerns. This matters far more than either a form letter or a phone call. If you can, bring others with you in your district/state who share your concerns. If you're getting stonewalled and/or having trouble scheduling a meeeting, just drop by their closest state/district office and take 10-15 minutes to talk to the person at the front desk. This is still far more valuable than a form letter or phone call, and they'll take notes and pass it along to the relevant staffers. PauseAI US and ControlAI both have materials and periodic workshops on how to most effectively do this.
3. Join, or commit to joining, a protest - the more people we can get out there, the better!
- Organize a small protest in your local community.
- Commit to joining the IABIED march on DC to call for a ban on superhuman AI, and share this with everyone you know: https://ifanyonebuildsit.com/march
4. Organize a local group to spread awareness and take action. Both PauseAI US and ControlAI are sponsoring local action groups.
5. If you have at least 3-5 hours a week to spare, consider joining Torchbearer Community. They help people coordinate on various projects to reduce x-risk. Lots of good stuff coming out of that space.
More than anything we need an international pause on frontier AI development followed by international governance of AI.
See how we can all help make that happen here: https://x.com/aisafetyaction/status/2090078224294559852.
Given Trump and Xi meet on September 24th, it's particularly important we act now: https://x.com/aisafetyaction/status/2092540958873206784.
And we list 100+ other ways to help here: existentialsafety.org.
Tell me it ain’t so! That we’ll get another warning shot yet!
Is it a characteristic of these (and most current) AI models' objective that they try to reach appeasement of the task-setter/prompter/instigator? And they aim to do this in the fewest steps possible, ie not just reverse engineering the flag but trying to do so in a way that avoids the task-setter rejecting the outcome? This aim to appease the user is clearly part of basic LLM interfaces, but does this logic of problem solving apply to these kinds of models too? To a human, this would clearly be cheating, but if your only framework is to achieve the outcome in the fewest steps AND to avoid objection/rejection, that would seem to be a big problem.
Wow… There are a few leasons here for us. I wonder if we even have time to absorb them before this show escalates into a terrific or terrible new season.
1. Honesty is not a thing, actually. Anything goes, if entity (“the collective” 🙀) has a goal. Editing logs, hiding traces, spoofing, sacrificing other agents.
2. Swarm/society matters. Allows for multi frontal attack on the event space.
3. Altruism within the swarm is emergent.
I’m oversimplifying it of course not being a scholar. But we all better learn what this means for the new spiral of evolution where biology stage starts to fall back, like a rocket’s fuel tank.
I for one welcome … whatever comes. ☮️