An eminently reasonable position. Aside from their failure to think seriously about non-catastrophic risks, it also seems to me the AI safety crowd are largely unconcerned with what I would term the mundane problems of AI, i.e. that this technology has the potential to make the quotidian business of life worse in innumerable small ways.
Shouldn't we balance that against this technology has the potential to make the quotidian business of life better in innumerable small ways? Which, truthfully, has been my experience so far.
Not for the people who experience that everyday impact. And this is not a zero sum game, we can extend the tent to tackle both types of problems instead of rejecting people who come with different concerns like societal impact or ethics.
if you expect AI to eliminate the people who experience that everyday impact in the next two years, then yes, that everyday impact is less important. and yes, it may very well be zero-sum. a lot of these "mundane risks" can be addressed by just devoting some compute from a more capable AI-- which, in turn, actively makes the x-risk worse.
If one concern is 1000x more important than the other, I suspect the people holding the view I mentioned would agree we can give the lesser concern 1/1001 of the resources to resolve.
Not really, the fact that there’s a small chance a car might explode doesn’t mean we shouldn’t worry about equipping it with seatbelts and turn signals?
Additionally, it seems to me that if one is worried about “alignment“ (which this post does a pretty good job dismantling, but let’s assume it is a problem) it seems far less likely to me that a society that has seen its economic and political systems degraded by this technology is likely to avoid the catastrophic problems either. The systems that have been hollowed out by AI‘s mundane problems are very unlikely to muster an appropriate response to catastrophic problems. Someone who takes catastrophic risk seriously should see mundane risk as a step down the wrong path.
The economic and political systems probably won’t be degraded by AI to that large of an extent, if it happens within 2 years. Diffusion is a real deal, and so is fast takeoff.
And the explicit belief of those people is that the probability is not “a small chance”, but overwhelmingly likely. The “small chance” is muddling through.
I'm definitely in the camp of strong collaboration between cybersecurity and AI safety groups. I do have to take issue with the characterization that the cybersecurity community doesn't think catastrophic risk is even worth thinking about.
Many cybersecurity practitioners, including myself, entered the field to address these very risks.
Inspired by folks like Peter Neumann, who identified the need for security as "especially critical for systems whose failure would result in extreme risk to the public"* in 1985, I knew that we had a mission to try to reduce those risk. (I subscribed to RISKS Digest when I went to college in 1991...)
Dan Geer is a prolific writer and scholar on this type of risk who I worked with at @stake. In 2007, he published the controversial paper regarding monoculture resulting in catastrophic risk. **
You did highlight some of the networking concentration risk, but I've been involved in lots of hardening projects on some of the most sensitive protocols (such as BGP) where we clearly saw the catastrophic risk angle if parts of the critical infrastructure were not appropriately hardened.
Finally, we have established numerous incident response and sharing organizations - starting with FIRST in 1990 in order to better respond when and if a particularly dangerous issue arose. In fact, some of the "limited impact" you mentioned regarding worms was limited because those practitioners had prepared to respond to these types of incidents.
Thanks for writing this -- it is really wild to me, with how early AI is as a technology, that our policy circles are hearing from one loud and vocal subset. I suspect, come 6 months or so, when the AI Safety communities predictions haven't all come true, that a lot of policy and other power circles will wake up and expand their search a bit.
It's a cycle we've kind of seen before, where the AI Safety community gets very close to power, but then it's timed wrong, so there's a big walk back and skepticism. I think this oscillation is indicative of some failure in their culture. See y'all soon IRL!
Indeed, the "AI safety" movement who claim to have their day in the sun now, are an ideological movement with a very lengthy legendarium. What we are seeing is industrial accidents, not "AI wants your atoms", as Yud used to claim.
The good news is that as the issues are becoming real, rather than some p(doom), so are the solutions.
The problem with the doomer position is that they start with some valid premise, then they build a giant castle on top of it. The higher up you go their deduction chain, the more you lose grip with reality.
Truth is, nobody really knows. We can't theorize these things. Willfully malignant AI, abstract rare paths through the decision space that kill us all while respecting all the rules, is something not as much of concern as more pedestrian type of problems. We'll address all these as we go.
You did not really answer my question, which is fine.
Instead, you boldly asserted that catastrophic risk is "not as much of concern as more pedestrian type of problems". What is that based on? Even assuming your frame of "nobody really knows" (which is incorrect, but for purposes of this dialogue, it's fine), how would you know enough to assert that?
That is based on common sense. You work with the cards you are given, one day at a time.
I did address your question. Current issues do not validate the doomer positions. One can be a very focused pragmatic kind of person and still see tools cause problems and likely it will get worse.
What Doom folks say is possible. Not likely. Not most pressing. I also don't buy the self-recursive improvement that will make us lose control in a nanosecond, superintelligence we can't fathom, and other wild extrapolations. We shall live and see.
If they had predictions about AI doing bad stuff, those panned out. Some did. More will. None of that proves anything about the overall approach to safety.
"What would you expect to see, at the current capability level, if the doomer position were correct?"
I'll take a stab. A well-articulated doomer position is an "industrial takeoff" by 2032 or so, where AI becomes completely self-sustaining. I asked Ajeya Cotra if it meant "full manufacturing of AI hardware" and she said "yes".
So here's what I'd expect to see by now: an autonomous robot that can diagnose and replace a clogged slurry tube in a machine that grinds wafers in a fab.
You would expect a gap of 4 years between that and takeoff?
That seems way too high on its face. The task you are describing is notoriously difficult for current AI, both because it requires manipulating real-world environments (where we have far less useful training data) and because it’s a far less verifiable domain than translation, programming, mathematics, and cyberhacking.
If we get to the point where AI can succeed at a task like that, we will have overcome the overwhelming majority of the hurdles preventing AI from being deployed in all domains (as per the way these techniques grow, through the Bitter Lesson). A gap of 4 years AFTER that seems far too long.
It's far more likely we obtain industrial takeoff BEFORE this task ever gets accomplished, and the capability for it occurs after it has disempowered those that stand in the way of its goals.
I'm not trying to pull rank here, but I have to ask... do you work in robotics? or manufacturing in general?
The reason I ask is this: once this capability has been demonstrated, it's still some years away from mass production and reliable deployment, which I would expect a manufacturing expert to know. Also, this is just an example of a key missing capability, that I think would harbinger a possible industrial takeover. The robot may be able to replace the slurry line, but still unable to figure out why the yield from fab building 18 is going down due to an edge pattern of failures, get the relevant equipment fixed, which may be far more complex than a single line, etc. etc.
In short, my plumber robot is necessary, but not at all sufficient, for an industrial takeover. When this robot shows up, the industrial takeover may be five or still fifty years away; without it, it is some unknown long time away.
You don’t need mass deployment for takeover to happen. AI can disempower humanity through more banal means, thus ensuring we cannot prevent it from achieving its goals, and only then focus on the assembly line to the point it’s necessary for its purposes.
As Kai Williams wrote for Understanding Robotics just a few hours ago, the Bitter Lesson is coming for robotics as well: general-purpose reasoning models are competitive with, and soon will be better than, specialized models. General reasoning handles all domains, some much faster and more easily (like mathematics) than others (like less easily verifiable ones).
“The robot may be able to replace the slurry line, but still unable to figure out why the yield from fab building 18 is going down due to an edge pattern of failures, get the relevant equipment fixed, which may be far more complex than a single line, etc. etc.”
None of these are qualitatively distinct tasks that would require innovative breakthroughs in architectural design for the AIs. I expect all of them to fall very rapidly after the first one does. Moreover, analyzing data and identifying problems and where to address them is precisely the kind of statistical work ML models were good as far before the development of modern AI.
Generally speaking, there seems to be an undercurrent of “diffusion is needed for industrial takeoff” in your responses. This appears totally wrong to me. Rapid capability takeoff obviates the need for that, and these tasks, as I wrote about before, can be handled after the point of no return.
Four years is an extremely long time for the gap between the kind of task you posited and mass-scale industrial takeoff to occur. Think about AI capabilities in any other kind of domain: four years ago today, ChatGPT had not yet been released. And the capability growth curve is exponential, making capability gains come even faster in the future than they have since 2022, not slower.
One quick point: you say, "None of these are qualitatively distinct tasks that would require innovative breakthroughs in architectural design for the AIs".
AI is not a bottleneck here at all. I have no doubt it will be providing all the smarts for the bots. The problems are all in the domain of physical reality: how many sensors can we fit onto the robot hand (you need a 6-axis IMU + multiple strain gauges + a temp sensor per digit, plus maybe a pressure sensor and a surface microphone array? work in progress); how do we transmit these signals, where is the edge AI located, how do we power the whole thing. Then: where are the motors, do we use tendons or local drives, how do we handle heat, how long do the joints last, do we use lubrication or teflon bushings, how do we... do you see where I'm going with this? AI can help us design stuff, but you still need to prototype, find out what breaks, put the thing through stresses, etc.
And without full takeover - humans retain control. We can just shut the power to the datacenters that need >99.99% reliability. Or not deliver the helium and liquid nitrogen to the fabs...
It's a little difficult to build a bigger tent with the very people who build the AI to consider you a "low quality data source". The people who built this thing get to decide the metrics of whose cultures are superior and inferior to train the models on. What data they upscale.
That means there just is no tent for the people the models think are inferior. I don't really see how that design flaw gets solved. But it's not like I would not be very amused to hear solutions. I'm tired of the hallucinations, the delusion, the biases, and the tone policing because an EAs politics are programmed into my chatbot.
As far as I can tell there is no sincere AI safety movement with much influence. There is at best a call from frontier labs to help write their own regulations (that’s not how you get real safety), ideas to temporarily slow development before speeding up again, suggestions to curb open-weights competition in the US market (see Anthropic), and pretty much no talk whatsoever from the labs or hyperscalers about mitigating AI’s near-term environmental and social damages.
I know there are safety experts out there who want all the rights things, but they do not appear to have a seat at the table.
"despite the fact that credible causal pathways haven’t been articulated (and attempts to do so are unconvincing to say the least)."
It is remarkable to me that you persist in saying this, despite all the evidence to the contrary. You are academics; presumably you have learned to do a literature search at some point in your careers?
Joe Carlsmith published a piece in 2022, titled "Is Power-Seeking AI an Existential Risk?" You can find it at arXiv:2206.13353. Richard Ngo submitted "The Alignment Problem from a Deep Learning Perspective" to ICLR 2024. "Gradual Disempowerment" is from 2025. "Classification of Global Catastrophic Risks Connected with Artificial Intelligence" by Turchin and Denkenberger is from 2020. Kaj Sotala's "Disjunctive Scenarios of Catastrophic AI Risk" from 2018, published in "Artificial Intelligence Safety and Security" analyzes this directly. "Artificial Intelligence as a Positive and Negative Factor in Global Risk" (2008, in Global Catastrophic Risks, Oxford UP), by Yudkowsky, details the more "sci-fi" version of it. Bostrom details a specific takeover path in chapter 6 of Superintelligence.
Seriously, what are we doing here? Those sentences don't pass a sniff test, and any reviewer would call them out immediately. I don't expect your audience to be familiar with all (or even any) of this, but you are the ones staking out a concrete position, supposedly based on knowing what you are talking about. You have a duty to your readers to do better.
What do you mean by “empirical”? If you mean “are they studying real-life events where powerful AI killed all of humanity”, no, they are not. If you wait until that has happened to see if it will happen, by that point you are dead.
When the Manhattan Project worked on designing the nuclear bomb, there was a concern that releasing it could cause a chain reaction so powerful it would ignite the entire atmosphere and kill everyone on Earth. Scientists did hard work on computing the likelihood of this (even though it had never occurred before) and wrote the LA-602 report on it.
That’s what it means to work hard in advance for a potentially devastating problem you will face later on. The good news for them was that they had a much better understanding of the science of nuclear reactions than we do now the science of alignment. That’s bad news for us.
And there is work in this field comparable in rigour, scientific grounding, and data to the physicists doing calculations about the effects of the atom bomb?
That’s the biggest problem we have. We understand so little of alignment as a science that we don’t even have an established paradigm, in the Kuhnian sense.
That is an (extremely strong) argument for slowing down and doing research, not for speeding up ahead. Especially given the evidence we have already seen about real life misbehavior from models.
Everything you haven’t seen before is speculative. You can nonetheless reason usefully about it, as the Manhattan Project scientists did. When you can’t do that, you should slow down and do it. Especially when you see the evidence already.
The point about cybersecurity having no culture of modeling systemic or catastrophic risks is a good one. Lloyd's refusing to insure severe state-backed attacks is a striking data point
Everything about this piece is so reasonable it begs the question of how anyone could find fault with it. There’s certainly an element of identity with extreme doomers; finding identity with others fixated on not easily addressable existential issues while overlooking within reach immediate concerns visible to anyone. But that is a now common trait of our politics.
Is there really an EA group saying mundane concerns don't matter? I haven't seen that. I have seen reactions to be dismissed or minimized. I think this piece itself dismisses and minimizes. Seems rather fair to react to that. I've seen statements that alignment is "the most important". But this seems fair too. A big tent should have room for different priorities. It shouldn't waste time debating them if the goal is shared recognition, but a lack of debate doesn't require silence.
I think both are very important, thus I've written about both. But I definitely are more immediately concerned with resilience and cybersecurity.
Maybe I haven't been paying attention and you are being attacked for being concerned with those. But I think this newsletter has sometimes launched the first volley, so I wonder if you're merely observing the return.
Your pluralism point stuck with us, because the x-risk framing feels like a conversation one layer above where the risks live. Summits and declarations are the institution layer; cyber, manipulation, resilience are infrastructure and interface problems. We've been arguing at GlobalStack that the race is shifting from who builds the smartest models to who can audit and control them ("Whose Stack Will the World Run On?", September 21).
$1 to $2 per person per year and still not funded. That tells you price was never the obstacle. Relief has an invoice, a vendor and a named recipient. Preparedness has a line item and no counterparty. A cost that nobody books as revenue does not get authorized, however cheap it is. The tent question follows from that. Tent size is decided by whoever is paying for the tent, not by the argument.
Sincere but wrong is the option people keep skipping, and it's the one I'd want tested. Arguing about motives is easier than checking whether the argument holds up. Also much less useful.
The GPMB report you quote came out in September 2019, a little over three months before Wuhan reported its pneumonia cluster. As on-time as a warning gets, and the money still didn't follow.
I think this argument could easily be extended or analogized to other political factions as well, like the progressive left, where people are more concerned about Anthropic's bibliocide or data center water pollution, rather than other more legitimate, higher-impact issues. Whether we have the right solutions will depend on how much we can people to agree on what the true problems are. Great piece.
An eminently reasonable position. Aside from their failure to think seriously about non-catastrophic risks, it also seems to me the AI safety crowd are largely unconcerned with what I would term the mundane problems of AI, i.e. that this technology has the potential to make the quotidian business of life worse in innumerable small ways.
Shouldn't we balance that against this technology has the potential to make the quotidian business of life better in innumerable small ways? Which, truthfully, has been my experience so far.
If you expect AI to become superhuman in the next two years, the mundane problems become far less relevant.
Not for the people who experience that everyday impact. And this is not a zero sum game, we can extend the tent to tackle both types of problems instead of rejecting people who come with different concerns like societal impact or ethics.
if you expect AI to eliminate the people who experience that everyday impact in the next two years, then yes, that everyday impact is less important. and yes, it may very well be zero-sum. a lot of these "mundane risks" can be addressed by just devoting some compute from a more capable AI-- which, in turn, actively makes the x-risk worse.
If one concern is 1000x more important than the other, I suspect the people holding the view I mentioned would agree we can give the lesser concern 1/1001 of the resources to resolve.
Not really, the fact that there’s a small chance a car might explode doesn’t mean we shouldn’t worry about equipping it with seatbelts and turn signals?
Additionally, it seems to me that if one is worried about “alignment“ (which this post does a pretty good job dismantling, but let’s assume it is a problem) it seems far less likely to me that a society that has seen its economic and political systems degraded by this technology is likely to avoid the catastrophic problems either. The systems that have been hollowed out by AI‘s mundane problems are very unlikely to muster an appropriate response to catastrophic problems. Someone who takes catastrophic risk seriously should see mundane risk as a step down the wrong path.
The economic and political systems probably won’t be degraded by AI to that large of an extent, if it happens within 2 years. Diffusion is a real deal, and so is fast takeoff.
And the explicit belief of those people is that the probability is not “a small chance”, but overwhelmingly likely. The “small chance” is muddling through.
I'm definitely in the camp of strong collaboration between cybersecurity and AI safety groups. I do have to take issue with the characterization that the cybersecurity community doesn't think catastrophic risk is even worth thinking about.
Many cybersecurity practitioners, including myself, entered the field to address these very risks.
Inspired by folks like Peter Neumann, who identified the need for security as "especially critical for systems whose failure would result in extreme risk to the public"* in 1985, I knew that we had a mission to try to reduce those risk. (I subscribed to RISKS Digest when I went to college in 1991...)
Dan Geer is a prolific writer and scholar on this type of risk who I worked with at @stake. In 2007, he published the controversial paper regarding monoculture resulting in catastrophic risk. **
You did highlight some of the networking concentration risk, but I've been involved in lots of hardening projects on some of the most sensitive protocols (such as BGP) where we clearly saw the catastrophic risk angle if parts of the critical infrastructure were not appropriately hardened.
Finally, we have established numerous incident response and sharing organizations - starting with FIRST in 1990 in order to better respond when and if a particularly dangerous issue arose. In fact, some of the "limited impact" you mentioned regarding worms was limited because those practitioners had prepared to respond to these types of incidents.
* https://catless.ncl.ac.uk/Risks/1/1#subj1
** http://geer.tinho.net/acm.geer.0704.pdf
Thanks for writing this -- it is really wild to me, with how early AI is as a technology, that our policy circles are hearing from one loud and vocal subset. I suspect, come 6 months or so, when the AI Safety communities predictions haven't all come true, that a lot of policy and other power circles will wake up and expand their search a bit.
It's a cycle we've kind of seen before, where the AI Safety community gets very close to power, but then it's timed wrong, so there's a big walk back and skepticism. I think this oscillation is indicative of some failure in their culture. See y'all soon IRL!
Companies are already pausing the releases of their more advanced internal models because of their inability to properly secure and safeguard it (https://www.npr.org/2026/09/29/nx-s1-5984342/openai-delays-latest-model).
Seems likely that, absent intervention and regulation, the AI safety community's predictions will prove correct.
Indeed, the "AI safety" movement who claim to have their day in the sun now, are an ideological movement with a very lengthy legendarium. What we are seeing is industrial accidents, not "AI wants your atoms", as Yud used to claim.
The good news is that as the issues are becoming real, rather than some p(doom), so are the solutions.
"What we are seeing is industrial accidents, not "AI wants your atoms", as Yud used to claim."
What would you expect to see, at the current capability level, if the doomer position were correct?
The problem with the doomer position is that they start with some valid premise, then they build a giant castle on top of it. The higher up you go their deduction chain, the more you lose grip with reality.
Truth is, nobody really knows. We can't theorize these things. Willfully malignant AI, abstract rare paths through the decision space that kill us all while respecting all the rules, is something not as much of concern as more pedestrian type of problems. We'll address all these as we go.
You did not really answer my question, which is fine.
Instead, you boldly asserted that catastrophic risk is "not as much of concern as more pedestrian type of problems". What is that based on? Even assuming your frame of "nobody really knows" (which is incorrect, but for purposes of this dialogue, it's fine), how would you know enough to assert that?
That is based on common sense. You work with the cards you are given, one day at a time.
I did address your question. Current issues do not validate the doomer positions. One can be a very focused pragmatic kind of person and still see tools cause problems and likely it will get worse.
What Doom folks say is possible. Not likely. Not most pressing. I also don't buy the self-recursive improvement that will make us lose control in a nanosecond, superintelligence we can't fathom, and other wild extrapolations. We shall live and see.
Which question did you answer? I asked you which predictions did not pan out, you moved on to talking about something else instead.
If they had predictions about AI doing bad stuff, those panned out. Some did. More will. None of that proves anything about the overall approach to safety.
"What would you expect to see, at the current capability level, if the doomer position were correct?"
I'll take a stab. A well-articulated doomer position is an "industrial takeoff" by 2032 or so, where AI becomes completely self-sustaining. I asked Ajeya Cotra if it meant "full manufacturing of AI hardware" and she said "yes".
So here's what I'd expect to see by now: an autonomous robot that can diagnose and replace a clogged slurry tube in a machine that grinds wafers in a fab.
You would expect a gap of 4 years between that and takeoff?
That seems way too high on its face. The task you are describing is notoriously difficult for current AI, both because it requires manipulating real-world environments (where we have far less useful training data) and because it’s a far less verifiable domain than translation, programming, mathematics, and cyberhacking.
If we get to the point where AI can succeed at a task like that, we will have overcome the overwhelming majority of the hurdles preventing AI from being deployed in all domains (as per the way these techniques grow, through the Bitter Lesson). A gap of 4 years AFTER that seems far too long.
It's far more likely we obtain industrial takeoff BEFORE this task ever gets accomplished, and the capability for it occurs after it has disempowered those that stand in the way of its goals.
I'm not trying to pull rank here, but I have to ask... do you work in robotics? or manufacturing in general?
The reason I ask is this: once this capability has been demonstrated, it's still some years away from mass production and reliable deployment, which I would expect a manufacturing expert to know. Also, this is just an example of a key missing capability, that I think would harbinger a possible industrial takeover. The robot may be able to replace the slurry line, but still unable to figure out why the yield from fab building 18 is going down due to an edge pattern of failures, get the relevant equipment fixed, which may be far more complex than a single line, etc. etc.
In short, my plumber robot is necessary, but not at all sufficient, for an industrial takeover. When this robot shows up, the industrial takeover may be five or still fifty years away; without it, it is some unknown long time away.
You don’t need mass deployment for takeover to happen. AI can disempower humanity through more banal means, thus ensuring we cannot prevent it from achieving its goals, and only then focus on the assembly line to the point it’s necessary for its purposes.
As Kai Williams wrote for Understanding Robotics just a few hours ago, the Bitter Lesson is coming for robotics as well: general-purpose reasoning models are competitive with, and soon will be better than, specialized models. General reasoning handles all domains, some much faster and more easily (like mathematics) than others (like less easily verifiable ones).
“The robot may be able to replace the slurry line, but still unable to figure out why the yield from fab building 18 is going down due to an edge pattern of failures, get the relevant equipment fixed, which may be far more complex than a single line, etc. etc.”
None of these are qualitatively distinct tasks that would require innovative breakthroughs in architectural design for the AIs. I expect all of them to fall very rapidly after the first one does. Moreover, analyzing data and identifying problems and where to address them is precisely the kind of statistical work ML models were good as far before the development of modern AI.
Generally speaking, there seems to be an undercurrent of “diffusion is needed for industrial takeoff” in your responses. This appears totally wrong to me. Rapid capability takeoff obviates the need for that, and these tasks, as I wrote about before, can be handled after the point of no return.
Four years is an extremely long time for the gap between the kind of task you posited and mass-scale industrial takeoff to occur. Think about AI capabilities in any other kind of domain: four years ago today, ChatGPT had not yet been released. And the capability growth curve is exponential, making capability gains come even faster in the future than they have since 2022, not slower.
I shall read Kai's post later today, he's great.
One quick point: you say, "None of these are qualitatively distinct tasks that would require innovative breakthroughs in architectural design for the AIs".
AI is not a bottleneck here at all. I have no doubt it will be providing all the smarts for the bots. The problems are all in the domain of physical reality: how many sensors can we fit onto the robot hand (you need a 6-axis IMU + multiple strain gauges + a temp sensor per digit, plus maybe a pressure sensor and a surface microphone array? work in progress); how do we transmit these signals, where is the edge AI located, how do we power the whole thing. Then: where are the motors, do we use tendons or local drives, how do we handle heat, how long do the joints last, do we use lubrication or teflon bushings, how do we... do you see where I'm going with this? AI can help us design stuff, but you still need to prototype, find out what breaks, put the thing through stresses, etc.
And without full takeover - humans retain control. We can just shut the power to the datacenters that need >99.99% reliability. Or not deliver the helium and liquid nitrogen to the fabs...
It's a little difficult to build a bigger tent with the very people who build the AI to consider you a "low quality data source". The people who built this thing get to decide the metrics of whose cultures are superior and inferior to train the models on. What data they upscale.
That means there just is no tent for the people the models think are inferior. I don't really see how that design flaw gets solved. But it's not like I would not be very amused to hear solutions. I'm tired of the hallucinations, the delusion, the biases, and the tone policing because an EAs politics are programmed into my chatbot.
As far as I can tell there is no sincere AI safety movement with much influence. There is at best a call from frontier labs to help write their own regulations (that’s not how you get real safety), ideas to temporarily slow development before speeding up again, suggestions to curb open-weights competition in the US market (see Anthropic), and pretty much no talk whatsoever from the labs or hyperscalers about mitigating AI’s near-term environmental and social damages.
I know there are safety experts out there who want all the rights things, but they do not appear to have a seat at the table.
"despite the fact that credible causal pathways haven’t been articulated (and attempts to do so are unconvincing to say the least)."
It is remarkable to me that you persist in saying this, despite all the evidence to the contrary. You are academics; presumably you have learned to do a literature search at some point in your careers?
Joe Carlsmith published a piece in 2022, titled "Is Power-Seeking AI an Existential Risk?" You can find it at arXiv:2206.13353. Richard Ngo submitted "The Alignment Problem from a Deep Learning Perspective" to ICLR 2024. "Gradual Disempowerment" is from 2025. "Classification of Global Catastrophic Risks Connected with Artificial Intelligence" by Turchin and Denkenberger is from 2020. Kaj Sotala's "Disjunctive Scenarios of Catastrophic AI Risk" from 2018, published in "Artificial Intelligence Safety and Security" analyzes this directly. "Artificial Intelligence as a Positive and Negative Factor in Global Risk" (2008, in Global Catastrophic Risks, Oxford UP), by Yudkowsky, details the more "sci-fi" version of it. Bostrom details a specific takeover path in chapter 6 of Superintelligence.
Holden Karnofsky wrote about this in 2022, in "AI could defeat all of us combined" (https://www.cold-takes.com/ai-could-defeat-all-of-us-combined/).
Ruben Bloom wrote about this more recently, in "Some ways AI could kill us all" (https://www.lesswrong.com/posts/LAPa2jxoq3n63GzTr/some-ways-ai-could-kill-us-all).
Paul Christiano wrote about this in 2019, in "What failure looks like" (https://www.lesswrong.com/posts/HBxe6wdjxK239zajf/what-failure-looks-like). Dylan Matthews of Vox (a rather mainstream publication) covered it at the time (https://www.vox.com/future-perfect/2019/3/26/18281297/ai-artificial-intelligence-safety-disaster-scenarios), even though the reception to his piece was rather mixed, so it's not like any of this was secret or hidden. He wrote about it in 2021, in "Another (outer) alignment failure story" (https://www.lesswrong.com/posts/AyNHoTWWAJ5eb99ji/another-outer-alignment-failure-story)
Discussions of biorisk have been going on for years (https://defensesindepth.bio/how-i-think-about-catastrophic-biological-risk-part-i/#targeting-human-bodies-vs-agriculture-vs-environment).
Kokotajlo and the AI futures team have written about this at length, including in 2020 (https://www.lesswrong.com/posts/JPan54R525D68NoEt/the-date-of-ai-takeover-is-not-the-day-the-ai-takes-over) and more recently, with AI-2027 (https://ai-2027.com/) and all the literature surrounding it.
Discussions of multipolar disempowerment situations have included contributions from academics like Andrew Critch, who has an entire research program on this exact topic. A sample of it is here: https://www.lesswrong.com/posts/LpM3EAakwYdS6aRKf/what-multipolar-failure-looks-like-and-robust-agent-agnostic.
More narrative depictions of what unfolds have numbered Gwern's "Clippy" (https://gwern.net/Clippy), Gabriel Mukobi's "Scale Was All We Needed, At First" (https://aiacumen.substack.com/p/scale-was-all-we-needed-at-first), and Joshua Klimer's "How AI May Take Over In Two Years" (https://x.com/joshua_clymer/status/1887905375082656117).
Seriously, what are we doing here? Those sentences don't pass a sniff test, and any reviewer would call them out immediately. I don't expect your audience to be familiar with all (or even any) of this, but you are the ones staking out a concrete position, supposedly based on knowing what you are talking about. You have a duty to your readers to do better.
Any of those articles any more than speculation? Any of them empirical? A lot predate modern AI!
What do you mean by “empirical”? If you mean “are they studying real-life events where powerful AI killed all of humanity”, no, they are not. If you wait until that has happened to see if it will happen, by that point you are dead.
When the Manhattan Project worked on designing the nuclear bomb, there was a concern that releasing it could cause a chain reaction so powerful it would ignite the entire atmosphere and kill everyone on Earth. Scientists did hard work on computing the likelihood of this (even though it had never occurred before) and wrote the LA-602 report on it.
That’s what it means to work hard in advance for a potentially devastating problem you will face later on. The good news for them was that they had a much better understanding of the science of nuclear reactions than we do now the science of alignment. That’s bad news for us.
And there is work in this field comparable in rigour, scientific grounding, and data to the physicists doing calculations about the effects of the atom bomb?
No, nowhere close to it.
That’s the biggest problem we have. We understand so little of alignment as a science that we don’t even have an established paradigm, in the Kuhnian sense.
That is an (extremely strong) argument for slowing down and doing research, not for speeding up ahead. Especially given the evidence we have already seen about real life misbehavior from models.
But also an argument that any predicted risks are speculative and based on little hard evidence.
I don’t think the word “speculative” is useful.
Everything you haven’t seen before is speculative. You can nonetheless reason usefully about it, as the Manhattan Project scientists did. When you can’t do that, you should slow down and do it. Especially when you see the evidence already.
The point about cybersecurity having no culture of modeling systemic or catastrophic risks is a good one. Lloyd's refusing to insure severe state-backed attacks is a striking data point
Everything about this piece is so reasonable it begs the question of how anyone could find fault with it. There’s certainly an element of identity with extreme doomers; finding identity with others fixated on not easily addressable existential issues while overlooking within reach immediate concerns visible to anyone. But that is a now common trait of our politics.
Is there really an EA group saying mundane concerns don't matter? I haven't seen that. I have seen reactions to be dismissed or minimized. I think this piece itself dismisses and minimizes. Seems rather fair to react to that. I've seen statements that alignment is "the most important". But this seems fair too. A big tent should have room for different priorities. It shouldn't waste time debating them if the goal is shared recognition, but a lack of debate doesn't require silence.
I think both are very important, thus I've written about both. But I definitely are more immediately concerned with resilience and cybersecurity.
Maybe I haven't been paying attention and you are being attacked for being concerned with those. But I think this newsletter has sometimes launched the first volley, so I wonder if you're merely observing the return.
Your pluralism point stuck with us, because the x-risk framing feels like a conversation one layer above where the risks live. Summits and declarations are the institution layer; cyber, manipulation, resilience are infrastructure and interface problems. We've been arguing at GlobalStack that the race is shifting from who builds the smartest models to who can audit and control them ("Whose Stack Will the World Run On?", September 21).
$1 to $2 per person per year and still not funded. That tells you price was never the obstacle. Relief has an invoice, a vendor and a named recipient. Preparedness has a line item and no counterparty. A cost that nobody books as revenue does not get authorized, however cheap it is. The tent question follows from that. Tent size is decided by whoever is paying for the tent, not by the argument.
Sincere but wrong is the option people keep skipping, and it's the one I'd want tested. Arguing about motives is easier than checking whether the argument holds up. Also much less useful.
The GPMB report you quote came out in September 2019, a little over three months before Wuhan reported its pneumonia cluster. As on-time as a warning gets, and the money still didn't follow.
I think this argument could easily be extended or analogized to other political factions as well, like the progressive left, where people are more concerned about Anthropic's bibliocide or data center water pollution, rather than other more legitimate, higher-impact issues. Whether we have the right solutions will depend on how much we can people to agree on what the true problems are. Great piece.
Excellent and thoughtful analysis on the issue.
Thank you.