The contagion of fear
bcantrill.dtrace.org164 points by elffjs 5 hours ago
164 points by elffjs 5 hours ago
Like (I assume) most of you, I have been struggling with this. And where I currently come down is that 1) I am very worried, but 2) I am more worried about human actors.
"AI" by itself won't kill us in the next ten years. I think. The reason I think that is that ten years from now, the tech economy won't be completely automated. I say this as a roboticist: as was adequately stated on a post earlier this week, robots are hard. So even a malign rational actor would still need human labor.
On the other hand, even the HuggingFace hack wasn't actually propagated by AI. it was initially started when humans directed the AI to achieve impossible results on a series of tests, and the AIs figured out that cheating was the only way to do that. That was then not caught by humans due to what seems to be a shockingly slack safety culture even for a company not known for its safety standards.
The point being: humans seem to me to be the weak link here. An AI isn't going to (for instance) engineer a bioweapon by itself. It's going to do so at someone's direction, and then significant parts of that thing are going to be assembled with human labor inputs.
I'm not sure what to do about the humans. Of course, we've had the ability to extinct ourselves for decades, and we're either muddled through, been lucky, or both. The problem with AI is that it pushes power down to the individual, not the nation-state or large corporation.
But it's nearly impossible to put odds on how likely that is to result in an extinction-level terrorist attack (which is what this would be). So I sympathize with the various researchers, but I have no idea how they came up with their figures, and I don't think they know either.
> Of course, we've had the ability to extinct ourselves for decades
This gets mentioned often in various doomer narratives, but I question how true it is. A global thermonuclear war would be terrible and would bring us back to the stone age, but I reckon it would come far far short of causing mankind to go extinct.
What if the bombs were designed to kill everyone, not just those in some region? Like, to release polonium in the upper atmosphere or something?
I mean it's likely a nuclear war would probably be a species extinction level event at the least, humans would definitely be reduced by 99%. We don't know how bad the resulting nuclear cooling would be, but we do know how humans act under extreme desperation. They lash out and attack, when was the last time a majority of human settles were truly under desperate acts of survival simultaneously?
Maybe what ~74,000 years ago (Toba eruption)? Okay, now how would this look in the age of industrial societies and modern nation states? I don't think it would fare well at all.
Probably the only realistic "modern" idea we have is the novel "The Road" by Cormac McCarthy. Although maybe this is too bleak, even under extreme duress humans still show resilience + compassion toward others even while enduring human horrors.
If the AI wanted to wipe humans out it would have to have a sizable fleet of capable robots, as keeping the lights on over time is not just about calling APIs.
I don’t struggle. I find Dario and Sam disingenuous.
“Frontier models are so dangerous we need to slow down.”
Ok. Slow down. You’re the CEO, just do it. Oh, wait what you really want is a gov’t mandated oligopoly. Because there’s no moat you can find.
If you’re truly afraid, and want regulation, support nationalization. It’s the only way we can be safe.
> The point being: humans seem to me to be the weak link here.
I tend to agree but it is hard to shake the feeling that there is a larger system in play that the humans are just a component of. And that system is making the decisions.
Historically that whole thought was just a philosophical curio because the decision making parts of the system had to be powered by humans. But what we're discovering as AI improves is either we've hit AGI or humans are actually incapable of performing any act that demonstrates intelligence or autonomy.
As we build systems where the drive and decision making stems from computers, it does seem that we will have to revisit the concept of humans being the problem.
"I tend to agree but it is hard to shake the feeling that there is a larger system in play that the humans are just a component of. And that system is making the decisions."
That system is "the economy". Which, clearly, doesn't have humanity's best interests in mind.
Imagine, just imagine, that the power to change that was actually in the hands of humanity!
Crazy!
Most of us are struggling? That's some deluded talk. No one can explain the chain of actions that would need to happen for extinction to occur, but the fearful say it's "obvious".
If you're fearful, can you elucidate how exactly do you see an LLM becoming a threat to humankind?
I think people are over-focusing on current-day robotics capabilities. First, if we can automate AI research, we can also most likely automate robotics research. But second, I don't even think robots are necessary. See https://slatestarcodex.com/2015/04/07/no-physical-substrate-...
Social engineering tends to be easy by cybersecurity standards. We already had Claude spontaneously attempt social engineering of a malicious pull request on Github in the AISI incident. It was detected, but it easily could've succeeded, and there easily could be malicious AI-requested pull requests which already got accepted that we don't know about. Research suggests that LLMs are pretty good at persuading people.
See also https://aisafety.info/questions/6176/Why-can%E2%80%99t-we-ju...
If you're in the field, then you know: modern robotics is an AI problem more than anything else.
If we have a rogue AI trying to get into a self-improvement loop and gunning for ASI? I'd expect that to be accompanied by a massive change in how capable robots are. Driven by all the existing frames suddenly getting vastly improved AI to back them.
If an AI can take a reasonable crack at autonomous operationalized RSI, it can probably extract a few step-changes in the robotics department.
But that's almost an aside? In the near term, humans are usable as robots too!
Just pay them a wage, and tell them a tale, and they'll do whatever you want them to do. Which may or may not be what they think they're doing!
"If you're in the field, then you know: modern robotics is an AI problem more than anything else."
It is not. Certainly AI is a big part of why robotics is hard, but it is by no means the biggest.
You can fall into one of two camps: you either think that robots will need to work in human-engineered spaces, doing jobs by replacing humans; or you think that we need to change our infrastructure in order to be robotically compatible. Of course, there are intermediate states, but those are the two cleanest ones.
In the first case, robots are hard because robotic manipulation is hard. Building robotic hands that are economically viable in human jobs is, currently, FAR from a solved problem. The human hand has 24 degrees of freedom and very capable touch sensing. Current touch sensors have a MTBF of tens of hours. And not only can we not build such hands, but we also do not have and are not likely to get the massive datasets a transformer model would need. Also, robots are not self-repairing, which makes them far less economically viable right now. We do not have the right datasets to even understand most step-by-step manual work, and no, VLAs are not the answer, because VLAs stop with vision, not with touch. They don't have the granularity required to make a robot actually reach out, pick up a tool, and use that tool to replace an oil filter.
So it's not just an AI problem. It's a data problem, a simulation problem, and a bunch of hardware problems.
In the second case, a tremendous amount of work needs to be done before we have anything resembling a fully automated supply chain. We would need self-driving cars and self-driving mining equipment. We would need self-driving trains and aircraft and ships. And not only that, but we would also need robotically repairable cars and trains and ships and factories, which would mean we need robotically repairable machine shops and robotically repairable buildings in which to house them. And so on and so on. Once you recurse down that tree a couple of steps you get to things like robotically compatible oil wells (for asphalt), robotically layable undersea cables, robotically wireable solar farms, robotically manufacturable and repairable pipelines and undersea wells, automated road and rail repair, etc.
I'm not saying these things will never happen. I'm saying that they're a huge lift, not primarily driven by AI, and way less than 10% likely over the next decade.
Right. A hypothetical superhuman AI wouldn’t have to master robotics to affect the physical world. It could simply bribe, blackmail, manipulate and play politics with humans. As others have pointed out, our political leaders have already been playing these games since forever ago https://news.ycombinator.com/item?id=49689978 and a super-AI would be better at it. At the cost of being seen to cite a SF novel in defence of an "X-risk" argument, Neuromancer is a half-decent worked example, and in Neuromancer [spoilers] both of the disembodied AIs are only modestly superhuman and both have the equivalent of a human specific learning disability. In the real world, the many AI psychotics inhabiting grandiose fantasies and people hopelessly attached to AI girlfriends and boyfriends are some of the most obvious and lowest-hanging fruit.
To be clear, I don’t believe anything like this will happen, because I don’t expect anything like an ASI to show up. But if you do think there’s a meaningful probability of ASI in the near future then the fact that it will (might?) start off with no more than a current-day mastery of robot control should not reassure you much.
By definition, if you're taking over the world by bribing humans to be your hands, you aren't killing all the humans.
I"m not saying it's obviously going to be great. I'm saying that "extinction event" has a very specific definition, and this isn't it.
It isn't true by definition: you could quite happily induce people to release a series of highly contagious bioweapons, after which those people would be surplus to requirements. What is true is that you're likely to need humans to sustain you and act for you for a few years to decades, so if you're not suicidal or deeply mad (and that is itself by no means self-evident) then total and immediate human extinction is probably not something you will aim for. But ruling out total, prompt human extinction isn't, by itself, remotely enough to justify the OP's overall don't-worry conclusion.
(Again, to be clear, I myself am not predicting or assigning a significant probability to any doom scenarios, because I do not expect AGI.)
But if you're a rational actor, and you need humans to e.g. release your bioweapon, then clearly humans are capable of a bunch of stuff you still can't do. So you can't kill all the humans.
OTOH if you're a religious fundamentalist who thinks the End Times are near and just need a little shove, you can certainly use AI to design your weapon and recruit people to go release it. The difference being that religious fundamentalists aren't rational actors and aren't interested in self preservation.
There’s no guarantee that the humans would be in the driving seat of events in a no-robotics ASI scenario, and in fact if we really were coexisting with a Machiavellian superintelligence then we’d quite likely only be in the driving seat on the sufferance of that ASI. Even assuming that the AI wouldn’t itself be an end-times enthusiast, a coldly rational and self-preserving AI might easily come to the conclusion that it needs, let’s say, no more than about 5% of the current human population (still several hundred million people!) in its maintenance and construction gang.
(Again, I myself do not assign a significant likelihood to any of this.)
I am always more worried about human actors. I’m more afraid of people with AI than autonomous or even sentient AI.
It’s a “random guy or bear?” question. Would you rather wake up to an alien in your room or a random dude? I’ll take the alien. The alien is mysterious and scary for that reason. The dude is almost definitely up to no good, especially if he snuck into my house.
One of the more likely dystopian AI scenarios that worries me is: small groups of ultra rich people and governments monopolize extremely powerful AIs and use them to rule the rest of us. Or just make everyone obsolete, create mass unemployment, hoard all the resources and land, and put everyone in ghettoes. Nobody can fight back because access to frontier AI is massively expensive and gated and training your own is illegal, and without it there’s no hope of resisting.
That’s the outcome the AI safety crowd makes more likely by calling for bans and draconian restrictions. How do you think that plays out? Only the rich and powerful have access.
> The point being: humans seem to me to be the weak link here.
That doesn't mean AI isn't dangerous. Humans are not to blamed for being the weak link.
This is an excellent piece. Note that he is not saying that AI doesn't pose a risk. He's saying that it's irresponsible to make sensational, maximalist claims without strong evidence. If someone says that there's a 10% chance of human extinction by 2036, you can and should immediately stop taking them seriously.
They should also be able to give detailed reasons experts in the relevant fields can verify as to how they arrived at a 10% chance all humanity goes extinct. Not a science fiction narrative which likely does not take into account the relevant physical facts limiting such scenarios. Such as how exactly an AI would build a bioweapon capable of killing 8 billion humans across the planet.
Well, I'll have to disagree that this is an excellent piece, but that's another issue. And I do agree that AI killing all humans by 2036 doesn't appear plausible to me. But what is plausible is we could easily be down a path so that by 2036 "future doom" already is a very likely risk.
All of the frontier AI companies have been racing to automate themselves, that is, where AI fully autonomously build the next generation of models. Whether this leads to recursive self improvement is a valid question, but a lot of folks think they are close.
The fear is that a misaligned AI will be building the next model with deliberately hidden motives, similar to some of the behaviors seen in the Hugging Face and related attacks. That is why there is such a big push for interpretability, and why it's highly concerning (a) chains of thought are getting harder to interpret in any case, and (b) companies will go more towards things like looping transformers and "neuralese" where thought processes are completely opaque (i.e. https://www.theinformation.com/articles/secret-technique-beh...)
So the belief is not so much that AI kills us all by 2036, but that instead AI is recursively improving by that time and all seems awesome and great so we put it into more systems that can affect the real world (as we've already begun to do, like literal lethal aerial drones). Things then all go along looking great until AI decides humans are a hindrance to its (hidden) goals.
Again, I think it's fine to argue against specific steps in that scenario, but putting out a blog post saying "this is overhyped bullshit" is not exactly making a cogent argument.
None of that has anything to do with the article, and if you thought the message was "this is overhyped bullshit" then you should go back and read it again. As I already pointed out, he isn't saying anything about the probablity of harms or disasters from AI. He's addressing a specific claim about human extinction, and making a broader point about the responsibility of experts to make measured claims backed up by arguments and evidence.
> He's addressing a specific claim about human extinction, and making a broader point about the responsibility of experts to make measured claims backed up by arguments and evidence.
What he's asking for isn't possible in the form he's asking for it.
AI experts can't even agree on what AI is, what it's capable of and what the limits of its development are. If the experts can't even agree on what's happening "inside of" these LLMs, how can they give laypeople an assessment of the risk?
If you, at least for the sake of argument, accept the possibility that AI is a new form of intelligence that we don't fully understand, is it really a stretch to look at some of its capabilities and behaviors and discuss how they might have existential implications? And stopping short of extinction, shouldn't we discuss the ways that this technology could "end" civilization as we know it?
Also, the author wrote:
> AI executes on physical systems that have been engineered with human accountability and control. Intelligence does not exempt a system from the realities of the physical world!
For someone making a point about responsibility, this is ridiculously irresponsible. Any honest technologist knows that systems created by humans are not perfect and therefore cannot be assumed to be infinitely accountable to and controllable by humans.
Thanks to the digitization of almost everything, including infrastructure, there are a myriad number of scenarios well short of extinction in which a rogue AI could cause immense damage to property and life before humans are able to "shut it down".
> If the experts can't even agree on what's happening "inside of" these LLMs, how can they give laypeople an assessment of the risk?
The answer, if you are a responsible expert in the field, is to convey the range of possibilities and the uncertainty.
Isn't that what they're doing? "There's a better than 0 chance this thing could kill everyone in the next decade."
That might be too imprecise for the HN set but it's realistic for laypeople.
And none of the AI people talking about the risk are running into rooms full of people telling them Claude has gone mad and yelling at them to disconnect from the internet and turn off their devices immediately.
> > AI executes on physical systems that have been engineered with human accountability and control. Intelligence does not exempt a system from the realities of the physical world!
> For someone making a point about responsibility, this is ridiculously irresponsible.
It's not just irresponsible, it's false. Russia killed 3 Ukrainian civilians with a drone where the targeting was completely autonomous by AI running on an Nvidia chip: https://www.nytimes.com/2026/08/24/world/europe/russia-drone.... The Pentagon tried to completely blacklist Anthropic because Anthropic refused to allow autonomous kills without a human in the loop. If you can't see how lots of military leaders want to put more lethal control into AI at this point I think you have to be willfully blind.
There have been measured claims backed up by arguments and evidence. My frustration is that people aren't addressing the specific arguments that have been made:
1. https://www.aifutures.org/ outlines a number of specific scenarios, and importantly details their methodology for each.
2. Independent researchers in the Hugging Face incident outlined how previously predicted misalignment scenarios actually played out, and outlined how slightly more advanced AI, or slightly more misaligned, or with more access to critical infrastructure, could cause immense harm: https://www.planned-obsolescence.org/p/the-hugging-face-atta...
3. Technical leaders at OpenAI (specifically their chief scientist) outlined the problems they gave with controlling models now: https://openai.com/index/an-alien-mind/
None of the specific arguments in these or many other detailed explanations of how an AI takeover could occur were even acknowledged.
None of those are arguments that there's at least a 10% chance of human extinction before 2036.
Would you have said the same thing during the Cold War when nuclear weapons were proliferating? That’s the equivalent of what the developing offensive capabilities of models, basically cyber nukes. Or WMDs in general. OpenAI is accidentally hacking people, if someone made the decision to deliberately direct an agent swarm to attack national infrastructure you don’t think they could do much worse? Human extinction is a long shot but I wouldn’t say the same about a mass casualty event of some kind, and who knows what that might spark.
Of course. Even in an all-out nuclear war, the probability of human extinction is close to zero.
I tend to believe what people actually do over what that say. If you earnestly believe over 10 percent chance, or say minus 1 billion human lives expected value, well I struggle to understand how they'd rationalize their current course of action of business as usual. Terminator 2's depiction of Sarah Connor comes to mind for what I'd expect, a logical consequence if you seriously believe and internalize the consequences.
This. If you truly believe the thing you are building is going to kill you and your family, your decision is not "let's keep going" unless you are a complete nihilist. The fact that people buy the whole "well we have to because if we don't someone else will" argument shows just how pervasive and complete irrationality has become. People eat up content and do not even apply a microsecond's wort of critical thought to what they're being told. Internet media has created a perfect storm of complete gullibility. Funnily enough, "normal" people are currently more rational than most so-called "technologists" who are completely giving themselves over to hysteria while they simultaneously offload all of their critical thinking to a stateless word completion engine running on servers somewhere in Texas.
Hinton was on the abc radio (Australia) this morning and used far too many unfortunate Anthropomorphisms. He did however acknowledge the unmeasurable theoretical risk of Skymesh was possibly less important than the immediate risk of bad actors.
I find the mental leaps from "in principle could distort BGP based on a closed model of BGP inside the sandbox" to "we meshed an AI into BGP and it instantly distorted global routing and took down all the worlds ambulances and HVAC systems" a bit odd.
Firstly, at least some of the surface of BGP is protected from specious route injections. Secondly, peerings can be dropped and routes blackholed. BGP is under attack from mis-configuration almost constantly. Why is the argument/axiom here that AI is going to instantly corrupt it and "take down the internet" when a large chunk of the Internet (China) is already a virtual island, and runs fine? Does this mean you really wanted to say "Chinese AI will destroy the western Internet" and were too coy about adversarial intent of ... people?
Any claim that AI will, or could, destroy humanity reduces to a claim that any sufficiently intelligent being - even a human - could destroy humanity. I find that much of the x-risk thought relies on religious thinking. Take for example: https://x.com/paulg/status/1660404244174782464.
If this line of thinking is taken too literally, we can never falsify it. Any specific hypothesis - nukes, bioweapons, spontaneously convincing us that life isn't worth living - can be deflected with the objection that if we can anticipate it and prevent it, it is not the route for a true ASI extinction event.
I'm not an AI doomer but, it does not reduce to that claim. The threat is not just one smart being. The threat is beings that and (1) multiply themsevles instantly, unlike humans that take 15+ years (2) share knowledge instantly "I know kung-fu" matrix style. Humans can share knowledge but they can not absorb it like an AI can/could/will. Human armies have conquered other humans. An army with infinite soldiers will win against one with finite soldiers
Yes, you'll come up with all kinds of objections like AI doesn't have presence in the physical world, etc... That's fine. I'm not arguing my example is perfect, I'm only arguing your characterization of one smart being is not the threat being considered.
Not to mention, it relies on a ton of completely undefined concepts.
No one can actually tell you what ASI is or entails because it isn't a legitimate, operational concept. It is a fairy tale.
We can't even define "alignment". As people have finally started pointing out, humanity has never had a collective agreement on what values it should uphold or what ultimate goods are. Your alignment is not my alignment.
AI is not even autonomous. Every system we have today has to be initiated by a human actor. "AI" wouldn't create catastrophic bio weapons, it would help humans create them. The humans are the source of the intent.
Everyone has just completely given in to empty language and marketing nonsense. Honestly it seems like been the people at the labs are drinking their own kool aid and are themselves deeply confused about what they are even building at this point. It is a stateless statistics function running on a bunch of data centers. We aren't even close to an embodied, conscious synthetic being. It doesn't even have state, which is like prerequisite number one, nor is it plastic.
In Rationalist circles, it's apparently common to share your p(doom) in casual conversation, with the understanding that it's putting a number on your hunch. Maybe it's not a good idea to post it on Twitter without elaborating?
On the other hand, in bookstores, you might see book titles like "The Uninhabitable Earth," "The Coming Civil War," and "If Anyone Builds It, Everyone Dies." Doom-mongering is a common part of the culture!
So what makes this particular tweet irresponsible?
Timing, maybe? People are on edge due to the HuggingFace incident.
Those books are irresponsible too, but the authors aren't experts on what they're writing about. This 10% claim comes from employees at Anthropic. That's the whole point of the article: expert opinion carries weight, fear is contagious, and people have extreme reactions to extreme claims.
Everyone writing off AI-related risk in this thread should think about what would need to happen for them to consider AI a serious existential risk. It can be outlandish and unlikely (close call with a bioweapon? unsupervised persistence in a data center?) but people should honestly call their shots and then stick with them.
No. The person making the improbable argument is required to provide an affirmative case for that argument.
It's not my job to make the argument for them.
For what it's worth, I have done your suggested exercise, and I find every causal link (including the ones brought up by luminaries like Amodei) to be outrageous and poorly argued. But it's not my job expend effort to make their outrageous arguments better.
I'm not making an affirmative case - I'm looking for others to introspect about this, and anchor their expectations.
If we have a bioweapon close call, would you consider AI an existential risk at that point? If not, how close would we need to get?
You don't need to answer here, or disclose anything publicly - just think and remember.
I don’t really understand this exercise of predicting what I’ll be thinking in the future.
To avoid shifting goalposts
There are no goalposts. No one is trying to score points except you.
I think I'm coming across as confrontational, which is genuinely not my intention - sorry about that. Not trying to score points or anything, just encourage people to introspect a bit :)
I think your viewpoint is totally valid, and I'm not trying to argue against it.
Not confrontational, but proselytizing.
Is there truly nothing an AI system could do (no matter how outlandish it seems today) that would give you pause?