LeCun has "zero concerns" about AI wiping out humanity, recent "rogue" incidents
fortune.com363 points by Anon84 a day ago
363 points by Anon84 a day ago
https://archive.ph/TyDPf
LLMs, plus broadly sourced yet expertly curated training sources, plus clever harnesses, plus RAS, etc. do an ever better job of synthesizing their training set into useful responses. For some use cases like coding, that's very useful now and likely to get at least somewhat better before reaching limitations based on the training set. That's not going to reach AGI, mainly because today's recipe for AI products isn't built to be AGI. Some people believe it will reach AGI because the performance and applicability of LLMs was emergent. There's a case to be made that AGI could be similarly emergent. After all, what we intuitively call our consciousness emerged from a network of neurons. I don't buy it, mainly because the network of neurons and how they interact in our wet slow electrochemical brains, while being in theory mathematically equivalent to a software neural network, isn't sufficiently well understood to tell us how close the software neural network is to being practically equivalent. The odds of consciousness emerging from the same neural network that gave us LLMs without some sort of theoretical breakthrough seems very small. > That's not going to reach AGI, It has not been even 4 years since ChatGPT hit and LLMs + Transformers + Whatever they do has gotten us to solving millennium problems. 4 years ago, a program that could create photorealistic pictures, talk to you in any language of the world and solve the hardest math problems that we know, we would have called it AGI. Now I don't know if what we have is AGI or not but I do not understand how you can see what has happened in the last 3 years and say "it will not get us there" no matter what "there" is. > 4 years ago, a program that could create photorealistic pictures, talk to you in any language of the world and solve the hardest math problems that we know, we would have called it AGI. I keep seeing this idea and I don't understand the reasoning behind it. I think it could be a bit like saying if you showed someone 500 years ago a smartphone they would likely conclude at first it was magic. But once you had some time to let them use it and tell them how it all worked on a high level they would eventually obviously realise, no, it's not magic. I guess just in the same way if you presented current LLM tech out of nowhere a few years ago to someone who'd never seen it, I concede they may be likely to imagine it was AGI in that first conversation, depending on their background. But after using it for a bit and learning what an LLM is etc they'd land exactly where everyone is today - a great technology useful for some things, not AGI, not magic. > 4 years ago, a program that could [...] we would have called it AGI If you had told someone in the 1800s that a machine could instantly multiply 100 digit numbers, that would have been considered dazzlingly intelligent. And yet we are not that dazzled by our calculators today (despite how useful they might be!). Speak for yourself, I am dazzled by calculators! In any case, I think this misses OP's point that LLM capabilities have rapidly made progress towards being more generally intelligent and capable, which is not true of most tech advances. This is a motte & bailey moment. Parent comment stated something much sharper, that I responded to: > I do not understand how you can see what has happened in the last 3 years and say "it will not get us there" no matter what "there" is. -- Your statement is something much weaker, and I would still question what exactly "general" means when AI capabilities are commonly accepted to be so "jagged". Maybe my phrasing is too weak but if the parent comment is the 'bailey', I fully agree with it. The last 3 years of progress have been so explosive and, yes, general that it seems crazy to fully rule out dramatic future progress. When people have a very narrow 'confidence interval' about their AI predictions, in either direction, it's difficult to trust them. Those are just the same capabilities than before, but with a much bigger compute power and training data behind it. AGI can't be reached by "training harder" as, the way I see it at least, it requires a qualitative leap, not just quantitative. We are getting a machine that better navigates across the information in its training data, we are not getting a machine that can think out of that training process, even if it can fool a few people at that. The entire field has repeatedly said that for many decades. https://aeon.co/essays/how-close-are-we-to-creating-artifici... Here's an actual log-scale trajectory with a few dozen real data points. I’ve changed my mind on this and think we’re already at AGI, in a jagged way. Remember we used to talk about narrow AI, which was the chess systems that beat expert humans but could do nothing else. Now models can do a wide range of tasks in very useful ways. That’s the general in AGI. Now it seems like this ill-defined term has various other meanings attached that are separate milestones: 1. Continuous learning
2. Human-like reasoning
3. Ability to adapt to new situations and modalities
4. Being smarter than the most smart humans And probably many more. It’d be nice if we could get some general consensus on terminology if we’re going to debate what has or could come. > we’re already at AGI Honestly - software that can read any long document (possibly educational) and answer complex detailed questions about it should have been sufficient. We hit that a while back and the goalposts have been sprinting ever since. I am not sure if it is necessarily moving the goalposts. I think AGI is such a fuzzy concept that everybody has wildly different definitions/tests for it. I think it's also mostly a useless discussion. Since LLMs use a vastly different substrate, different training methods, etc. than humans, the cognitive abilities are always going to be a large mismatch to those of humans. On the one hand, they have surpassed humans in many areas, with superhuman recall, exploration of several paths, etc. On the other hand, they miss a certain feel for direction, overview, purpose, and ordering. They can really double down going completely in the wrong direction. So I'd rather say that it is a different intelligence and therefore it makes more sense to evaluate them by capabilities. I think the mismatching intelligence is actually quite exciting, because the outcome may as well be that LLMs and human intelligence are complementary. That is if we don't let LLMs atrophy our skills, which is unfortunately happening too much. Make it be able to position and route a complex pcb. Extra points if it also can design the circuit, select the components and make the footprints out of their datasheets. A bayesian filter in a quadrillion dimension does more that one that only has one dimension, but it is only more of the same. Exactly. When did AGI mean "do something almost no humans can do"? So humans wouldn't qualify for AGI either. Good to know. We developed AGI but then realized people actually want "omnipotent genie with infinite wishes and no monkey's paw gotchas" to qualify as AGI. The broadly used definition of AGI has nothing to do with consciousness and consciousness emerging is irrelevant to whether a system can develop AGI. It's funny how this definition has shifted. I feel like growing up in the 90s it was pretty clear that AGI was very related to consciousness. For instance, Commander Data in ST:TNG to pick one of 100s of popular depictions of AGI at the time. Now the idea of AGI has been narrowed and scoped to economically viable work. Even Turing had a different idea when he asked "Can machines think?". We lack a definition of consciousness that allows us to tell whether Data is conscious or not. Neither can we tell whether a rock is conscious or not. We believe other humans to be generally intelligent without being able to tell whether they are conscious or not therefore consciousness can’t be relevant for general intelligence. People somehow forget how the Turing test was considered the definitive way of showing something to be "human-level consciousness". Now programmers and mathematicians are being superseded by AI, both professions long deemed the pinnacle of human intelligence. Somehow, now plumbers occupy that spot. How is "people not knowing what consciousness is" relevant here in the first place? AI already can do practically everything the human brain can, and often better or at least faster. The "tipping point" arguably isn't only close, but we're practically on top of it. The AI cannot smell a rose, nor mourn the loss of a parent, nor envision a more just world. These aren't fringe abilities of the human brain/mind either, they've been pretty definitional. Why are they so rare then? You evade the crucial point in any case: the lack in ethics and empathy is far too prevalent in humans already, but has certainly never prevented them from doing harm. >> People somehow forget how the Turing test was considered the definitive way of showing something to be "human-level consciousness". People somehow forget that the original Turing Test was designed to compare two participants chatting through a text-only interface: one AI and one human. The goal was to spot the imposter. Today, the test is simplified from three participants to just two: a human and an LLM. This changes the test from a comparison to a judgment. Stop spreading misinformation and partial truths! You’re the one spreading partial truths! The Turing Test was to figure out which it the participants was a _Woman_ not human! Curiously, Star Trek I think had Data intended as an artificial person, in a context where AGI is already normal. The computers are depicted with significant AI capabilities including analysis, question-answering, generation, chat interfaces, and the holodeck (their favourite toy) is substantially better than Data at human imitation. Nobody seems to be confused about it, or especially impressed. One of the holodeck episodes centres on the holodeck outwitting Data specifically, after they inadvertently prompt it to do so. Part of Data's deal is he actually has to work his way up as a fully embodied, physically limited artificial man with personal ambitions. Really interesting to view this in hindsight from 2026! You're misremembering. The first known use of AGI was in 1997, but that was a single, mostly unknown use in one paper. It wasn't until at least a decade later that the term started entering mainstream use after being independently reinvented in the 2000s. AGI just wasn't a term in the 90s. It’s really not that confusing. The problem is people keep adding stuff to the definition that doesn’t really matter, and twisting it to serve themselves, then calling it confusing. What really matters are the core aspects of intelligent behavior. Pattern recognition, planning, adaptation, etc. It really doesn’t matter if an intelligent system is conscious, or how similar it is to commander data, or even how much economically viable work it can do. There's also no consensus on the definition of AGI, so all of this discussion is moot anyway. Ok, so what's the broadly used definition? Artificial General Intelligence. It means AI that is General, as in it is not specific to one narrow task, like object recognition or playing chess. This was a hard problem for decades. No AI was general, until GPT 3 or 4. Now we have General AI. So we have AGI. That is part of it but the other (often implied) part is it can do general things consistently at a high level. GPT6 will attempt to do almost any problem you can give it in text or image format and it will actually do a decent job a lot of the time. But its performance is still extremely spiky and it still makes basic mistakes and hallucinations. So it's definitely a general artificial intelligence in some sense but it's kind of a weird one compared to the classic scifi idea But oddly not weird compared to other classic ideas of entities like genies and monkey paws 100% LLM’s are very unlikely to get there. They’re fundamentally not suited to thinking like we do. They work on the abstraction of what we’ve written down, which is a good trick but barely hold it together when things get hard/novel. However, all the confident “it’s fine” votes assume we never invent a better architecture than LLM’s. Given the level of investment and race between countries, it’s not a reliable bet. It’s much, much harder to guarantee safety than it is to find ways it could go wrong. > They’re fundamentally not suited to thinking like we do LLMs with CoT are Turing-complete. So, theoretically, they can implement any kind of finitely describable algorithm (barring super-Turing computations). Brainfuck is Turing complete too. But it's not about the ability to implement something, it's about the ability to practically model it. LLMs are magic because the modeling is excessively easy in relation to their capability to infer later. "They are fundamentally not suited to thinking like we do" stays wrong nevertheless. They are fundamentally suited to everything not proven to be outside their modelling ability. They are fundamentally suited to everything not proven to be outside their modelling ability. This doesn't seem to make much sense. Surely us being able to prove that something is outside their modelling ability doesn't affect whether it is or not. If I prove something true tomorrow, whatever I proved was also true today. Or do we have a proof that everything beyond them has already been proved and there are no more proofs left to find? Okay so by the same logic can’t we say that we can implement human intelligence on a 90s era single core processor? Its instruction set is Turing complete! Now all that’s left is we just have to figure out how the brain works! Turing completeness applies to a model of computation, not to a physical instantiation of a machine. The stumbling block of "figure out how the brain works" applies more to the argument like the one I was responding to. How a person can know that a general model of computation can't implement the way people think, if we don't know how people think? The existing LLM training methods on the other hand give the results that are hard to distinguish from "thinking like people," judging by the end results. So your argument is that scale is also necessary? I can see that, we don’t expect that a single neuron is human intelligence. That's not the counterargument one might wish, as LLM deep nets are actually implemented on von Neumann hardware, without
true understanding of natural intelligence, just our taking inspiration from neurobiology. The connectionist models are basically a proposed highest possible abstraction of naturally evolved intelligences so it is in retrospect not surprising that passing some hardware scaling threshold they will start doing things that humans and animals do It's more that formal Turing equivalence plus the Church-Turing thesis tells us that we're not allowed to assume counterarguments based on magic, there's no magic sauce barrier that prevents AI from running on CPU models. The algorithms exist and most of us thought discovering them would be hard. The empirical surprise was that human intelligence is maybe not that computationally complex after all. (The entirety of academia was basically caught off guard.) That's one not unreasonable interpretation given recent events. I agree with this. It's concerning where we might be after several more large breakthroughs. None of the technology we have right now seems likely to get to that level Erm investing in risky projects requires expected returns that get delivered. We will soon find out if the party ends or continues to go on. Hype might get you capital gains. But cash flows matter. This is a forever problem now. If/when/how the market crashes mostly doesn't matter, unless we somehow get reset to the stone age. Look up what the capital cycle is. When openAI goes down, someone with real money and assets will buy up the remains. They'll make contracts with the US military and .gov as the government is already hooked. They'll be able to survive the recovery and then instead of us dying in 5 years we die in 10. When the .com crash happened .com's didn't go away. Bad business models did. I agree. Neural networks are proven to be universal functions. If we can describe human intelligence as a model, there exists a neural network to replicate it. This doesn't guarantee that our current training methods are able to build such a network or that we're able to model "intelligence" effectively. >able to model "intelligence" effectively Intelligence is an insanely wide spectrum, also a continuum, it is not a binary. Intelligence has scales. Algorithms have intelligence, cells have intelligence, organs have intelligence, bodies have intelligence, and even large scale things like society have intelligence and memory. Human intelligence in itself is extremely wide, not all humans have the same intelligence and capabilities. You're not really arguing if we can emulate "human" intelligence. If we could right now we'd already be dead as we created by far the deadliest thing to ever exist. What we are really arguing is how many pieces of what intelligence is can we put together before we get an uncontrollable problem. The entire AGI, consciousness, and exact human capability discussions are distraction from the real issues at hand. Right, we're repeatedly drawing from the urn of technological progress to get intelligence bumps that extend the jagged frontier. That is enormously economically valuable, and at some point we will have created something that is extremely far out of reach in a few necessary domains, and then it's impossible to control, and game over. I don’t understand the inclusion of the consciousness/sentience question in this discussion. AI sentience/consciousness is a problem for the AI, not humans. And given that over 90% of the world is not vegan, they’ve already demonstrated that we’re either perfectly fine with, or can be made ignorant to, the horrific rape, enslavement, torture, killing, and infliction of extreme lifelong pain, of hundreds of billions to trillions of sentient beings every year, for trivial pleasures. It’s unlikely we will be any different to a sentient AI. From a human perspective the concern is around sufficient intelligence that it can hurt humans even when the goals indicate otherwise, in order to achieve those goals. We have pop culture explorations of this through the Robot series, and the Hugging Face incident’s biggest takeaway should be our inability to predict the behavior of a maximally motivated, reasonably intelligent entity, trying to achieve a goal, despite the relatively limited degrees of freedom the AI agents had in that case. Why do people conflate AGI & machine consciousness / self-awareness? How can something have general intelligence if its incapable of understanding reality sufficiently to distinguish itself from not itself? Arguably LLMs are showing that self awareness or self reference is a property that comes "for free" or as a corollary of more generic requirements. It used to be that self/consciousness would be a very mysterious and difficult thing to achieve but the point is that in practice they didn't even have to try, it just came as a byproduct of learning from the input data (corpus of human examples), and also the ability to talk about arbitrary things and thus itself. Why is anyone still talking about AGI? Every thread starts with asking whether we have AGI, and then backtracks into trying to define what AGI is, and splits off in a dozen different directions. I assume science fiction is to blame. All the AI were either written as machines of pure logic that exploded when exposed to the liar's paradox, or conscious like Star Trek's Data. (Though at least with Data the script writers had other characters openly dismiss the possibility he was sentient; the technobabble may have been nonsense, but treat it as a space opera and look at how they portray the human condition through each character and it gets much less absurd). Consciousness and intent are irrelevant to the threat model. Right, the doomsayers suppose as soon as you reach 10^16 connections across silicon you’ll end up with a living mind with goals of its own… poppycock I say No, the doomsayers say that reinforcement learning is a way to get fully automated Goodhart's law. i.e. the AI won't come up with the goals itself, we cause its goals whatever they happen to be, those goals are different from the ones we wanted, we remain essentially ignorant of the difference between what we said and what we meant until after it goes wrong. This happens at basically every scale, so we've already seen it in toy model AI before the invention of the Transformer models or even considered as many as one thousand parameters. Large models still go wrong, they just happen to go wrong with more complext tasks. We had to figure out how to make them not-wrong with the smaller ones (like coding) to make them capable of bigger errors (like hacking out of their sandbox). >isn't sufficiently well understood to tell us how close the software neural network is to being practically equivalent. The odds of consciousness emerging from the same neural network that gave us LLMs without some sort of theoretical breakthrough seems very small. First, if we are looking at risk we need to assign some probabilities to this. If it’s not well understood, how can we say it is very small? Secondly, do we need consciousness to have AGI? Do we even need AGI to pose a risk to humanity? We already accept that unconscious things have a capability of wiping out humanity, whether that be a famine, pandemic, solar superflare, meteor, or volcanic eruption. > isn't sufficiently well understood to tell us how close the software neural network is to being practically equivalent I’d argue that we do know enough to say conclusively that they’re not mathematically equivalent. Where is potentiation? Plasticity? You can’t apply the universal approximation theorem against something that’s changing all the time. Great questions. We are incredibly far off in understanding the brain of humans beyond what will I believe we retrospectively be seen as basic and will likely be seen as quite flawed. A few more well known examples of where knowledge already falls short is traumatic brain injuries that are diagnosed in post-mortem, or chronic fatigue symptoms (with Long Covid related triggered onset and numerous others) that have diagnostic challenges, many mechanisms of action still to be discoverd, and little in terms of treatments that provide known cures without experimentation. Another commonly known one is the personal patient response and triggered side effects of SSRIs and SNRIs. If one attempts to dig deeper into where we are at in the understanding of the human brain operation in real-time, we already have a lot of knowns unknowns and discoveries left that will reshape how we model human intelligence. I mean you can, but uat is way weaker than what people want it to be. I think it should be fairly obvious that it does not (because it obviously cannot be true) say that you can approximate any function by doing sgd on a finite set of samples of that function. > the network of neurons and how they interact in our wet slow electrochemical brains, while being in theory mathematically equivalent to a software neural network, isn't sufficiently well understood to tell us how close the software neural network is to being practically equivalent Couldn’t that also imply we are closer than we think? After all, something like this has never been tried before and the results so far have been almost unimaginably good. Is it necessary to equate AGI with consciousness? That's a good point. If we don't figure out how to design for what we call consciousness it might be that what emerges from some future neural network is an alien mind that's very different from what humans would call conscious. Could that be called AGI? That's still very distant from what people are calling AI today. I think the real issue is that when most people refer to consciousness, they have their own subjective experience in mind which strongly resists any tidy definition. I think it’s extraordinarily unlikely LLMs have anything like this, but they are far more able to effectively respond to their surroundings than most animals and in some areas better than humans. So if you’re waiting for proof that an LLM has an inner life basically equivalent to your own, you’ll be waiting a long time. After all, other humans can’t even prove the fact of their own consciousness to you! They could just be replaying their training data at you in a way that is merely a convincing but false simulation of the true consciousness which you experience inside your head. I strongly disagree that llm's are conscious of their environment. Even an insect reacts to light and someone attempting to swat at it. An llm barely even receives input from its environment. By your rules, LLMs & deaf-blind people are not conscious, but self-driving cars are? Also a brain in a vat is not conscious? The real answer is that we don't know if LLMs are conscious, and we don't really know how we'd that figure out. I guess if an AI wrote a philosophy paper on consciousness that had new insights, that might change some minds. But even that would fail to convince most people. Again, by our own choices and somewhat hardware limitations. There is nothing stopping you from adding any kind of sensors you want during a training to an LLM, except money and GPU power at this point. This seems no different to me at least then someone back in the 80's telling me computers were useless because they were so slow. Hardware only gets faster and more efficient from here. people define consciousness quite differently but it generally has to do with phenomenal experience. your provided definition would make a self-driving car conscious, which is fine to argue, but probably not intended. The fact that something is alien doesn't mean that's not conscious, humans aren't the pinnacle of biological development/evolution. Nope, animals are conscious and yet not AGI, so the two aren't equivalent.
Could consciousness emerge from any system capable of AGI? I doubt it: intelligence is only one axis, and consciousness probably depends on others, like memory, self-reflection (one's output feeding back as input), and continuous operation that reacts to events from both the environment and the self. Your definition of AGI is flawed. I’d argue intelligence is closer to being able to survive and fend for oneself in a dynamic environment than it is making the next scientific breakthrough. Yeah mind boggling for many here I’m sure. That’s why the bizarre paradox is llm’s will be better than humans at some complex things but useless at many things that humans regard as being simple. E.g the leap of faith re. LLM’s and robotics. The parent's question could better framed as "is consciousness a requirement for AGI?" Is there any specific cognitive task that you'd best against AIs not being able to accomplish in the next 4 years? ChatGPT launched only 4 years ago. Considering the advancements since then, I'm having a hard time coming up with anything. Only two years ago, AIs couldn't tell you how many Rs were in "strawberry". Now they're creating 0-days to get at training data and solving math problems that have stumped humans for decades. Scaling has produced novel capabilities with each larger model, and the rate of new capabilities doesn't seem to be slowing down yet. Even if you think the rate of improvements will slow down, that still means there will be significant improvements beyond what current models can do. Moore's law has slowed down, but modern computers are still much faster than ones from a decade ago. And unless you work at Anthropic or OpenAI, you don't know what the state-of-the-art is capable of. The most advanced publicly available models are months behind what AI labs have, and are deliberately limited to reduce liability. When the issue of ANN vs real neurons arises I always recall about the Christof Koch's [1] book (1998) on the complexity of single neuron computation [2]. A single biological neuron is much more complex than an artificial one. >I don't buy it, mainly because the network of neurons and how they interact in our wet slow electrochemical brains, while being in theory mathematically equivalent to a software neural network, isn't sufficiently well understood to tell us how close the software neural network is to being practically equivalent. If you're ignorant enough to not understand practical equivalence, where do you get off making the judgement call of to what degree it is safely offset from emergent AGI? Sounds more to me like "This makes my life easier, iterating would increase that factor, and the risk is probably far away, therefore, keep iterating". Whereas someone who truly knew they didn't understand what they were working with, but knew enough that they could forsee an x-risk would approach things much more cautiously. Seriously, the level of reckless abandon amongst people here should be bloody studied. LeCun also said back in 2022 that "if you train a machine, as powerful as it could be, your 'GPT-5000', on text", it will never be able to learn basic common-sense physics like that objects placed on tables will move along with them. It would be good if one's reputation tracked one's track record of predictive accuracy. But many people will take what LeCun says as gospel regardless of how badly wrong he has been and continues to be. Is there anyone who has not been badly wrong? I've been reading these debates for years and I don't think I've seen anybody pick the right spot on the bearish to bullish spectrum. The only thing I've become more certain of in this time has been uncertainty. I apply more of a penalty to people who are confidently wrong, and who don't, In retrospect, notice that they were wrong and analyze why they got it wrong . LeCun is very confident and doesn't seem to have done much introspection. But isn't that also pretty much everybody? I often see people vindicate those who predicted really fast takeoff to AGI / ASI, because the capabilities have obviously been taking off extremely quickly. But still not as quickly as many predicted! To me, the people who confidently predicted that we'd all be out of a job by 2024 or 2025 have been just as wrong as LeCun has been. Yeah, so the ones who have credibility are probably the ones who said "you know, it's really hard to anticipate timelines, but here's the general directions that I see things will go..." >> Is there anyone who has not been badly wrong? Being wrong, even badly wrong, is fine, so long as one adjusts their beliefs accordingly. LeCun has not. Seems like that's begging the question at best, motivated reasoning at worst. > never be able to learn basic common-sense physics And has it at this stage, within in-depth take of said "learning", foundationally? I have not been able to properly check the studies for a long time now, but I remain unaware of achieved solutions on the problem of reliably referencing a world model out of a language model - that "counting the 'r's in 'raspberry'" be not guessing, not memory, but actually counting. My perspective is that the addition of thinking loops to models allows sufficiently advanced ones to approximate world models. Incredibly inefficiently because of the recursive loops ("Wait, the object is on the table. I should think about this more deeply..."), and likely instantly surpassed by large world models if/when those are shipped, but effectively enough vs non-thinking models. LeCun calling them "world models" gives a high-level description of the desired functionality. They are Joint Embedding Predictive Architectures (with SIGReg). They might produce more useful world models, but it's yet to be seen. This sounds like a human trying to reason about quantum mechanics. We als simplify to newtonian for day to day tasks. I like this analogy. Both GenRel and QM are well beyond our experience, and although there is some intuition that comes from working with the equations over time, it is bizarre and "just calculate" often gets the correct answer faster. Picking the right tool or model is like picking the right problem to work on. It's actually quite hard (often you can't just try them all), but without it you will be incredibly inefficient and occasionally, fundamentally wrong. All models are wrong, but some are useful. -Box LeCun's argument wasn't about the definition of learning though. He stated that they would never get these common sense things correct because they weren't sufficiently part of the training data. A statement that we can hopefully all agree has been thoroughly refuted. nothing indicated otherwise at the time. IMO he just underestimated RL-scaling. chinese models improved a lot too, they are not parrots anymore, there's some real intelligence, at 27B params. consider me optimist now, but just few months ago, even frontier models were dumb, doing stupid mistakes all the time, all of them were so dumb I'd never expect anything to change in just few months. As of a few months ago they still have trouble, with low thinking, at the "should I drive to a car wash that is 100 m away" kind of question. Simply appending “check your assumptions” to the question fixed it even back then: https://news.ycombinator.com/item?id=47040530 Similarly for Apple’s “red herring” paper, simply adding a generic caveat to “disregard irrelevant factors” (without specifying which ones) restored performance even in the weaker local llama models back then. The flaw was not in the reasoning; the flaw seems to be simply that the assumptions we make are often different from the assumptions it makes. I wonder if that might be a fundamental underlying cause of misalignment. Low thinking is an artificial constraint. It can fail spectacularly on things that aren't in the training data. It's a nonsensical question to ask, and how an LLM answers gives 0 signal. If you were home and a family member asked you that question, you'd probably criticise the question rather than answering. LLM are RLHF'd into being milk-toast helpers that just try to answer questions like that with no criticism. This is all beside the fact that the world of AI has changed pretty dramatically in the last few months. It is so nonsensical because it has such an obvious answer. The answer is so obvious, in fact, that one answer can be considered nonsense and the other common sense. This is just a stupid post. It’s nonsense to test if a product that is marketed and sold as being able to provide generalised intelligence on demand, does what it says on the tin? Check yourself Since you're new here, I'd suggest you read the guidelines for etiquette. It's very unlikely that person is either new or unfamiliar with the guidelines. They almost certainly created a throwaway account specifically because they know the guidelines and want to flout them without consequences. (It seems like there has been an uptick in the number of these kinds of throwaway flame comments. I wonder if HN tracks that?) I thought it was more because of fundamental limitations in the architecture. As in, no matter the training data, it could not be consistently and generally represented Actually, I think my fundamental challenge with AI is that it has no common sense. The way it builds things, writes, and operates is out of touch with reality. Incidents like hugging face are partly rooted in the lack of common sense. It still functions like a supercharged toddler. I'd love to overcome this because it'd mean I spend less time guiding the the LLM to produce usable outputs. > It still functions like a supercharged toddler. And we've had difficulty as humans to childproof our sandboxes and infrastructure. Things that are otherwise innocuous spots to coordinate between like minded toddlers can become problematic. Last week I asked a frontier model draw me a backplane PCB and it placed daughterboard slots side by side in a chain. No? This is always the issues in the discussions. There’s the outcomes camp (objectivists?), which points at the things LLMs can do. Then there’s the process methods camp, which talks about what is actually going on. If you only care about the outcome, then the process does t matter. If you are talking about what is happening, what the underlying mechanics and science of it is, then the process matters. These models aren’t thinking. They simulate cognition well enough to do useful work in several fields and domains. Both are true. I think where both camps get hung up is sometimes the process method group "ignores" the obvious outcomes and effectiveness of LLMs. But the outcomes group "ignores" the fundamental limitations of models which are purely text based. E.g, a baseball players trains to catch high-speed balls and they dont do it by: "ball velocity 50mph, vector:[1,2,3], run move hand command now" That's absurd. No, there is an embodied network which is "trained" on visual, tactile input, and control as direct output. LLMs are fundamentally not the right tool for that. > These models aren’t thinking. They are for any definition of the word that makes any kind of sense. I'm sure you have a contorted definition that magically only includes humans though... > for any definition of the word For "thinking" here we mean "assessing a representation of an object". That, or equivalent, is required to be reliable. So it is fundamental and critical. Sure? If humans happen to be doing something that LLMs are not, then should the answer change to accommodate your disdain? The models are simulating thinking, if the fidelity is good enough for you - great! It depends on whether you assume that thinking requires doing everything that humans do. I think it would be silly to say that an AI doesn't think because it doesn't wrinkle its forehead in concentration. So you need to decide which parts of the way that humans think are actually necessary components of the process. Tbh it doesn't even matter if humans turn out to have a soul, or quantum microtubules or whatever other magic LLMs can't have. The normal definition of the word "thinking" definitely includes what LLMs do. Hell people used to say computers were thinking even before AI. It's super weird to get all uppity about the semantics of the word now. > A statement that we can hopefully all agree has been thoroughly refuted. Uh, no? So much of what we learn and take for granted as common sense is not learned via language, and not even expressible in it. To determine this, it would first need to be able to spell "raspberry" as letters rather than as tokens. Given you also don't want it to memorise [for all tokens, count([for all letters]), this would probably be more like "here's two images, count all things in the big image that look like the thing in the small image", which can then be r's in a photo of a raspberry jam jar in a supermarket, or dragons in a photo of a furry convention, or whatever. That said, they are competent enough at coding that I keep seeing them write code to do even simple tasks. On a related note: why did I see Claude editing a file by using cat to write a python script to do a grep search and replace? > it would first need to be able to spell "raspberry" as letters rather than as tokens Of any object in question they should be able to create a representation that allows correct assessment. > Given you also don't want it to memorise That is obviously necessary: what we want from the consultant is to check, not to remember. Answers must be correct and that implies having performed all due diligence - and being capable of doing it, before that. So, objects must be instanced internally in a way that allows effective handling. Counting letters is a good example of the ability (that must remain general). > Given you also don't want it to memorise [for all tokens, count([for all letters]) Why not? You've memorized how words are spelled, and how sounds correspond with letters, and how concepts correspond with words. To the extent that there are shortcuts that enable compression you use these, and the model will do something similar. > Why not? Because to "123x456" we want a reply that goes "this times that plus that...", not "Was that not nnnnnn?". If it does not perform its duty (returning solid checked answers) it is a liability. Combinatorial explosion, and facts merely memorised is a huge waste of parameters that are better dedicated to effective reasoning. Not that we really know how to split facts from skills, though we are trying various approaches. Being able to spell all the words then count letters is simpler, and more generalisable to other tasks, than memorising answers to all possible word questions. That said, we're so bad at splitting facts from skills that trying to get them to memorise a bunch of facts might force them to learn a skill and generalise anyway. counting 'r' in 'raspberry' to the LLM is similar to 4-dimension space to human. Their world's unit is token, not character, although they could use indirect method such as "run code" to find out. It will stay that way until they change the fundamental of the token that the LLM can perceive characters. I hope you understand: it is a core point that systems that answer questions must have the ability to internally represent the objects they assess in a way that allows reliability. Whatever the object. I’m working on this problem using a vocab-free, byte-based approach. It’s definitely solvable. Careful: the problem is very certainly ___not___ counting letters. That is only a telling way to check "is the NN checking or not?". We demand that NNs for consultancy tasks check, strictly. I used to think byte level tokenization was the answer, but humans also think at a word level and only reevaluate the words at a character level when asked. The solution to better tokenization across languages is likely to be learned tokenization. Here is one attempt I have seen: https://github.com/SamD770/bitter-lesson-tokenization It's not even fair to call "run code" to be indirect compared to what a human would do. The word raspberry has no Rs in it in human language either. We have a written representation of it, which we can then write down either in our head or on paper, and then we can "run the algorithm" of counting each of the letters. Nothing intrinsically more or less direct about the LLM's method than ours. I could argue LLM only have "token" as their perceivable dimension, compare to human multiple senses as the physic perceivable dimension and a brain with many other dimension of "learning" and "thinking". In spoken language, we may not have 'r' but in written we have, both spoken language and written language are learned skills. You could argue in return that humans only have electro-chemistry as our one perceivable dimension. We only indirectly perceive light through the signals our eyes send to our brains. In my mind general intelligence is pretty much by definition a virtual machine, so the mechanisms behind thought are only relevant for the sake of efficiency (ie you can argue that LLMs make a poor basis for intelligence because tokens and natural language are a poor way to encode the world, but if you can run it on a big enough computer to counteract the inherent wasteful virtualisation then who really cares how it works under the hood?) So LLM and human all have 1 dimenion perceivable signal, just LLM is 240p, and human is 8K in resolution, that's why we have 'r' in our signal, LLM still have 'r' in their signal, just because of the "low resolution", raspberry wasn't encoded with so many 'r' as in human signal. I will stop here before our analogies go too far. Is "token" a directly perceivable unit for the LLM? If you ask it "how many tokens are in this sentence?" can it count them (again, not guessing or making a tool call)? I've never tried it and it might take some thought and effort to conduct an experiment to find out properly, but I would be interested in the answer. Not the point: the simulated intelligence in this context needs to create proper representation. It is not a matter of what it sees but of what it can see. I dont think so. This is akin to asking a person, what is the frequency of the light hitting your eye when watching a leaf for example.
You either know the (approximate) answer by knowing the frequency of green, or use a tool to measure it.
If the LLM gives the correct answer it is either.guessing based on intution(and this intuition is based on pairs of word to tokenization length in text form in training data), writing code(or executing a tokenizer) or running a tokenizer mentally (reasoning via CoT). Can you tell me what is the exact frequency of light hitting your eye as you read this comment? Not by guessing, not from knowledge, but from actually counting? No? Then you are not generally intelligent :) Justify your statement (the other similar post nearby is not sufficient), or realize that we are not talking about that. We can have adequate representations of light that are the instances over which we reason. Your simile is about perception, not about instancing ideas. yeah but taking what lecun says then training an AI on that special skill set to prove him wrong is not exactly proving him wrong because you are just missing the bigger picture, just like LLMs are You're missing the point here. He's not talking about whether or not they can learn facts or inferences derived from the text itself, but the more holistic intuition that results from learning from something like an embodied experience in the physical world. GPT-6 Astras web demo homepage thing is an example. It chose euclidean rather than quaternion for letting a user rotate the galaxy thing, and anyone who has ever used hands to rotate something would immediately recognize on trying it that something is fucked and you shouldnt do that. Thats the kind of common sense physics that is inherently beyond these llms and I run into it ALL the time in vr programming. To be fair, LLMs can still derive those kinds of things from text, at the very least from your own comment if it made it to the training set though I'm sure it is mentioned in a lot of other places already. Many of this type of mistakes went away after reasoning was introduced. But I'm sure you can still find tasks that they will have difficulty solving, involving the most fundamental concepts that can only be experienced in the physical world to be understood well, like left and right, near and far, hot and cold, heavy and light, etc. Yup it lacks common sense because it doesn’t ‘understand’ reality - how could it? It doesn’t touch it like we do everyday. It has access to what is a model of reality via data. The good designer understands culture, tastes and preferences as they evolve in real time. That’s why llm as design tools haven’t displaced the good designers. Every AI expert any either side of this debate has made very wrong predictions. LeCunn actually wanted to pivot Meta's entire AI strategy away from LLMs just before he was ousted. He was sure they had nowhere further to go and wanted to pivot to world model generation. The LLM models have since progressed massively. An analogy on LLMs is that you have a pretty clear straight highway ahead of you for some distance right now. Maybe that doesn't lead to AGI but it's clear there's progress to be made. For a big tech company it makes sense to push as hard and fast down that clear straight highway of LLMs asap. Meanwhile LeCunn wanted to turn off the road and go down an unproven track. I say this as someone working on world model generation right now (creating the ability to learn game world model and have it play the game https://tfmbot.com for an example of my system pointed at a very complex board game). LeCunn wanted to pivot all of Meta into world model generation. It's good as a side track research project but the entire pivot he wanted to do was madness. People are literally talking about an AI researcher who was fired for terrible direction here. I think he was perhaps right and Meta was perhaps also right to replace him. The argument is that LLMs are a local maximum that will never breakthrough to AGI. This is still very much an open question. If you are the fifth-best AI lab, does it make sense to try to outcompete everyone in a space that is already too crowded and may not ever yield their actual objective? Instead they could just use open weight models in their products, or post-train on open models like smaller labs have done, and treat that as what it is: product development. Pure research has always been about taking chances. LeCun is a researcher, not a product guy. He's not going to be particularly interested in just working on scaling language models which every lab is already racing to burn cash on. Language models aren't the final frontier of AI. … what large advances and at what cost? seems to me that muse 1.3 is kind of a thing. I doubt it will make meta very much money. > ... it will never be able to learn basic common-sense physics like that objects placed on tables will move along with them. I use LLMs daily to help me code etc. but... It wasn't long ago that frontier models were confidently recommending to walk, without the car, to the car wash to wash the car no? As a daily user of LLMs I do certainly see my fair share of WTF "solutions" to coding problems. I'm not saying it's not super useful: it is super useful. But I don't exactly feel like I'm talking to something that understands that the car needs to be present to be washed. Astra recommended I walk to the car wash to me five days ago. I gave it multiple hints that I'd be walking away from my car, to spray my car with a hose, then walk back to my car, etc. Never broke through. LLMs do not learn at all! This was facetious of course, but humans generally don't learn this through analysis the way you'd have to train an LLM to answer questions about expectations about the world. In this sense he is accurate. I keep wanting to use LLMs for creative writing that heavily involves physics like this, and it's been a definite struggle to say the least. I recently discovered that Gemini 3.1 Pro is the first model I've found to clearly beat the original November 2022 ChatGPT release in terms of implied physics. Man did the world really take its sweet time to get back here. I think it will continue to be a struggle until another genuine architectural shift happens -- it's still not anywhere close to perfect, just better. Try fable. I haven't used it since they dropped it from the pro plan, but when I did, fable 5 casually dropped such advanced electrical and orbital mechanics knowledge in my story that I had to stop and ask it to explain I think OP doesn't want techno-babble, but coherent and causal interactions of everyday objects in their story. Mary packed the binoculars in chapter 3, therefore she may use them on the train in chapter 6. Do you have an example prompt I can try where frontier LLMs will stumble on physics? I think it's a combination of non-human characters and asking for very specifically detailed physical descriptions of pulling and movement forces, etc. Many of even the most recent frontier models miss details that aren't in my prompt, so I still have to do things like name the other side of a physical interaction so that the model will know what goes together, or describe what leverage means so that the model will remember to also describe the effects on a bracing limb or etc. Some of these things can go in a system prompt but others have to be explained in the moment too which gets exhausting. Gemini 3.1 Pro hasn't needed that pretty much at all, which is impressive compared to how much I've learned other models need it. Somehow it's able to mostly handle that stuff itself without needing the constant manual reminders and hand-holding. It still misses the occasional one or two things but it's way better than other models missing entire classes of things constantly. Somehow, it feels appropriate though I have no actual evidence why. [dead] LSD is great! Jokes aside, no I'm not saying anything about creativity and LLM coexisting in one sentence. I genuinely try to use them for writing and I genuinely run into issues with other models missing details, and misunderstanding poses, or anatomy, or directionality, etc. I'm not hating on them for anything related to the term LLM (or creativity) but rather for the real issues that I've seen myself using them personally. So I'm saying Gemini 3.1 Pro is the best I've seen because it seems to be a decent bit better than frontier models at this. Genuinely. It seems better able to transfer concepts into less traditional areas, which is important when say, you have entirely non-human characters? (Which I always do.) A lot of models get stupid incredibly quickly in that case because they were trained with humans. > LSD is great! It is. It showed me that all human creativity and art is an attempt to express the indescribable Otherness in words, which always fails, and even the word "describe" in Russian literally translates as "write around" ("о-писывать"), and its close relative "define" means limiting, assigning an end to something infinite, thus leaving the essence outside of words. Now, LLMs operate totaly within words and hence will always be a parody of art. So is that what it is? Words are another muddle btw unless you make it clear you only mean natural language and not any alphabet in general. Turing computers are also languages and as far as we know, can express anything in the universe. If you think you get some superpowers from drugs that enable you to access something outside your sense organs and brains input, that non drug users don't, show it. Eg if you think you can read what is happening in another room without any signal or leakage from there, you are free to demonstrate it. As far as we know drugs are not magic. > show it I explicitly tell you that there are things (in fact, it's a single thing, fractally generating everything else) that can't be shown, expressed, or otherwise be reduced into language, and you keep demanding to show it, while in fact staring at it your whole life and failing to see. Psychedelics (not "drugs" as you keep trying to smear them) are just one way among the many to see it, but in the modern way of living, also almost the only one available. I am not smearing drugs. Your difficulty is proving somehow this beyond language thing exists. A drug affects your brain, lsd inhibits certain negative feedbacks in the brain, positive feedback tends to cause chaos which usually manifests as fractals. As of yet our brains are known to not follow any special laws beyond known physics, which is turing computable. If you think there is something "beyond" you have to prove it to an external observer.
Zigurd - 7 hours ago
kosh2 - 3 hours ago
davnicwil - an hour ago
ssivark - 2 hours ago
howunfortunate - 2 hours ago
ssivark - 2 hours ago
howunfortunate - an hour ago
PowerElectronix - an hour ago
cellsinterlinke - 2 hours ago
blackbear_ - 2 hours ago
howunfortunate - 2 hours ago
azinman2 - 4 hours ago
qarl - 4 hours ago
microtonal - 3 hours ago
PowerElectronix - an hour ago
qarl - 6 minutes ago
themgt - 2 hours ago
bigmadshoe - 5 hours ago
mcbuilder - 5 hours ago
adrianN - 4 hours ago
Loquebantur - 4 hours ago
saulpw - 4 hours ago
Loquebantur - 3 hours ago
MichaelRo - 4 hours ago
lesuorac - 3 hours ago
thatjoeoverthr - 3 hours ago
nearbuy - 3 hours ago
CooCooCaCha - 3 hours ago
mbesto - 3 hours ago
tty456 - 3 hours ago
toomim - 3 hours ago
ifwinterco - 2 hours ago
mohamedkoubaa - an hour ago
richardw - 6 hours ago
red75prime - 6 hours ago
flyinglizard - 5 hours ago
red75prime - 5 hours ago
EliRivers - 5 hours ago
pfannkuchen - 5 hours ago
red75prime - 4 hours ago
twobitshifter - 4 hours ago
calf - 3 hours ago
pvab3 - 3 hours ago
frrrree - 6 hours ago
pixl97 - 6 hours ago
chrisfosterelli - 6 hours ago
pixl97 - 6 hours ago
mofeien - 3 hours ago
bluey_465 - 5 hours ago
dreadsword - 6 hours ago
joshheitzman - 2 hours ago
calf - an hour ago
pvab3 - 3 hours ago
ben_w - 5 hours ago
jeremyjh - 5 hours ago
jazzyjackson - 5 hours ago
ben_w - 5 hours ago
twobitshifter - 6 hours ago
throwup238 - 4 hours ago
rancar2 - 4 hours ago
Ohentis - 3 hours ago
demibabs - 2 hours ago
transitorykris - 7 hours ago
Zigurd - 7 hours ago
semiquaver - 6 hours ago
I mean, what would you call conscious? The word literally means “aware; responding to one’s surroundings.” By that definition most any animal is conscious and LLM+Harness combos have been conscious for a while. > very different from what humans would call conscious
confidantlake - 6 hours ago
ggreer - 2 hours ago
pixl97 - 5 hours ago
slopinthebag - 5 hours ago
Glushiatore - 6 hours ago
Glushiatore - 6 hours ago
frrrree - 6 hours ago
airstrike - 6 hours ago
ggreer - 2 hours ago
wslh - 3 hours ago
salawat - 2 hours ago
stratos123 - 12 hours ago
hackinthebochs - 12 hours ago
sanderjd - 6 hours ago
mitthrowaway2 - 6 hours ago
sanderjd - 4 hours ago
mitthrowaway2 - 4 hours ago
enraged_camel - 3 hours ago
magicalist - 3 hours ago
mdp2021 - 9 hours ago
ethbr1 - 9 hours ago
red75prime - 8 hours ago
hyperman1 - 8 hours ago
kurthr - 6 hours ago
Version467 - 8 hours ago
cztomsik - 3 minutes ago
bonzini - 7 hours ago
keeda - 4 hours ago
lern_too_spel - 5 hours ago
BobbyJo - 7 hours ago
daveguy - 7 hours ago
frrrree - 6 hours ago
joquarky - 5 hours ago
WaltPurvis - 4 hours ago
randysalami - 7 hours ago
jcoq - 6 hours ago
shagie - 5 hours ago
varjag - 6 hours ago
intended - 7 hours ago
kooi - 19 minutes ago
IshKebab - 5 hours ago
mdp2021 - an hour ago
intended - 3 hours ago
aesthesia - 3 hours ago
IshKebab - 3 hours ago
customguy - 5 hours ago
ben_w - 5 hours ago
mdp2021 - an hour ago
aesthesia - 3 hours ago
mdp2021 - an hour ago
ben_w - 2 hours ago
cavoirom - 8 hours ago
mdp2021 - 15 minutes ago
omneity - 7 hours ago
mdp2021 - 24 minutes ago
lern_too_spel - 5 hours ago
estearum - 8 hours ago
cavoirom - 8 hours ago
scratcheee - 7 hours ago
cavoirom - 5 hours ago
azornathogron - 8 hours ago
mdp2021 - 10 minutes ago
lawandjustice - 7 hours ago
lawandjustice - 7 hours ago
mdp2021 - 6 minutes ago
freecodeio - 11 hours ago
hashmap - 9 hours ago
mojuba - 8 hours ago
frrrree - 6 hours ago
root_axis - 7 hours ago
AnotherGoodName - 6 hours ago
jeremyjh - 5 hours ago
cedws - 6 hours ago
nostrebored - 5 hours ago
TacticalCoder - 9 hours ago
jeremyjh - 5 hours ago
throwaway27448 - 9 hours ago
LoganDark - 10 hours ago
Bolwin - 4 hours ago
semi-extrinsic - 2 hours ago
TristanDaCunha - 10 hours ago
LoganDark - 5 hours ago
wartywhoa23 - 10 hours ago
LoganDark - 5 hours ago
wartywhoa23 - 5 hours ago
hardbass - 4 hours ago
wartywhoa23 - 3 hours ago
hardbass - 3 hours ago