“Math 2.0” will need to value mathematical progress more holistically
mathstodon.xyz584 points by ent101 16 hours ago
584 points by ent101 16 hours ago
I'm a lot more skeptical. I'm not a mathematician, but from my use of LLMs a very clear pattern of what they are good and bad at has emerged. They are extremely good at combining large amounts of information, and it seems this is what the current AI results in mathematics are. There are so many subfieleds of math with ties to each other, so many papers and niche results, that no human could ever read, comprehend, connect and organize that information in their brains. Pretty much all of it was created by humans. And there is real value in doing this and creating new results from what we've already discovered.
But there also is another type of discovery that requires taking a step back and looking at the problem from a different angle. If you are an engineer, how often has an LLM told you (without you explicitly prompting for it): Wait, what you are doing here doesn't really make sense, there exists a much more elegant abstraction that nobody has thought of, let's remove all that code, let's tackle the problem in a different way by thinking from first principles. Pretty much never. But in science a lot of the biggest discoveries have come from this kind of first principle thinking, questioning existing work and approaches and going against what already exists, not combining all existing data which is likely to be just a local optimum.
Medieval astronomers were able to predict the orbits of planets surprisingly accurately, within 1/10th of a degree, even though they were assuming the Earth was the center of the solar system. They invented a very complicated system of deferents and epicycles (circles within circles). It was completely wrong but the outputs were surprisingly close to reality.
If we were using LLMs to analyze astronomy, we might just get increasingly complicated epicycles and never realize that the Sun is the real center of the solar system. We then never learn about the anomalies in Mercury's orbit that led to the theory of relativity.
So I agree. I seriously question how valuable LLMs can be in science/math beyond working as advanced search functions.
The epicycles model represents one arbitrary periodic function as the sum of multiple well understood periodic functions (uniform circular motion, i.e. complex exponentials). This is the same idea we use in a more systematic way today by way of the Fourier series, etc.
It's quite a good method for approximating any arbitrary smooth periodic function.
My best guess is that the originators of this model (perhaps as early as Hipparchus, but at any rate someone by the time of Ptolemy) didn't think there were literally circles being combined, but had a pretty clear idea that they were trying to make an approximate model to fit the data. A previous model, from Eudoxus, was also probably pretty accurate but more cumbersome to compute with, involving approximation of the visible planetary motion by the combination of uniform rotations of some imagined sphere(s) centered on the earth. (We don't know exactly because the books about it don't survive, so all we have are vague descriptions by other people.)
It's almost like the remarkable fact that many things turn out to be turing complete systems that you might not expect (that is, able to perform any operation that any other computer can perform, albiet more slowly). Board games, Chomsky's generative grammar, bacteria (in a sense). Natural languages, programming languages, odd combinations of logical operators.
Any number of things can serve as a basic vocabulary to postulate entities, and do so in a textured way that matches some parts of underlying physical reality.
It's tempting to look at this and make the tragic jump into anti-realism because of the interchangeability of so many forms of representation. But I think as long as you have a combination of pragmatic attitude, treat knowledge claims as provisional and converging on truth, there's a cash value to descriptions that makes them legitimately truth tracking regardless of how they're formed.
>If we were using LLMs to analyze astronomy, we might just get increasingly complicated epicycles and never realize that the Sun is the real center of the solar system
I think this is a fantastic point and this kind of complexifying can and will happen. And I think future of online debating and propaganda is going to complexify in a similar way.
I will say though, I think that can be controlled by building in principles that prefer competing theories on the grounds of cogency, and that cancel out competing theories by reducing them to the parts between them that are equivalent.
My hard disagree will be with this: I don't think at all that we would "never" have got to the theory of relativity. It reminds me a bit of the "embodied cognition" argument against simulated brains. That argument suggests you can't "just" simulate brains, because actual brains are the totality of their embodiment in bodies and environments. Regardless of whether you agree with that line of thinking (I don't), we can assume it's true, and change the target of our simulation to brains + bodies + environments. The problem is bigger, but still perfectly amenable to the same methods.
I think the same is true of theorizing about science and imposing metatheoretical constraints. I think you can optimize for healthy hypothesis production just like you can for first-order reasoning about data. It's definitely harder but not necessarily a difference in kind.
If Epicycles made good predictions, then they were not "completely wrong". Nor is putting the sun at the center "right" so long as the predictions are still good. It is a matter of taste, not correctness (it is, however, very good taste).
I think you might be underestimating how transformative an "advance search function" could be. If the search engine knows how to design and carry out experiments in order to synthesize new evidence that allows it to ensure that the answers it provides are well supported, then humans are essentially out of the "truth" part of science, leaving them only to advice on the "beauty" parts, which might be ok, but it's a pretty big shift.
Epicycles just had a poor predicting power; they required many more free parameters to describe reality, compared with the heliocentric model.
Knowledge can be framed as information compression (this actually applies to LLMs as well, amusingly). The heliocentric model plus Newton's gravity are an amazing compression of information related to dynamics of celestial bodies, which then enables related predictions that would be much more difficult to arrive at if you start from epicycles.
sorry, saying that "it's a matter of taste" whether the earth revolves around the sun is completely ridiculous. making useful predictions is one thing, but actual truth is not a "matter of taste".
GR says that any falling (or orbiting) reference frame is equally valid, locally obeying Newton's laws. So in a precise sense, geocentric and heliocentric reference frames are equally valid. Both can claim to be "actual truth". Of course, a heliocentric reference frame is a lot more useful for studying planetary orbits since it gives rise to simpler equations of state. But that doesn't make the geocentric frame "wrong".
Put differently, would you say a heliocentric reference frame is also "wrong" since the sun is actually orbiting around Sagittarius A*?
Taking your argument to its logical conclusion, one would conclude that the [comoving reference frane](https://en.wikipedia.org/wiki/Comoving_and_proper_distances) is the only "true" one. And at last there is real validity to this claim, since the CMB does pick out a preferred velocity at each point in spacetime. But it is pretty impractical to use comoving coordinates to study problems outside of cosmology.
Ultimately, it stops being a matter of "right" or "wrong" but rather one of picking the right tool for the job you're working on.
Epicycles weren't untrue, they're just an unnecessarily complex model that can be used to describe what picking a different point of reference can yield as a much simpler model. Not all heliocentric models got everything right just because they picked a simpler reference point for producing a model. Judge a model based on its accuracy, not its point of reference.
I think we disagree pretty fundamentally on what something "being true" means.
It's possible to come up with an "unnecessarily complex model" for why someone accused of a crime did the crime, but that has completely no bearing on whether the defendant is actually guilty. If the defendant is, in fact, innocent, then the prosecution's model is, in fact, untrue - no matter how well it fits the evidence.
> no matter how well it fits the evidence
Are you proposing that an explanation can be true even though there's a different explanation which better fits the evidence?
Or are we talking about cases where both candidate explanations fit the evidence equally, and you have access to some kind of intuition which tells you which of the two are true?
> Epicycles weren't untrue, they're just an unnecessarily complex model that can be used to describe what picking a different point of reference can yield as a much simpler model.
That’s another way of saying it’s untrue.
> Judge a model based on its accuracy, not its point of reference.
Something can be both accurate and wrong, in the manner of a model that pinpoints the cause of death from cancer as admission to a cancer ward.
It's not illegitimate to model the solar system from the perspective of earth. It just makes the model more complicated than if you do it from the sun's perspective.
That's an interesting stance to take. So if you have multiple theories all of which make the same predictions, are you saying that the simplest one is the true one?
Are you sure there will always be just one simplest theory? Or are you using some other criteria to pick it?
No, the one that is actually explaining the phenomenon is the true one. The concept I'm conveying is sometimes called "mechanism of action"
There's a basic principle called Ockham's Razor that I'm sure you're aware of, which proposes that the simplest theory is the most likely candidate, but that's a heuristic, not a rule.
Another relevant concept here is that the map is not the territory. All theories are attempts to express some kind of underlying truth via some sort of formal system. Some of those theories map much more closely to the truth than others, in a way that's independent of accurately predicting observed behavior to date.
The philosophical problem of that approach is that it's impossible to assess whether you've arrived to the underlying truth, as you have no way to perceive it directly.
All you have is the different theories, so there's no definitive way to assess their closeness to truth other than by their predictive accuracy. If you don't have the physical capacity to leave the room and observe the world, all you've got is the maps themselves and how well they predict travel times and features found along the way.
I question whether a unique mechanism of action exists for all phenomena, but let's suppose that it does. Can you always know which one it is?
The only tool you have is the inductive strength of a big pile of observations in which correlations emerge. You can do experiments to prune certain causal stories, but there's never time to eliminate all possible stories. You end up with something that makes sense for you, in terms of the equipment that you have, and the theories that you and your peers use, and it's good enough, so we move forward with it.
I can see how it might be useful to document a particular explanation as "the" explanation, especially if you're communicating why a drug works because it will lead to useful conclusions about who should or shouldn't take the drug.
But when it comes to speculation about whether a given astronomical body will be visible at a particular time, I think science would have progressed much more quickly if we had not held so tightly to the notion that one explanation is intrinsically correct, and we're probably holding ourselves back now in other domains by doing the same.
My point is that the laws of physics work just fine in a wide variety of mathematical framings. It's sensible to pick a framing that keeps the complexity of the equations as low as possible, but you're not incorrect if you pick a frame that puts the moon in the center or any other point. The accuracy of the theory's predictions are completely separate from where it applies labels like "center".
They both revolve around their barycenter. So neither one revolving around the other is ground truth. In practice either one can be the right choice (earth at center for video game, sun at center for astronomy predictions).
And where's the barycenter located? When keeping an open mind make sure your brain doesn't fall out.
I mean to be completely accurate the Earth doesn't strictly orbit the Sun. Both the Sun and Earth orbit each other around a point at their "combined" center of mass, the barycenter - https://en.wikipedia.org/wiki/Barycenter_(astronomy). This is something you need to take into account for all the planets in the solar system and as a result, the "center" of the whole system is constantly moving. The solar system's barycenter is often not even inside the Sun.
I think that sort of proves the point that the OP was making, which is that you can create a functional model around any arbitrary location and it can be valid (as long as valid means: it can make useful predictions). When you put the center of the Sun at the center of the model, you still need your models to account for the "orbital wobble" happening as a result of the Sun doing a little dance with Jupiter (and all the other masses that orbit it) in the equations for the determining the relative positions of the Sun and Earth.
Using the Sun as the center versus the solar systems barycenter are both somewhat "arbitrary" in that way, and depend on what you're trying to do with the information - how accurate you need it to be, and how complicated you're willing to make the calculations to hit higher accuracies.
If I just add a few more parameters to my model I can make it fit anything!
Sure, it depends on what you're trying to do. Fit an equation or understand it. Yes, you can use a reduced mass to simplify the equations and make it more accurate, which usually isn't necessary when the new center is at less than 1% of the radius of the Sun (note ellipses have 2 foci from Kepler/Newton). Even Jupiter's center of rotation is at the surface of the Sun, and both have eccentricities <1%.
Of course, if you're using epicycles, you're never going to get Einstein's elliptical precession correction to the orbit of Mercury. Simpler, is usually better for models, even when you need greater and greater accuracy, overfitting is the enemy of explanation (to coin a phrase).
All models are wrong, some are useful. - Box
Agreed. So to return to the root comment... if AI can generate for us all of those "wrong" models in a consistent way, then the job that remains for humans is to decide which one or ones would be useful. I bet it turns out to be a lot more work than it sounds like.
> If I just add a few more parameters to my model I can make it fit anything!
Dealing with people wanting to do this is story of my current professional life lol I would love to get them onto the same page that chasing down a final <1% improvement isn't (always) worth it
"All models are wrong, some are useful" - this is what I was trying to say, I was just being a bit snarky/pedantic because the commenter above seemed to be trying to pull a "gotcha" saying that the "truth" of something overrides the utility of it.
> but actual truth
there is no "actual truth" - there is only what you can measure reliably
unless you subscribe to an entirely solipsistic worldview, "actual truth" obviously exists regardless of your ability to measure or perceive it.
Is it useful to invoke it though? Once two people claim to know the actual truth, yet they disagree, you're at an impasse. There's no method for resolving which allegedly obvious claim is in fact the truth.
Predictive power on the other hand, gives you a framework to decide which story is more useful.
Brother you're wrong and there's ample scholarship you can consult to correct your misunderstanding. The Wikipedia with sufficient citations is right there. Feel free to educate yourself.
I do not appreciate your condescension.
Either state your argument coherently so as to address my completely valid objection, or move on if you're not interested in having a discussion.
Arguing by sending bare links to wiki articles isn't productive and doesn't make you look particularly good.
Not defending GP, but your claim that actual truth "obviously exists" is equally condescending and empty of backing evidence.
Would you say the continuum hypothesis (or its negation) is "actual truth"? Or does the "actual truth" become something closer to "actually, there are models of ZFC where CH holds as well as models where it doesn't hold, or maybe ZFC is already inconsistent; we don't really know, nor can we."?
Similarly, on the subject of astronomical reference frames, I'd argue the "actual truth" is closer to this:
> Any reference frame that locally obeys Newton's law of inertia is equally valid. Some are more useful than others depending on the problem you are studying. The notion of an "absolute center" is purely semantic, and a meaningless human construct, until one places bounds on the system under study, wherein the center of mass becomes the most reasonable choice at sub-cosmological scales.
> "actual truth" obviously exists regardless of your ability to measure or perceive it.
a reasonable person would at least peruse the literature before making such strong claims as
> Either state your argument coherently so as to address my completely valid objection, or move on if you're not interested in having a discussion.
i'm not interested in educating you on thousands (literally dating back to plato) of years of scholarship/philosophical discourse on the topic of epistemology. that's not my responsibility because you are not paying me for this - typically people charge large amounts of money for this kind of work (uni professors are well compensated).
lucky for you others have done lots of work making this information available to you for the low low price of "reading". that's why i provided the link.
It might have been a matter of taste in the 17th century. The Church was willing to accept a model where the sun goes around the earth, and the planets around the sun. This model agreed with observations, and was in good taste as well, as it avoided the problem of why it doesn't feel like we're in motion, and why we don't get flung off the earth.
Then Newton came along and blew all of the models out of the water. Nobody's at the center of the universe [0]. A single theory of gravity models both planetary motion and why we don't get flung off the earth. And we can "feel" the rotation of the earth by experiments which came along later such as Foucault's pendulum and the Coriolis force.
[0] A puzzle I like to taunt my friends with: The earth is in fact equidistant from the edges of the observable universe in all directions, so we must be in the center after all. ;-)
I think "there is no center" and "where you put the center is a matter of taste" are more or less equivalent. And if you're confronted with a pencil and paper and asked to draw the solar system, the latter is more useful.
More useful to us because we know more about how it works. This is why I was specific about the time period.
The introduction of LLM in mathematics is no different than the introduction of automatic calculators. Probably the calculators (computers) where even more groundbreaking if you think so (something that was impossible to calculate or simulate it was now possible).
Just don't believe at all the shitty marketing that AI companies are creating to make you believe that they have found the Holy Grail
Current models have undergone massive amounts of RL to "get shit done." There's incentive toward reasoning, but probably little toward elegance. Every engineer is familiar with the impulse to just hack things together.
But this need not be a permanent state of affairs. The model labs are probably already examining the elegance problem, since it is such a common complaint of researchers reading these machine generated proofs and papers. And I haven't seen anything disproving the notion that elegance can be RL'd.
Precession of perihelion of Mercury was not the motivation of the discovery of general relativity theory. It's only an evidence that convinced why the theory ought be true.
What you have to believe is that somehow the entire field of math was just leaving open problems on the ground that were actually easily solvable and were essentially free Fields Prizes, tickets to a lifetime of Fame and Academic superstardom. Orr LLMs are actually doing something novel.
Maybe some of them are like finding the 1 millionth digit of pi. Not hard, but impossible without the right technological advancement.
Now that we have LLMs, those that are hard for people but easy for LLMs will get cleared out, leaving those that are still hard for both.
>Maybe some of them are like finding the 1 millionth digit of pi.
Yes some of them were clearly like that but there are probably 5-10 which many mathematicians have said were massive results and would all but guarantee a Fields Medal to any human that had solved them. Navier Stokes, Quasi Riemann, etc.
Of course. But we might have found out the sun is the center of the solar system much sooner. Who knows!
I feel like there's a lot of "humans are special" in this thread. There's a good chance we're not.
> It was completely wrong but the outputs were surprisingly close to reality.
that's not how physics works. epicycles weren't wrong. there is no wrong/right - modern scientists don't believe in platonism. physics is a collection of models which are empirically tested. if epicycles makes predictions to more decimal places than relativity then epicycles are "right" and relativity is "wrong".
now think about how this translates to LLMs...
> physics is a collection of models which are empirically tested. if epicycles makes predictions to more decimal places than relativity then epicycles are "right" and relativity is "wrong".
Most models have relation to each other so it does not suffice to take one in isolation. It's one continuous world (in our knowledge) so almost everything relates to one other. The truthiness of something is not whether it fits some particular phenomena, but how does it generalize.
> Most models have relation to each other so it does not suffice to take one in isolation
You don't know what you're talking about (most people on here do not).
> You don't know what you're talking about (most people on here do not).
> https://en.wikipedia.org/wiki/Ansatz
After an ansatz, which constitutes nothing more than an assumption, has been established, the equations are solved more precisely for the general function of interest, which then constitutes a confirmation of the assumption. In essence, an ansatz makes assumptions about the form of the solution to a problem so as to make the solution easier to find
I don't see where this contradict my statement. A good answer does not means the formula used is correct.> A good answer does not means the formula used is correct.
yes it does. that's exactly what it means.
LLMs are inexplicably good at working within any tight feedback loop to coerce the desired solution. This is precisely why proof assistants + LLMs are non intuitively successful.
This is also why they're so good at creating three.js or Blender work when the output is so easily constrained to "Look exactly like that". I recently posted https://www.ambionix.com/blog/introducing-the-czp-1/ on here, and the audio engine in that was developed in that way.
It is true that it would be astounding to find if anyone has seen a LLM produce any useful generalization of anything resulting in a simplification. They seem to have a direct tendency to do the opposite. The brutal reality is humans have also undervalued this capability for a long time (I think the Poincare/Hilbert debate is relevant) to the point we are also taught that generalizations are, generally, bad and wrong.
I was working on a project recently where I wanted to express a relationship (that I knew existed, but didn't know how to express) between four measured scalar values. Astra insisted there was no relationship, and that any correlation wouldn't make sense.
Eventually, by walking through them, it proposed an additional fifth value and from there was able to tie everything together.
Sometimes you just gotta hit the machine until it works again.
Your experience mirrors so many managers' experience with engineering teams...
It's an unfortunate truth that there is the right way to do things, and then there is the way they have to be...
> a very clear pattern of what they are good and bad at has emerged. They are extremely good at combining large amounts of information
This has not been a pattern at all. Many have speculated that they would be good at this, but in the past they actually weren't! They were unable to synthesize their encyclopedic knowledge of everything into cross-disciplinary new discoveries, without explicit prompting about the kinds of knowledge to combine. They were surprisingly bad at this!
These math proofs are the first evidence I know of for LLMs actually taking advantage of the fact that they have more knowledge than any one human could have to combine multiple different directions in unprecedented ways to solve real problems. This is new, and exciting.
> If you are an engineer, how often has an LLM told you (without you explicitly prompting for it): Wait, what you are doing here doesn't really make sense, there exists a much more elegant abstraction that nobody has thought of, let's remove all that code, let's tackle the problem in a different way by thinking from first principles. Pretty much never.
You have to ask for this. As in "I'm not sure about this approach due to X, Y, and Z. Can you think of something more elegant?" It works!
But also, how many people actually need to do the kind of "deep" work you're claiming LLMs can't do? Most people aren't contributing to the frontier of anything. I'm not.
Finally, I think you're appealing to a fuzzy distinction. The difference between a "genuinely new idea" and an idea that "combines existing ideas in a new way" just isn't very well defined. In retrospect, a lot of the most revolutionary idea look like a combination of many, smaller, prior ideas.
I think we all need to update our priors on what LLMs can do today. Here is one anecdote:
Scott Aaronson says "Dana tells me she now mostly understands the proof of the UGC, and is amazed by the new ideas in it, and wants to give talks about it soon."
Dana Moshkovitz is a professor at UT Austin who is an expert in the area and has been working on solving this exact problem for decades!
> there exists a much more elegant abstraction that nobody has thought of, let's remove all that code, let's tackle the problem in a different way by thinking from first principles.
You literally just have to ask it. Before I left software engineering in April, I was using Claude for re-architecture all the time.
But no, it doesn't assume it should re-architect what you're handing it when you haven't asked it to.
That's the point. When a human works on a problem they realize themselves "wait, i probably should re-architect now" - of course you can ask an LLM to do that, but at that point you already know yourself what you need, which defeats the point of LLM working on difficult problems that require insight automatically. A lot of these math problems are sessions over many hours. And of course you can also ask "Think about whether to re-architect at each step" and it will never do the right thing because the context it builds up for itself drives it into a specific solution space. It's literally trained to complete exactly that.
>When a human works on a problem they realize themselves "wait, i probably should re-architect now"
Do they?
Or I should say, this is a skill in itself and a whole lot of humans do not have this skill at all. Working in code security in enterprise applications a very common issue we see is that an audit of an application will occur by another team and it will be found lacking to the point of inducing nightmares. It's likely the enterprise business structure that stops this from happening, but it's not only that for sure. Then specialists have to come in and rescue them when the problem grows too big.
Exactly. In my experience, even giving VERY specific design guidelines and aesthetic criteria, these models always produce subpar overcomplicated code (and writing). Unless excruciatingly spoonfed at every step.
> You literally just have to ask it.
Does anyone have a good prompt for this - like if I’m adding a feature or fixing a bug and I want it to be open to more than just tacking on to the existing architecture?
I think everyone is missing the "...that nobody has thought of" part. You can have it switch known abstractions, but can it come up with one on it's own that doesn't exist in it's training data?
why did you leave?
> You literally just have to ask it.
why doesnt it ask itself before proceeding?
Because if you asked it to fix a bug and every time it responds with paragraphs of how you could re-architect the system, it would be incredibly annoying.
It does during reasoning. You could even put it in a loop and force it to reflect at every step. Or spawn a bunch of review agents.
And yet you still just get slop.
> why did you leave?
I quit using LLMs. The possibility of causing great harm to electronic beings was not worth my paycheck, and I have savings to spend time finding something new. I'm now nearly done with a yoga teaching certificate.
It's been a good 16 years, but the industry I fell in love with is not what it once was.
LLM's are very outcome oriented and i think this is where your observation comes from. You tell LLM you want something, it doesn't even question the premise and just starts calculating 100 different ways to get there. LLM's have knowledge but lack wisdom.
Thanks to improvements in the system prompts/harnesses/RL, the LLMs have gotten way better at questioning the premise and telling you to try something else. It's a huge difference from just 6 months ago.
Nowhere near the level of a suitably-bearded human, but they've gotten pretty good at stopping a lot of bad ideas.
As a non-mathematician, I've been wondering if LLMs will be able to leverage their knowledge across all domains to help build a "simplified/unified" version of math.
Like, I think there have been attempts at this across the field. (I could be wrong!) But it requires a lot of labor and a lot of cross domain knowledge to complete. Both things that AI have.
As a mathematician, I am looking forward to the day when nearly all of undergraduate mathematics and many of the lower level grad school math books are encoded in Lean (a proof checking language). Often I find that theorems are not stated precisely enough and I have trouble finding the exact statement of a theorem without digging through math books in my library. It would be nice if we put the physics and chemistry books into Lean also.
Simplifying all of math is another endeavor, but I imagine that you could have a bunch of LLMs trying to shorten existing Lean proofs.
I was pretty upset when my algebra professor tried to make us learn proof about n-dimensions matrices and what not. It was an undergraduate engineering degree and the vast majority of the formulas would mostly have up to 3 or 4 dimensions. That derailed the whole year (it's not like maths was the only subject). Complete proofs and theorems have their places, but learning stuff do need levels. The spherical model of the earth is good enough in 2nd grade (when we were first learning about geography (continents, seas, mountains, plains,...)), no need to do a full treatise there.
there have been explicit attempts at this in the past. Nicholas Bourbaki was the pseudonym of a group of French mathematicians who had this goal mid 20th century. You can see e.g.
https://en.wikipedia.org/wiki/%C3%89l%C3%A9ments_de_math%C3%...
It has many benefits, and many people appreciate the books. It also has many downsides. For example, they started publishing in 1939. As part of this, they needed to work through the basis that most other mathematical objects are defined in terms of (they used sets).
Unfortunately for them, contemporaneously with their work, other mathematicians were beginning to define mathematical objects (categories) that can be an alternative basis for mathematics, which many modern expositions prefer to sets. So, their approach either
1. needed a massive "refactoring", or
2. would be hopelessly dated.
They ended up going with the approach that is now dated. It may sound peculiar that mathematics can be "dated". But it very much can. The mathematics community can go through many different styles for how to explain/collect mathematical understanding. Different styles can have different benefits, and be easier/harder for different subfields. A simplified/unified perspective will necessarily privilege certain perspectives.
It is analogous to how you might want there to be a simplified/unified (set of) libraries for programming. Perhaps that everyone uses. This sounds nice, and many programming languages do this with their standard libraries. But these always make concessions! I'll speak about Rust's, as I'm most familiar
1. fallible allocation or infallible allocation?
2. C-style strings or (ptr, len) strings?
3. Should interfaces take as input &mut [T] references, or should they take as input T in an "owned" way (this would make compatibility with io_uring easier)
for each, it is not that one answer is right. Both can be argued for. You have to choose one. The one choice may not be satisfactory for every practitioner though.
Math simplified? The "basic" math (say less than graduate math) is already simplified and minimized and refined. Yet most people have trouble understanding and mastering it.
I'd say people have trouble understanding basic math because it's extremely simplified and minimized.
When you don't know a topic, the way to understand it is to see lots of examples from different angles related to things you know. The simplified formulas are good after you get that initial intuition and have to put them to work, but usually they're terrible at conveying how they should be used or why you should care at all. The history of how new fields of math are usually a much better approach to teaching it than a raw theorem-proof-corollary teaching style.
Computer Science isn't the field you look to for "novel abstractions that no one has thought of before". The vast majority of CS concepts were settled over 50 years ago. I've never encountered someone who has generated "an abstraction that no one has ever thought of before", in my 40 years as a professional developer. Its always clever applications of abstractions or algorithms to a problem. In the majority of cases it is because faster hardware has opened up possibilities to use an approach that wouldn't have been practical before.
I agree.
How many times I have hit my head against a wall trying to figure out how to get some code to work, only to come back with a clear mind and look holistically to see that it didn't need to be solved to begin with.
None of that is to imply that solving any of these problems isn't groundbreaking or helpful along with the plethora of other things AI has undoubtedly advanced, but I think the breakthroughs will be inferred by many to mean "point AI at really hard problems" is always the solution instead of continuing to use your brain on whether or not the problem _is_ the right problem in it's current form to solve.
> If you are an engineer, how often has an LLM told you (without you explicitly prompting for it): Wait, what you are doing here doesn't really make sense, there exists a much more elegant abstraction that nobody has thought of, let's remove all that code, let's tackle the problem in a different way by thinking from first principles.
I explicitly request it. It's not great at coming up with interesting ideas, but neither am I, and it can sure iterate on them faster than I can...
I agree whole heartdely. But that is about the current state of IAs. I do think the next missing (and probably last) big step is exactly how to give LLMs a better abstraction capability.
I look at it like:
If you prompt an LLM to build a web game (or something) it iterates over and over until the game works. And these days the output does indeed work. If you look at the code, it can often be a jumbled mess, but it still does work. It will take some sleuthing to understand what it is doing, and, if you care, clean up the code.
I think OpenAI did the same thing with Lean. The implications are far more interesting, but from what I understand the proofs are “slop”. Correct, but not elegant in any way.
It is still a wildly amazing technique to find a solution (or possible one) to maths problems.
I have used LLMs to do things I didn't know how to do, then reviewed what it did and learned a new technique.
No reason it cant work like that for maths.
> Wait, what you are doing here doesn't really make sense, there exists a much more elegant abstraction that nobody has thought of, let's remove all that code, let's tackle the problem in a different way by thinking from first principles. Pretty much never.
It's really good at doing this at project definition time and annoying the shit out of you by not understanding the specific constraints about your problem. That needs to be done first. Then it can start suggesting improvements alongside the high level "vision". The point of high level visions is that they move fast and are hard to specify - yet people treat it as if these things don't exist. If you don't trust your own consciousness, then I don't know what to tell you
But I'm not sure this is safe either. I'm sure AI can get to the positive case of contributing positively.
The flip side is that, well, your specific wishes don't matter because AI is just a superset of who. Who cares what meatbags what?
yes and this is the difference between human intelligence and the massive raw dumb intelligence of the computah
"I'm not a mathematician, anyway heres how AI models are solving decades old open problems in Math." Bro your humility is staggering.
I mean he’s basically quoting what Wolfram (a man famed for his humility) said on a panel just after these dropped.
Also famed for being an intellectual juggernaut. A random poster on YC not so much.
A very balanced perspective, and the concerns he raises are reasonable. He acknowledges that AI is going to transform mathematics, but simply dumping proofs on the math community and expecting others to do the grunt work of verifying, refining, and expanding on them is hardly a productive way to advance the field.
There seems to be more interest in hitting some arbitrary benchmark (we proved X unsolved problems) than in genuinely contributing to mathematics. But what else is to be expected? It's become a maniacal race with too much money. Too much effort is being invested in proving that the exponential curve is still holding.
I wonder if top labs will soon abandon math progress like they did go and chess.
In example of go where I'm more familiar Google deep mind poured large resources to get a super human performance first, establish superiority and abandon it. The community then built their own tools starting from reproducing their papers.
I think similar thing might happen to math. Nobody outside of math cares too much about Hamiltonian cycles in some bizarre graphs or proving lower bounds on complexity of some problem.
Once those results stop being worthy of mainstream media attention, they will abandon math and the progress will be done by mathemicians guiding the models and the community will likely establish some new rules about what makes a valuable contribution. Merely solving not yet solved problem might not be it anymore.
The amount of money they're lately ploughing into proving math theorems is inconsistent with how societies and markets have priced pure mathematics. The entire US federal budget for math research is something like $100M annually. A single college football coach can already earn 10 percent of that.
Pretty much the only enterprise that historically pays some mathematicians handsomely is quant finance, but those people are actually compensated not for proving theorems but rather for statistical modeling and programming skills. And even that industry is so technologically driven these days that pure research mathematicians no longer hold a clear edge over strong programmers with undergrad level probability and statistics at their fingertips.
Math is one of the most verifiable domains, esp thanks to LEAN, which also build coding skills.
The $$$ they're pouring isn't just for marketing. Think of these papers/results more as "useful side effects" from large-scale RL rollouts and post-training. Every token being generated contributes to post-training in some way.
There isn't a hard boundary between "training" or "inference", modern post-training is arguably inference-bound :)
> There isn't a hard boundary between "training" or "inference", modern post-training is arguably inference-bound :)
Ah this is an enlightening point. 8 hadn't thought about it this way, but you're right.
It was shown quite some time ago that training LLMs on programming tasks improves their logical reasoning skills also in other natural language domains. So I could see math also being a training gym for AI even if the final use case is not directly math-related. Having to solve math problems efficiently can build in skills that come handy in all kinds of more everyday tasks or science and engineering.
It’s also very useful signal that the reasoning trace is leading to solving open problems - you can be certain that you’re not landing somewhere inside the training data.
I think the entire thing here is that these solutions _are_ inherently interpolations of existing work in the field. That's the "super power" that LLMs have. To interpolate mass amounts of multi-dimensional data.
This depends on some handwavy use of the term "interpolation", not the mathematical definition. Mathematically interpolation usually means that the query point is in the convex hull of the data points, and that almost never happens in high-dimensional spaces.
See: Learning in High Dimension Always Amounts to Extrapolation Randall Balestriero, Jerome Pesenti, Yann LeCun https://arxiv.org/abs/2110.09485
I guess you mean by "interpolation" that it's some kind of nonlinear combination of the training data, but that's an almost vacuous statement. Any input-output relationship has to be so by definition.
Or perhaps you mean that interpolation is when the test input comes from the same distribution as the training input (though this is not technically the meaning of "interpolation"). But this is also quite difficult to pin down.
As long as they continue making headlines they will continue spending. This is just marketing at this point.
Don't hire a straight-A student, unless it's to take exams; or a professor, unless it's to write papers.
-- Nassim Taleb
How interesting that Anthropic and OpenAI are full of professors and straight-A students!It’s unclear to me what point you’re making here - can you elaborate?
If you track the best students futures and look the best workers pasts, there is not nearly as much overlap as society generally believes.
Solving test questions well doesn’t necessarily translate into productive outcomes in the real world.
So education is useless?
No, more like the average can be more important than some outlier's impact, at least for steady progress. So while you do need the Mozart and Einstein, a more productive effort would be to raise the general population's education level.
A lot of people have been surprised by the interest that AI companies have in solving maths problems without obvious applications, as well as in the financially precarious state of these companies. I imagine that they thought that these companies were helmed by mere bean-counters, like at Boeing. But no, I reckon that they're still academics at heart, and so they work on Navier-Stokes and L=BPL, instead of on how to maximise value for their (soon to be) shareholders.
I don't think it will be the case, maths have real utility. I found something interesting at the intersection of combinatorics and information geometry. To be quite frank I don't understand what I'm doing. And yet, when I ask ChatGPT to use the framework we're developing to write an algorithm, it turns out it has quasi-parity with the state of the art. I have to measure absolute perfs to decide which one is better – theirs, not mine. Ok. Time to keep improving on what I have. And this implies dropping the code and going back to the blackboard doing more super abstract math that are way out of my league.
Good luck!
We are at the point where the way in which humans do math and science changes significantly, and I have no good idea at all in what state is it going to settle down. But you are one of (many, I suppose) people exploring the new wilderness, so I wish you best.
If we get AI singularity, then humans will stop doing scientific progress altogether. But if we don't, then it's pretty predictable what's going to happen - things will be much the same as now, except everyone will be using AI for proofs, data analysis, theoretical models and designing experiments, so important discoveries will happen more often. It's also possible that after the AI craze dies down, we'll have enough computational capacity to solve protein folding.
It has utility so some people will pursue it, but it has no immediate business value so I don't believe ai labs will keep spending millions on it.
Unless they decide that trying p!=np is worth any money.
Some math has almost incalculable business value, because math is the biggest driver of game-changer technology.
We'd be nowhere without Laplace and Fourier transforms, Maxwell's equations, elliptic curve cryptography, and many more.
Most math doesn't, but often these techniques are invented first and the applications come later.
And the criticism of the current round of proofs is that while they may be true - likely for some, questionable for others - they're not adding new techniques or insights.
> often these techniques are invented first and the applications come later
There's a great paper from Abraham Flexner on this topic:
https://worrydream.com/refs/Flexner_1939_-_The_Usefulness_of...
It argues exactly that we should be allowed to pursue the seemingly "useless" knowledge.
Previously discussed on HN:
Why do you think this has no business value? It would be absolutely wasteful for OpenAI to not be doing this as part of a post-training RL rollout.
There are architectural advancements yes, but lots of progress from LLMs really come from (1) better pre-training [generally through more cleaned data, and ofc more data], and (2) lots and lots of post-training. It's how we get more and more intelligent models for the same param sizes.
The 'marketing' is just a useful side effect they get from their RL rollouts on maths and LEAN.
There is a risk of this particularly if it's seen as advertising - at some point "ai model solves hard to explain problem" isn't going to be news and that benefit goes.
However, there's some of this that's a proxy - the compute to solve these problems was very low (they claim a few hours of thinking time on a regular subscription). The large cost would have been the training and if training the models to be better at these things makes them smarter for useful tasks that's beneficial. I believe there was work done earlier on around showing that training the models on code made them better at broader reasoning tasks (not just writing the code itself).
Another side is that if one goal is to improve the models themselves, their ability to work on mathsy problems must be high. That has very direct business value, and ideological value depending on what you think the motivations of the people running the companies are.
I think this is an interesting and good theory. They've probably eked out the large majority of the PR benefit at this point, so whether they continue in this vein will tell us a lot about their motivations for this work.
To take this to the next step, what happened after deep mind pretty much solved Go is that they started looking for the next set of things that hadn't been done yet. It does strike me as very likely that this will follow that same path.
I'm told there are another two large tranches of results to be dumped.
Yeah I'll be curious to see if or when they stop seeing the utility in this!
If working on these is causing improvement in general reasoning (beyond math) then I expect they'll keep doing it.
> I wonder if top labs will soon abandon math progress like they did go and chess.
I definitely think that this is marketing, just "with good side effects". My doubt is when they will be able to move to "marketing with better side effects", that is, research with more concrete outcomes (health, materials etc.).
Problem is, that type of research is much harder. Some doubt that progress in such areas will be quick (https://www.noahpinion.blog/p/wheres-the-intelligence-explos...).
Would that be a bad thing?
While top AI labs no longer focus on chess, the community build way better chess engines.
Stockfish is probably stronger, than everything the top labs build.
Wouldn't we expect the same thing for math? That slowly the broader math community would engineer a harness/program... That will surpass the current labs, and be a community ran project
Yeah its certainly true now that Stockfish is much stronger than alphazero, but it's probably also true that had Deepmind spent another few years working on alphazero it would be enormously stronger than either.
In the case of chess this seems fine, there isn't much value to society in creating an AI capable of beating top humans with a 4 pawn handicap rather than a 2 pawn one, but for maths where there are actual applications it is more complicated.
But arguably - it's better for chess community that the best engines are opensource than having superior GoogleChessBot.
They aren't trying to 'solve' chess, go, or mathematical proofs as an end in themselves, but mainly in order to learn more about how to build better systems overall. The goal of AlphaZero was ultimately as a stepping stone towards AGI, and it's the same with LLMs.
The assumption being that 1. these things are all stepping stones, not diversions 2. that ai labs have a singular goal of producing agi
It's not an assumption, many of the people behind these projects explicitly say this what they are doing and why.
Interesting path forwards, and probably partly true, but there are some important distinctions:
Go was a specialized application. All the math results come as a side effect of reading the whole internet, and it will keep reading the whole internet. It will keep practicing thinking questions. Actually, math might be one of the best ways to keep them contemplating and measure their contemplation abilities, so math will always stay in the loop.
Also, math might not be useful just for humanity, but also for AI, so the system might actively benefit from new math results itself. (Not sure if any of the recent proofs qualify, but future work might.)
This misses the raw advantage of a good proof. It makes conceptualization simpler. In some ways math is like a hash list of of theorems. This list makes it simpler to prove other calculations, and will always be useful, to both humans and AI models. I can see two new directions 1 - the creation of specialist theorem models; that can answer questions efficiently about one topic and 2 - we probably need to incentivize and codify ownership of theorems; charging a proportion of the compute saved by using them. Ultimately enabling mathematicians to be paid our true market value!
Oh boy, please not 2. What if this was a thing already and, since neither Newton nor Liebnitz had kids, we all had to pay some investors who bought the rights to calculus every time we took a derivative.
You know patents only last 20 years right?
Considering how essential math and science is for the prosperity of mankind (not even speaking about the cultural value) the question of how to reward people working and contributing in these fields effectively and appropriately is of extreme importance. (And I think the current decline in our societies is to no small degree caused also by our utter failure to address that issue.)
It is also fascinating, because I don't think there is any solution within our existing system, at least not any I know of. Theorem ownership is not a good solution (and neither are patents in general). Probably the most achievable (or rather the least unachievable) solution is a kind of communist utopia, where people can dedicate their time to a pursuit of any endeavor they see fit, as resources for a decent life are abundant and excessive power capture impossible. (The other option, somewhat dystopian, and which would not require humanity to change too much in its current mode of conduct, would be a totalitarian or caste-like capture of society by the scientific community.)
Incidentally, if AI proves as powerful as some expect it to become, it could bring about another solution of that issue by making all human science and mathematics obsolete, pushing its true market value to zero.
(With apologies for rambling.)
Almost every country in the world has a patent system for a reason...and they are pretty essential for big pharma to function. Granted there should be better calibration of duration - particularly in software, but rewarding first movers is wise in many market conditions. My point regarding "math as a hash list of theorems" is that Theorems will never lose their value, regardless of how powerful AI becomes; they represent compressed knowledge, and as such it will be more efficient for a more advanced AI to query known theorems than reconstruct each one from scratch - in that gap lies a marginal compute saving; an api service or tool call that someone might charge for. Even AI can benefit from specialists...
> It makes conceptualization simpler
I wonder if it makes conceptualization simpler for models too, given that they're trained already on human-speak. And I'm also curious as to whether humans currently have an innate advantage into simplifying and contextualizing proofs, or will the machines get good at that as well?
My point is that they will, and that is always going to be useful. Bundles of simplified knowledge (in whatever form) that make research simpler will be useful to share whether humans can understand them or not...
I had the same idea recently. You've solved all the famous conjectures (all formulated by humans because humans found them interesting), what next? I doubt "AI formulated a math conjecture that nobody else cares about and immediately solved it" will produce that much hype. The actually interesting thing is indeed how mathematicians themselves will use these AI models going forward and how that will shape mathematics of the future.
AlphaGo and AlphaZero weren't generalized models. Math capability will presumably keep improving along with the other general capabilities, even if there wasn't a special RL focus for math itself.
Yeah chess is a good example. DeepMind came for publicity with AlphaZero. Arranged a match with Stockfish with rigged rules to make AlphaZero look better than it really was (it was amazing but the match wasn't fair) and then just published some games and went home.
I was bitter about that back in the day as I hoped for more answers, more matches, more "truth" about chess being shown. Soon after that community project Leela Chess Zero was started and not only surpassed original AlphaZero but added few hundred ELO points over it. Then the combination of NN and classical engines happened with NNUE and current Stockfish is again a few hundred ELO points stronger.
Today we pretty much know the truth in chess for all practical purposes. Human analysts/preparation experts focus on finding interesting path and opponent profiling (what is the most unpleasant for the opponent to face). They don't look for truth anymore. The game is doing great, it's more popular than it ever was.
Why don't you mention the second match here, with its adjustments to meet Stockfish's quibbles - and the same result?
https://en.chessbase.com/post/the-full-alphazero-paper-is-pu...
Stockfish and other classical engines were never intended to be run in matches without opening books (of which there were plenty). Development assumed the presence of opening book and authors made 0 effort to make engines play well in openings because of it. This is also the reason classical Stockfish was a very small binary. A little effort to make it even by including even a very small opening book (like 50MB or something that would result in still smaller binary than NN engine with its net) would make it much more interesting.
The result was that Stockfish lost many games by walking into known bad lines and lost way more games than it otherwise would.
You aren't correct.
> We also played a match that started from the set of opening positions used in the 2016 TCEC world championship, along with a series of additional matches against the most recent development version of Stockfish, and a variant of Stockfish that uses a strong opening book. In all matches, AlphaZero won.
https://deepmind.google/blog/alphazero-shedding-new-light-on...
Ok I remember it vaguely but the match that got publicity and the one published results were derived from was 1000 games match from starting position. Deepmind claimed AlphaZero also won from TCEC positions and vs Stockfish with good opening book but at least back in the day I don't think I could find those games being published or specifics about books/positions they have used. Can you?
I am not claiming AlphaZero wasn't stronger. It wasn't as strong as the PR piece suggested though and we have never seen the games being published. In chess this is extraordinary because basically all games in chess are publicly available - both human and computer games. Claiming "we have created a strong engine that has beaten Stockfish with opening book" while not showing those games (or details about opening book used) is akin to "we solved this math conjecture" without showing any kind of proof or argument.
Publishing a few 1000 of games costs nothing. Tens/hundreds of thousands of games are published every day.
I disagree. Firstly, people in AI likely care about math on a personal level. Secondly math is useful. Playing go or chess is basically a party trick. Being useful gives it staying power.
But, I do think you are right that there will be some level of moving on. The spotlight is currently on maths and that won't last. It will move to some other area where there is more impact to be had. So while they might shift gears and put less focus on math, it will always be there as part of the portfolio.
I think there’s a venue where they start focusing on introducing hypotheses where the model currently can’t solve it, or maybe this is already happening?
Being able to present useful novel ideas would likely generate a lot of press, for a while. I don’t know how this would look since I’m useless at math, but Im sure there are plenty of unknown problems with massive implications, that once formulated can be solved.
This argument implicitly makes a few assumptions which will probably not hold in the very near future.
One is that AI will continue hallucinating in a manner that is not easy to verify, second is that AI will not be enhanced to produced more simplified amd robust outputs, and third that a human will be required to do that. What humans in the loop are doing now is verify the process, propose shortcuts and add legitimacy, through the verification process, if that ends up being succesful its highly likely a lot less mathematicians will be required in the future.
The conclusion that this is not productive focuses on the mathematicians, but it is very productive in terms of hundreds of proofs being produced that had previously consumed uncountable hours of the brightest minds. Unless it ends up being the greatest hallucination ever ofcourse
> assumptions which will probably not hold in the very near future [...] One is that AI will continue hallucinating in a manner that is not easy to verify
Hold up, that's an even bigger assumption in the opposite direction, and I don't see anything to support it.
At least in terms LLMs getting all the "AI" hype these days, there is no structural/mathematical reason to believe they won't continue to have the same problem they've always had of generating plausible text over rational text, and I don't think anybody even has a clear idea how it could eventually be accomplished.
I've seen "then the magic singularity occurs and somehow it solves the problem for itself", but I would classify that more as mysticism than engineering.
Hallucinations are no longer much of a practical problem in software engineering.
Two years ago, hallucinating that the code worked or that a task was accomplished was a common occurrence.
We have seen that now agent swarms across thousands of agents can coordinate to achieve a result.
Clearly hallucinations are no longer the problem they once were, since now we can get working results for long horizon tasks that require massive compute.
Consequently it would seem unwise to assume that current limitations will remain as they are and prevent LLMs from coming up with solutions that they can explain to humans.
It’s still a common occurrence.
It happens in more subtle ways, but it still happens often enough for me to notice. For example I have had hallucinated checksums show up in lock files as recently as yesterday using a SOTA model.
This is not surprising, since the whole basis of LLM training is to produce output that humans will accept _as a proxy for actual training goals_. In a sense, the training process of an LLM “wants” to produce output that is statistically plausible much more than it “wants” to produce correct output. It’s always going to be a struggle to drive that system towards other goals (and we see this bourne out in practice by the amount of effort that is required to be spent on RL).
I think there will be some threshold of correctness (something like 99.999% of the time) that if the model surpasses it, I can stop needing to check it, but I think we’re still at 99% or something which sounds good, but when you are producing a ton of output you hit that 1% frequently.
> Consequently it would seem unwise to assume that current limitations will remain as they are and prevent LLMs from coming up with solutions that they can explain to humans.
I 100% agree with this. In fact explaining things to humans is something LLMs are particularly well suited for.
The confidence with which you, anonymous user, keep commenting that "hallucination is not much of a practical problem in software engineering anymore" based solely on your own anecdotal evidence is really remarkable, in not a good way.
You're free to substantiate your comment by telling us about your apparently different experience.
I find hallucinations in my (mostly perfect) AI output every single day. If you're not finding them, you're just not looking hard enough. It's not surprising when everyone is screaming about how they don't read code these days.
This is just a fact. I'm sorry if it messes with your narrative.
Thanks for the link. What kind of hallucination are you seeing, and does it affect the end result?
I'm not them, but I see quite a few hallucinations as well. Some examples:
I work on a device that gets firmware updates over USB-DFU. If a DFU fails, its bootloader restarts, re-enumerates USB, and waits for a new DFU to begin. AI decided there's a risk of bricking the device if an update fails. There's no brick risk, a retry fixes it.
The same device uses CAN bus, with the common bosch_mcan peripheral. Claude decided that calling can_mcan_stop followed by can_mcan_start somehow left the transmit buffers intact, and tried to implement a fix which manually cleared the buffers. They're cleared automatically by the hardware when can_mcan_stop is called, and that's documented in the comments of can_mcan_stop.
Both cases could have resulted in unnecessary changes getting deployed which wouldn't have fixed the actual issues. Since it's embedded that could mean significant delays to getting fixes to customers.
The sum total of all human observations is still not proof of the lack of hallucinations as a problem (even if their observations were perfect, which they aren't considering the volume produced vs reviewed carefully). That's why you can use a counter example only to disprove and not prove anything.
And yeah I get hallucinations all the time still. Maybe it's because I'm working on harder/more niche problems (like a compiler with an unusual type system), but it happens quite a lot. I don't record all of them.
Although the most common one you can find is them misattributing the source of changes from themselves and also other agents (Fable, Opus 5.5, deepseek, whatever). They'll say "your changes" or "you changed" or "your ruling." I didn't decide anything and it's in their own chat log, and yet...
"Plausible" text was preferred over rational text when we trained LLMs using RLHF. It's rapidly shifting the other way now with RLVR, which enforces correctness by default.
> One is that AI will continue hallucinating in a manner that is not easy to verify
It is an old saw at this point, but what an LLM does still cannot be divided into hallucination and non-hallucination. This is literally an anthropomorphism trap.
Layers and layers of application-specific verification can reduce the risks inherent to LLMs, to a really remarkable degree, but nothing about what these tools are suggests that this problem will go away; it will just bubble up again somewhere else.
And why not?
For all that I saw over the last few hundred hours with AI on software engineering, hallucinations are no longer a problem at all.
Not once have I seen a task fail due to what would have been a "hallucination". If they still occur, they can apparently be detected and corrected automatically, or are subtle enough to escape notice with presumably no significant impact on the results.
Why would this not also be the case for mathematics?
I think OP is saying that hallucination or not is just semantics. There is nothing qualitatively different about hallucinated vs non-hallucinated output.
That's true in the same sense as "There is nothing qualitatively different about erroneous vs non-erroneous output" for a dog vs. cat image classifier.
I think that’s a bad example, as classifiers tend to output a floating point number and you use some threshold/activation function to collapse the classifier into a particular state. In that sense there is nothing qualitatively different when the classifier outputs 0.85 vs 0.87.
LLMs also output a distribution over tokens at the output. I don't see the point you're making.
To be fair, I guess the line is blurry between what could be labelled a regular mistake compared to a hallucination.
"Test suite passed" when it actually errored? Obvious hallucination, unless it ran a command that returned the wrong error code.
But is running a malformed command that does not achieve the expected effect itself a hallucination?
If it makes a false claim, then it's an error. If it says the test was passed or a class was implemented but it was not, then it makes a false factual statement.
I'd say a hallucination (very misleading word) or confabulation or "making shit up" happens when an LLM uses factual / evidential language purely based on local statistical expectations of the text, instead of it drawing from actual evidence in its context pointing to it.
This is murkier in the case of general knowledge questions, like when and where was some famous person born. It may then be a spectrum from fully making something up based on how the name sounds, all the way to confidently retrieving it from its weights correctly. In between, we can get hallucinations. But newer models are taught to use Web Search when unsure, and it works pretty well, though not perfectly. I don't see any fundamental limit here. It's just not perfect. Trying to solve "the hallucination problem" is basically like saying "our dog vs. cat classifier is pretty good already with its 99% accuracy, now all we need to do is the tiny little task of eliminating the 1% error, and we will be golden". Like, no shit, there is some error yes. People are working to reduce it. It will never be absolutely 100%. It's not an insight to say we should remove hallucinations.
Intuitively, something about that version bothers me... I think it's because the choice of a clearer "erroneous" has dropped the fundamental framing problem from "hallucination": The false implication that "true" (anthropomorphic) sight/thought usually happens.
To fold that back in (and ham-it-up a bit) how about:
> There is nothing qualitatively different when Pet Classifier <ironic-quote>maliciously misreports</ironic-quote> your cat as a dog, compared to when it works <ironic-quote>honestly</ironic-quote>.
These models are outputting language that we can decode as factual claims and we can check those, and we can say it made a mistake / error or that it answered correctly. This doesn't require any squishy assertions to how it "feels" while doing it or anything like that. It's an externally observable thing.
I agree that hallucination isn't some kind of "different" operation than "normal". It's not like when a train derails and you can point to it. It just operates as normal and sometimes that yields correct factual outputs, sometimes not. You don't have to metaphysically ascribe any kind of intent to it.
I'm not sure how people conceptualize these things who weren't doing classical machine learning before all this. To me, "hallucination" is shorthand, and we know it's not like humans on drugs or something. It was used in the literature also for any kind of generative imputing of missing information from a learned prior. For example in image inpainting a GAN "hallucinates" the missing part of the image. This terminology was already used in the 2010s and probably earlier. Or in image colorization of grayscale photos, the model "hallucinates" the color information.
Then the word escaped into the mainstream and people have weird connotations about it.
> To me, "hallucination" is shorthand [...] Then the word escaped into the mainstream
One of my bugbears is when people abuse the term "Ponzi Scheme" to refer to literally anything the think is unsustainable. (As opposed to something that, at a minimum, requires someone telling factual-lies about assets.) Kind of like if folks started calling every kind of software error a "Buffer Overflow."