Sharing AI progress in mathematics
openai.com1190 points by OfficialTurkey a day ago
1190 points by OfficialTurkey a day ago
https://github.com/openai/math
https://github.com/openai/math/tree/main/preprints
I was a graph theory junkie long ago and even moved to Budapest for awhile to study among the greats. While I was there I started working on Barnette's Conjecture which came to occupy my thoughts over the next 24 years of my life, on and off as I worked in many different fields. Last summer I even thought for a few days that I had actually solved it. But it's supposedly proven here - problem 180. I don't know what to think exactly. I spent thousands of hours on that problem. I really enjoyed it. Hearing that it is solved somehow makes me sad in a far-off way, like hearing an ex-girlfriend died suddenly in a car crash. I don't know, there's probably a lot of people feeling odd emotions tonight. There's no Lean proof for this one so I'm digesting the paper. On the surface it looks like an approach I considered 24 years ago and abandoned. I revisited the problem this summer, along with my partial solutions, when the previous round of stunning proofs came out. Several hours of work with Fable simply convinced me it wasn't yet solvable and reinforced how hard of a problem it was. > Several hours of work with Fable simply convinced me it wasn't yet solvable and reinforced how hard of a problem it was. This is the part that gives me the strangest feeling about it all, because you're not the only one with this experience. I've experienced this too on different problems, as have many researchers across many fields. I disagree with the Fields Medalists on the majority of their complaints. AI math is happening and there's no going back. However, on one point I increasingly agree: virtually none of this stuff is possible with technology any normal citizen has access to. I have no problem with AI models making revolutionary advances in math or science. Where I start to have a problem is when the AI models making these advances are tightly withheld, proprietary, and seemingly never released with these capabilities intact. This has been the case for all of 2026 so far. I suspect that this is in fact the source of much of the angst. None of this progress is reproducible outside of one or two teams inside OpenAI and Anthropic. It's becoming an incredible concentration of power that I don't know that we've ever quite seen before. Right now, it feels harmless because it's being used for wonky math problems that aren't (yet) practical for anything. But great power never stays harmless. History has taught us that countless times, in countless different forms. > I suspect that this is in fact the source of much of the angst. Why do you "suspect" this as if it's some hidden motivation when the very first paragraph of the advisory group's statement (linked from the OpenAI post) says: > At present, some frontier AI labs are testing advanced mathematical problems on proprietary models that remain inaccessible to the broader scientific community. Our recommendations are formulated with this practical context in mind. However, ideally, they would not do so. We want to state clearly from the start: we do not endorse this practice, and we ask them to stop testing advanced mathematical problems on proprietary models. Tao and others in that group have been strongly and publicly pro AI from the start. They are not advocating "going back". They're objecting to the strip mining of open problems using proprietary technology. OpenAI: At long last, we have created the Open Problem Strip Miner from classic Terence Tao tweet “Don't Create The Open Problem Strip Miner”. I don't think the strip mining metaphor is appropriate. Mining is a zero-sum game; if I mine something, nobody else can go and mine the same resources I did. Mathematical problems don't go away when AI finds a Lean proof. They create new opportunities for humans to study the solutions, learn new techniques from them, identify promising directions for future research, discover alternative/more beautiful proofs, and write expositions for other humans. Strip mining is very apt if you view the economics of the present system as "effort -> recognition -> career advancement". Even in strip mining, the resources that had been buried are now available for use in the broader economy. What's no longer available is the living that was to be had digging them out. The problem isn't effort, though. All of the things I mentioned constitute effort and could be rewarded. The job economy was created by mathematicians incentivizing the proof of difficult theorems above all else and valuing all other work at approximately zero as far as career advancement was concerned. Now they're pulling a 180 and claiming that math was never really about proving theorems, but that's contradicted by their revealed preferences. The strip-mining problem only exists if they continue with the status quo ante. Mathematicians aren't homogeneous. There are mathematicians valuing pedagogy, collaboration, bridge-building, theory building, along with those that chase the 'difficult theorems', to name a few, and there are lots of flavors within each class, with lots of blending and blurring. You infer that mathematicians prefer the status quo simply because it is the status quo -- with a little thought, you'll recognize that this is a fairly silly notion. There are myriad circumstances where the values of most practitioners differ from the status quo, which is nevertheless well-entrenched. This can arise from inertia, or from outside forces, such as broader cultural milieu, integration with larger institutions, or contending with economic realities. If you think that these do not and haven't historically played a role in determining the job economy and that math is a pure field where mathematicians could comfortably shape it according solely to their own ideals then you are naive And, in addition, many mathematicians are graduate students or postdocs hoping to line up a permanent job soon. For example, if you look at Terry Tao's blog, he has a tremendous amount of first-class expository writing. So, too (to some extent) do junior mathematicians -- but, unfortunately, this tends to not be highly valued by the job market. Grad students and postdocs have learned that to succeed they need to play by the existing rules of the game. Well, the board has just been yanked from underneath them. People like me can afford the sort of idealism and soul-searching that the parent comment describes, but junior mathematicians face a very unenviable set of circumstances. 1) I never said the problem was effort; I was trying to explain the strip mining analogy, and it's one of the two anchors that make the analogy work. 2) Mathematicians didn't create this economy; it was foisted upon them by the same managerial mentality that brought us "publish or perish" and "the monthly sales quota". 3) I can't tell if you honestly don't get why the strip-mining analogy resonates, or...? Here's another analogy: if we suddenly discovered personal teleportation, and marathon runners were complaining that it was ruining the sport, would you say "they're pulling a 180 and claiming that marathon running was never really about getting to a point 26 miles away as fast as possible, but that's contradicted by their revealed preferences"? The strip mining analogy is better though, because it captures the sense of irreversible goal-loss when a problem goes from being "unsolved" to "solved". As far as I understand, even with "publish or perish", peer reviewers decide what counts as an important enough paper to be published in a prestigous journal, and committees of peers decide whether or not, say, an expository article on arXiv or a textbook counts toward hiring or tenure. Again, as far as I understand, those things have largely not been rewarded in the past. I like your marathon example, but maybe not for the reasons you intended. The community of marathoners decides the rules of a marathon. You don't need a hypothetical teleporter; you're already not allowed to use a bicycle, performance-enhancing drugs, or shoes that don't fit the specifications. The rules are updated to adapt to changing technology. Yes, I'm arguing that the strip-mining analogy doesn't make sense because mathematics is in the same situation. There's nothing stopping peer reviewers and hiring/tenure committees from changing the rules about which kinds of effort confer recognition and career advancement. Strip mining is an extraordinarily appropriate metaphor. Imagine a mine has an unknown number of rare materials. And you know the general location of a few of the most valuable spots. But you don't know what may be valuable right next to it. If the pieces that we know are valuable are suddenly gone, the incentive to mine that particular area drops considerably, dropping the chance to discover potentially brand new materials that would have been found the normal way. > the incentive to mine that particular area drops considerably, dropping the chance to discover potentially brand new materials that would have been found the normal way. FWIW, I think the metaphor breaks down with this framing. This isn't really a problem associated with strip mining, what's left behind is generally low or negative value (toxic). I'd suggest a different metaphor, from Wikipedia: > This process involves the removal of all ground vegetation in the area, which is a detriment to the environment.[19] Topsoil may be placed over the tailing along with planting trees and other vegetation. Another reclamation method involves filling in the hole with water to create an artificial lake. Large tailing piles left behind may contain heavy metals which can leach out acids such as lead and copper and enter into water systems. This feels very similar to the issues with algorithmic problem "mining". It has the potential to destroy the human ecosystems surrounding these problems, leaving barren wasteland behind where nothing can grow or flourish. That's an empirical claim. I could equally well say that doing an automated search of the problem space and having a database of results and open problems will identify vastly more interesting and valuable areas. Again, the idea that math is some exhaustible material is a metaphor, not an established fact. I'm willing to change my view as new evidence comes in, but I think we're going to have to wait and see what the landscape looks like in a few years. > Tao and others in that group have been strongly and publicly pro AI from the start Unfortunately being "pro AI" means relinquishing any control over what the AI, or more importantly the company running it, might be doing. How is this different from literally any other part of the economy? We've relinquished control over just about everything we use or consume. We can't compete with larger enterprises for production of food, clothing, machinery, medicine, energy, services. Mathematics is just the latest thing to be industrialized. What keeps large companies under control is competition with other large companies. This competition causes the surplus value they produce to flow to consumers, not be hoarded via monopoly prices. Do we see strong moats that are going to cause monopoly in AI? I don't see it, and in particular I don't see it persisting if it exists transiently. You're right, and that's a bad thing. AI is nothing fundamentally new, but its extremity is making many people aware of the truth that's been there all along. There's no contradiction in that. > Do we see strong moats that are going to cause monopoly in AI? Ownership of the capital assets used to train and inference new models. Yes, we may end up with more than one firm. But as we see with big tech today, a small number of fantastically wealthy firms in "competition" does not an open market make. Is it a bad thing? We live in a society. We depend on the work of other people. We are not autonomous. Sure, we can try to be self-sufficient, and that would lead to a subsistence lifestyle much degraded compared to what we experience. Somehow you have to argue either that society itself is bad, or that math is somehow different from all these other human activities. I think the obvious fact that people prefer to live in places with large commercial organizations shows they don't really care about that, at least to the point of foregoing the benefits these organizations bring. No, it doesn’t. You can be in favor of something and opposed to a particular way of handling or implementing the thing. And the issue here isn’t what it’s being used for but who is able to use it. "AI" is largely a marketing term for a particular type of computer program that uses a statistical language model. Computers and computer programs are tools. Humans always remain sovereign over their tools. And incentives are sovereign over the humans. The humans leading the AI labs have every incentive in the world to move quickly without any restraint. The second law of thermodynamics always wins. Eventually. In the meantime, here in the human socioeconomic sphere, you might be dealing primarily with the Second Rule of Fight Club and Operation Mayhem. I am not sure how convinced I am by that argument. A gun is also a particular kind of tool, and it makes the person at the handle end sovereign, and the person at the pointy-shooty end subjugated. Regarding the advisory group, OpenAI claims to “have drawn on their advice”, which would include not dumping a bunch of AI slop, with the footnote that if they do do that, at least fund the process of digesting it. At the same time, there's a new note at the bottom of agmai.org stating how they've been in contact with OpenAI about this particular release, and they say that “we consider these discussions constructive, it is ultimately up to the mathematical community to assess the extent to which our recommendations were followed successfully”. So, what's going on there; is this British English for “they didn't follow anything at all”? Because from my perspective, it looks like they doubled down on the Navier–Stokes approach of trying to maximize PR gain while being as lazy as possible about actually contributing anything back to science, releasing only slop that may or may not be correct and may or may not be straight up plagiarism, as has been the case earlier. If I were on the AGMAI board, I'd feel terribly exploited when reading that press release, yet their response is modest. Hairer, if you're reading this: is there any indication whatsoever that AGMAI was anything but a cheap way for OpenAI to science-wash their press release? > is this British English for “they didn't follow anything at all”? Yes, but the subtext is even stronger. > AI slop, Now I know there are issues with the field and how just answering these questions may cause broader problems, but I feel like the posted results is far from slop. We can't just call any output slop, or it loses all meaning. If it was slop, it'd not be causing the issues the group are concerned about - they're not saying "the problem is we're getting loads of incorrect proofs thrown about that are nonsense". When you blanket a set of things with a pejorative, and it turns out that some of the members of that set are demonstrably and definitively NOT covered by that pejorative, and that all the pejorative means at bottom is "I don't like", all you've accomplished in the long run is to call into question any future legitimate use of that pejorative. It is tempting, especially when heated, to stretch an invective, but it will ironically only lead to the death of its utility over time. So the fact that the Library of Babel (i.e. all possible books) contains occasional gems means that you can't object to using it on principle? That would seem to follow from your logic. What about a filtered set "all well formed books"? Or "all well formed books that are plausible enough that they could convince a reasonable person, regardless of their accuracy"? It's generally taken that a cup of sewage in a barrel of wine makes a barrel of sewage. Surely a reasonable person could claim that a barrel of sewage was still sewage, even if it contained several cups of wine? I personally just find it hilarious how the complaints and excuses against AI have slowly marched and changed from 2023 to now. I've read some of the results papers (the Einstein condensate one and the pi exponential one). I'm not an expert but it definitely wasn't AI slop. The introduction sections were particularly well framed and informative. Also you can see in the papers where an idea is introduced but in the bibliography you can see where the foundational idea comes from. So the narratives are not unmotivated as some claim (proof without intuition claims). In the context of maths papers, the term has come to refer to papers having the shortcomings that are, for whatever reason, typical of LLM out, including things like using non-standard terminology all over the place, emphasizing easy steps while leaping over harder ones, having bizarre organisation, and, importantly, failing to properly cover existing work and as a result being hard to tell from plagiarism. The degree to which these issues feature will differ, but it is generally the case that converting the output to proper research requires significant effort, hence the AGMAI recommendations being what they are, and not performing that effort tends to come off as laziness or incompetence, so I can see how slop has become the popular term. > We can't just call any output slop, or it loses all meaning. The term “ai slop” is not supposed to discriminate good ai output from bad, the entire purpose of the phrase is a blanket term that delegitimizes all ai output. That is not how it is generally being used. That is exactly how I see it generally being used. Why else would people be dismissing work as AI slop without even reading it, discovering what it says, or even looking into how and to what extent AI was used in a project? Saying things like "if you didn't write it I won't read it" at the first whiff of an AI smell is absolutely said to delegitimize all ai output. Or in this specific case, why would someone call these proofs (no one is saying they are wrong) AI slop if not to delegitimize all AI output? This. There's a large subset of people who, seemingly consciously, try to delegitimize anything related to AI by calling it* "slop". * even pretty amazing advances like this one
jboggan - 18 hours ago
nilkn - 15 hours ago
omnicognate - 14 hours ago
alberto-m - 11 hours ago
fasterik - 6 hours ago
MarkusQ - 6 hours ago
fasterik - 5 hours ago
vector_spaces - 4 hours ago
impendia - 24 minutes ago
MarkusQ - 5 hours ago
fasterik - 5 hours ago
Catloafdev - 5 hours ago
AnIrishDuck - 2 hours ago
fasterik - 4 hours ago
pjc50 - 10 hours ago
pfdietz - 7 hours ago
achierius - 4 hours ago
pfdietz - 34 minutes ago
thfuran - 9 hours ago
curt15 - 8 hours ago
throw-the-towel - 8 hours ago
morpheos137 - 8 hours ago
59percentmore - 7 hours ago
bebimbop - 6 hours ago
pred_ - 11 hours ago
phillc73 - 10 hours ago
IanCal - 9 hours ago
cjcole - 6 hours ago
MarkusQ - 5 hours ago
hardbass - 3 hours ago
magicloop - an hour ago
pred_ - 4 hours ago
ModernMech - 9 hours ago
thfuran - 8 hours ago
ModernMech - 8 hours ago
anthonyrstevens - 6 hours ago