Caltech Mathathon – first hackathon ever devoted to research level mathematics
mathathonchallenge.com221 points by astroanax 16 hours ago
221 points by astroanax 16 hours ago
Hey guys, I'm one of the organizers. AMA.
- We are a team of undergrads at Caltech. We don't represent Caltech, any Caltech departments, or any of our sponsors.
- We don't receive monetary compensation. All the funding raised goes toward paying our judges and participants.
- Our goal is to promote responsible AI use. You can read more about our commitments here: https://mathathonchallenge.com/faq.html
Why not specify 2-3 problems? Won't you have people show up having already spent a bunch of time on their self chosen problem? Which sort of defeats the point of seeing what you can do in a short period of time?
Solving a problem is the easy part with AI. Selecting a problem worth solving is half of the challenge.
We allow prior work as long as it's labelled. We record chat logs, so it's easy to verify what's prior work. When we evaluate the significance of a result, we focus on the part produced at the event.
it seems like a person could cheat by finding an open problem that's AI-solvable ahead of time (by trying many problems), and then pretend to do it for the first time in the competition. You can't verify past chat logs so make sure they didn't already realize it would work.
just something to think about
We are giving each team $20k AI credits. I hope they will use it to solve problems beyond the reach of regular AI plans.
Is this hackathon only for those with formal math backgrounds?
I've seen a few instances of AI assisted advances math and cs this year that were _not_ published by authors with formal backgrounds in those fields (or even institutional affiliation). Which makes me wonder if they would have a place at the event.
Yes. To us, solving a problem with AI doesn't matter as much as selecting the right problem and understanding the proof. A participant must have a good mathematical intuition.
Hey! Any indication on what area of mathematics theses questions are from?
You pick your own problem! You can even formulate your own conjecture and then prove it. Picking an impactful problem is part of our judging criteria.
That sound cool!, do you have hints on the cash prizes the website said something like 2 M ???
Serious question: Are you sure solving abstract mathematical problems with AI is responsible use? You will likely put mathematicians out of jobs, and I doubt solving the Collatz conjecture is urgent or will save lives. It also robs a future Fields medalist of the pride of doing all by themselves.
All things AI seems to assume that more and faster is better, but there is no justification of that assumption. As a biological counterexample, a tree grown quickly will likely not be as healthy or strong as one grown slowly.
I dunno, I like math and I still think automatically solving problems is worth doing. If problem-solving is a hobby then it will still be a hobby after all this. If it's about doing things for humanity--which it often is since many thousands of people are paid to do it--then doing it more efficiently benefits humanity. If the value on the other hand comes from people being extremely good at math, rather than from them solving novel open problems, then we should pay them to do that, which is independent of whether the open problems are solved or not. There's just no version of this where "not solving the problems" is morally justifiable.
Now, I think it's the case that professional mathematics spends way too much money on open problems and way less than it should on pedagogy, exposition, mastery, etc. But that has always been a problem, even decades ago (I've been complaining about it my whole life). AI just finally puts pressure on the world to do something about it. I find it relieving, honestly. And I'm an AI skeptic in many other ways; it's not an AI-maximalism thing. I genuinely think the state of the field of mathematics has been something of a disaster for a long time (thanks, largely, due to the academic incentive structure which heavily favors novel results, no matter how esoteric).
FWIW even Astra hasn't been able to solve the problems I care about, which are less about proving theorems and more about understanding the right way to think about already existing stories (and thus permitting extensions to new contexts). However it's been a more than a capable interlocutor to test my ideas with and see if they actually have any content. It's also great for parsing possible mistakes in long technical arguments that at least my brain isn't wired to verify completely satisfactorily. Personally, I think it's good to know what is made trivial (meaning depending only on token expenditure) vs what remains a real hard kernel.
> you will likely put mathematicians out of jobs.
why is that a concern in this context? would you have asked the same about steam engines and horses?
this is a really cool concept, organized very well. and that is very commendable.
steam engines solved a problem people had
If mathematicians aren't solving problems people are having (which your comment seems to imply), then putting them out of their job with AI is not a bad thing. Of course mathematicians are solving problems, just in a very different way than other professions.
Don't respond to a strawman argument with another strawman. The post you are responding to ignored the reasons given in the second paragraph. They're just trying to score points by preaching to the choir, not engage with the concern.
what strawman?
do you mean,
> All things AI seems to assume that more and faster is better, but there is no justification of that assumption.
is good argument?
of course faster discovery without human in the loop is better. is that not what humans have been optimizing for the past few thousand years ? faster mobility, faster communication, faster medical recovery etc. everything modern civilization has to offer is because of a rush to get better and faster. for example, discovering penicillin 2 years early would've saved ~15 million people more.
why is that not worthy enough to pursue?
I don't think AI will solve more math problems in a world with Mathathon than the counterfactual by EOY. People will use AI in math anyways. What matters is: can we encourage them to do so transparently and with full understanding of their results? Can we change the incentives in academia to reward problem selection and verification over proof generation?
It’s a noble goal to change the incentives, but how will you prevent the headlines from this event being “students prove Collatz conjecture with Claude” and instead be “students give great explanation of Collatz conjecture proof”?
You're right that we can't. We'll be responsible in our press releases and award prizes based on explanation, but we don't control the headlines. However, this is already an improvement over the current state, where results are announced by headlines alone.
It's more responsible than the use of electricity for a messaging board for people to argue minutiae.
I get your point and agree to some extent, but you can't understand the proof without significant background in Maths so it will just allow mathematicians to solve issues faster than not have the opportunity at all.
Why do people need to understand proofs? If Amazon improves package routing with new advances in graph theory, my cat doesn't need to understand it to benefit from better shipments of cat food.
Similarly, humans don't need to be involved in scientific advances to benefit. We just need an aligned AI to take over the scientific thought for us. AI is already better than all but the top tier of humans at doing mathematics, it's writing most of the posts on the front page of this website, and it's doing the bulk of programming at many startups.
We can't put this genie back in the bottle.
Well, for one, most math proofs don't have any practical applications, so a proof that no one reads is basically a digital paperweight. You might as well suggest AI write novels for other AI to read.
The hope is that some of them end up being useful; otherwise, nobody would be funding math departments. Mathematics typically anticipates and enables new physics and chemistry.
If people are just doing math to kill time, I don't get why anyone would bother with AI. Do people really enjoy picking through a million lines of generated Lean code, if it's not for any practical use?
If you're interested in the topic enough to comment on it, you'll probably find it worthwhile reading a mathematician's perspective. Here's the prolific Terry Tao: https://mathstodon.xyz/@tao/117219548485446992
They don't actually say anything about why anyone should fund this, though. I don't get why a society should worry about progress in mathematics if there's no practical benefit expected.
Maybe there's two kinds of math that we need? Useful math and navel gazing, and we can hand the first to the machines, and let hobbyists do the second in their free to entertain themselves?
It's hard to know what math is 'useful' a posteriori. That's always been the argument for supporting basic research. This is not why I am a mathematician however. I think there's intrinsic value into understanding something of depth and meaning, but the societal setup we have now that mostly agrees this is valuable is probably a very contingent phenomenon that is unlikely to last much longer.
Hi check zeta.pukapasoft.xyz
am I eligible?
recent caltech grad here! and know some of the organizers well
caltech's cs department is very, very weak, and has struggled to recruit top people in the last few years, and the most recent AI faculty hires have had issues. this is very slowly changing but a lot of the motivation for htis was to create a way for students to get ml "recognition" and learn about ai since it cant be done through the school right now. really glad to see hn picked this up!
We aren't affliated with the CS department. We're a student-led initiative. Our goal is to promote responsible AI use in math.
Won't deny that this is an interesting idea, but I feel like waiting on the output of an LLM for 40 hours feels like it is completely antithetical to what makes classic Hackathons appealing / educative.
More generally, I don't think the shape of a hackathon (intensely working for a short timespan) maps at all onto the way LLM Math progress has seemingly been made so far; AFAIK it mostly involves picking out something for the Model, then having it run for a week with sporadic correction / encouragement.
1. IMO the hard part isn't prompting. It's selecting the problem and understanding the solution. It'd be especially exciting if a participant formulates their own conjecture, proves it with AI, then generalizes it to a new theory.
2. We have talked to mathematicians and frontier lab employees. We think 40 hours is enough to produce interesting results.
But it's not "waiting on the output of an LLM for 40 hours" any more than a regular hackathon is "waiting for my damn teammates to finish their part for 40 hours". From my experience using agentic coding for hackathons, the best teams are those that coordinate with the AI agents in relatively quick cadence, generally giving it small tasks and steering it often. Teams may want to run some long-running sessions too, especially closer to the deadline, but even then, they'd probably want to run and follow several sessions in parallel, and continuously inspect their work so that they have reasonable confidence that their main efforts will wrap up before the deadline. There is an art to it.
Have you done any math hacking with sol/astra or fable? It’s more fun than using them for coding. The models are great at the monotony, like constructing a Gröbner-basis, etc. But they’re all still absolutely awful at coming up with new ideas, new proof methods, or new constructive forms. So you spend all your time on coming up with novel hypotheses yourself and handing off the rote work to an agent.
It’s also quite fun to get instant results by finding isomorphisms into unfamiliar areas of mathematics that previously would’ve required some networking in order to build a collaborative relationship.
"A mathematician is a person who can find analogies between theorems; a better mathematician is one who can see analogies between proofs and the best mathematician can notice analogies between theories. One can imagine that the ultimate mathematician is one who can see analogies between analogies."
I wonder how models perform on finding analogies between analogies
Well, I suppose that explains monads. It’s one thing to see an analogy and quite another to make it the basis of an API.
> I wonder how models perform on finding analogies between analogies
Load-bearingly verbose, in my experience.
If the goal is to accomplish something then why limit yourself with available tools?
I’m not a full on AI optimist but it is absolutely the most powerful tool in a host of applications. From a Hackathon perspective, obviously in the 90s it was much more unorganized, but the same ethos existed. Use all available tools to accomplish the goal/task, it’s where a lot of incredible learning came out of. The same will hopefully happen in scenarios like this one
It seems like making progress on math is letting the AI run fully autonomously for a few days, occasionally asking it to keep going.
I'm not sure people need to organize a mathathon to wait for a computer to give a printout. They mainly need tokens.
That's how several major AI advancements have happened. I have seen no evidence that that is the fastest way to make progress right now. I expect that, much like chess engines, it will not take too long before AI is significantly better than AI + human. But right now, my bet is that we are still safely within the window where an AI + human mathematician team is still better than AI alone (at least for the case where the human has learned how to work effectively with the partner....something that this event could possible be good for teaching).
I suspect that the best progress will be made by a team that purely spends their time taking a list of open problems and promoting "solve <problem>", without actually trying to understand anything. Just keep as many problems in flight as you can across as many sessions as you can.
You can probably ask the AI to come up with a list of problems itself, and rank them by the likelihood of progress.
Then 20 years go by and you wake up one day with questions that you cannot get out of your mind: why did I start prompting the LLM for? Why did I need these random proofs for? What do i do with my repo with 2billion lines of Lean?
Are they actually autonomous? I’d say subject knowledge at the prompt stage plays a large part towards getting proper results
When Claude made progress on the Riemann conjecture, here are the kind of prompts used:
> Jarred's input was mostly limited to sending Claude messages of encouragement (mostly variants of “keep going” or “believe in yourself”).2 This seems to have helped Claude overcome some initial skepticism that it could make meaningful progress.
And left it for a long time. Jarred isn't a mathematician, he's the maintainer of a janky JavaScript environment.
Here's the transcript: https://www-cdn.anthropic.com/8a0d1add3c637b858a9a181e98c40e...
Prompting for some of the results was almost the "Computer, do a breakthrough. Make no mistakes." meme. Just someone telling the model to keep trying a couple of times.
Unfortunately we don't actually know what kind of prompting was done for the more prominent results.
It’s definitely not how I work, I’d need to read the model responses and set a direction for the model to go
Same here, it's one thing if I'm just screwing around, but if I'm trying to do anything serious, I need to at least have a handle on what it's doing, and, when thinking traces are available, keeping track of any logical errors in the model's reasoning.
you know the thinking traces are redacted and summarized using another model, right? The actual thinking trace looks something like:
7♣-removal-IS-the-prerequisite-for-10♠/9♥!!)-⟹-OVERLAP-(ii)+(iv):-{6♠ J♦ 9♥ 2♣}-=-FOUR--—-UNLESS-7♣'s-seat-8♥-...-and-2♣-drains-only-at-crack-:-⟹-2♣-celled-+-9♥-celled-simultaneously-UNAVOIDABLE-in-t8-dig--—-BREAK:-9♥-drains-to-10♠-THE-MOMENT-10♠-is-free:-t8-dig-order:-[K♣→t2]-[2♣→cell]-[7♣→8♥-:-8♥-WHERE:-post-chunk-9♠-:-chunk-⟸-K♣--done-:-ORDER:-K♣→t2,-CHUNK→K♣-(cap-4!!:-cells-then:-{6♠ J♦ 9♥}-FULL--chunk-cap-=-1--—-F-F-F-F-F.-—-chunk-BEFORE-9♥-celling?!-:-9♥-celled-at-t1-dig-⟸-needed-for-5♣-⟸-4♥3♣-⟸-t2-dig-⟸-K♣-seat-⟸-chunk-:-⟹-chunkAFTER-9♥-celling-FORCED--:-chunk-cap-with-{6♠ J♦ 9♥}:-1--—-—-J♦-THE-NEW-CANCER.-—-⟹-J♦-celling-DELAYED-till-after-chunk?! :-J♦'s-celling-was-for-J♥→Q♠-(5♦-access-for-4♣):-DELAY-4♣-resolution:-4♣→CELLearly-(as-always)-then-4♣-cell→5♦-LATER-when-5♦-frees-!!!:-cells-rotation:-4♣-celled-[t2-dig-…-5♦-freed]:-5♦-freed-⟸-J♥→Q♠-⟸-J♦-celled-:-⟹-{6♠, 4♣, J♦}-overlap-window-until-4♣→5♦-drains:-then-{6♠ J♦}+1-rotator-:-—-AND-9♥?!-9♥-celled-[t1-dig…]:-OVERLAP-{6♠ 4♣ 9♥}-before-J♦-even-:-⟹-rotator-slot-SINGLE:-timeline-:-(1)-{6♠}+2:-…-(2)-+9♥-(t1-dig):-{6♠ 9♥}+1:-(3)-+4♣-(t2-dig):-{6♠ 9♥ 4♣}-FULL-:-(4)-NEED:-t6-dig-(9♦8♠→10♣--no-cell;-8♥→CELL--FULL)--—-8♥-alternative-seat-pre-chunk:-NONE-—-.-⟹-⟹-THE-TRIANGLE-{9♥ 4♣ 8♥}-verdammt.-—-⟹-dig-t6-BEFORE-t2?!:-(3')-+8♥:-{6♠ 9♥ 8♥}-FULL:-J♥→Q♠-⟸-J♦-cell--FULL--AAAAAAAAAAAARGH.
Citation: https://www-cdn.anthropic.com/d00db56fa754a1b115b6dd7cb2e3c3..., section 6.2.2
You're not going to get a handle on what it's doing. The thinking traces are there to make you feel better about yourself.
Anthropic specifically calls out that example as
>"illegible reasoning in a few reinforcement-learning environments over long rollout"
Yet, I get the point that you're making: those tokens essentially are an internal scratchpad for the LLM which isn't required to logically lead to the output.
This video presentation of the paper you linked was interesting: https://www.youtube.com/watch?v=hUp3zh23aHw
I'm referring to models where the actual traces are available. Eg. I've been using a local Qwen3.8-Next-Flash lately.