GPT-6 Astra
openai.com1370 points by kibae 8 hours ago
1370 points by kibae 8 hours ago
System Card: https://deploymentsafety.openai.com/gpt-6-astra
Related ongoing threads:
OpenAI's GPT-6 Astra on ARC-AGI-3 - https://news.ycombinator.com/item?id=49555691
GPT-6 Astra makes major gains in the Artificial Analysis Coding Agent Index - https://news.ycombinator.com/item?id=49556147
Related: OpenAI begins rolling out GPT-6 Astra - https://news.ycombinator.com/item?id=49554273 How about we stick to that one for talking about the rollout, and this one for talking about the model? The ARC-AGI-3 scorecard is extremely misleading given that it clearly states itself that "with [the responses API] harness, we estimate Sol would score in the ballpark of ~30%." but it shows a score of 7.8% for GPT-5.6 Sol presumably since if they updated the percentage for GPT-5.6 Sol to the score it would receive with the responses API harness they used for GPT-6 Astra they'd have to do the same for the percentage they show for Opus 5 which would similarly be much higher. Regardless, the result is still valid as the original benchmark harness is definitely unreasonably handicapped, and if a harness alone can help the LLM saturate the benchmark with a near perfect score then the combination of the two must still be effectively AGI in the sense of passing the most famous benchmark designed specifically to measure AGI progress, after multiple iterations of progressively making it harder. I think it is fair to say that this is probably effectively AGI if the benchmarks are remotely accurate - even with Fable, I've been at the point personally where I am reasonably confident that there's essentially nothing that I am better than Fable at despite generally being substantively above average on human benchmarks. If Astra's this much better than Fable, I'm ready to call AGI here. For the many people who resist the AGI label possibly ever being achieved, I'd be curious to hear takes on what would make you think Astra is yet to be AGI, and what would still need to be achieved for this to effectively be AGI from this point forward. Take it from the mouth of the creator of ARC-AGI: When we released ARC 3, I got asked, "when do you think a frontier model will saturate it?", and I answered "in about a year, though it depends on how much it gets explicitly targeted" That was 6 months ago, so the progress that Astra represents happened about 2x faster than I anticipated. I think the speed of progress will surprise a lot of people, and what the new models can do will challenge the views of AI that people developed by using prior generations of models. I feel like AGI's definition got watered down, and these tests do not cover the original definition, what is your definition and thoughts on aligning with what all of us understood from the original claim? I feel like this test is just helping someone like Sam Altman pretend like he implemented AGI as originally pitched for an IPO when in fact, he has not. Shameful. > AGI is essentially the equivalent of a median human that could be hired as a remote co-worker... capable of performing any task that one would be satisfied with a remote colleague doing via a computer. - Sam Altman on AGI I can call a coworker right now and have a real time conversation with them without feeling like I’m talking to a frustrating machine. Most of them, anyway. But I guess that’s “moving the goalposts”. I can call coworker right now and have conversation so frustrating that I wish I was talking to machine instead. Here's another definition of AGI from Sam Altman: https://www.nytimes.com/2023/11/20/podcasts/hard-fork-sam-al... Sam Altman: Let’s say we make an A.I. that is really good, but it can’t go discover novel physics. Would you call that AGI? Kevin Roose (New York Times): I probably would, yeah. Would you? Sam Altman: Well, again, I don’t like the term, but I wouldn’t call that done with the mission. If the new AGI benchmark is "be Einstein/Feynman" then we've hit AGI. What if it can be Einstein, but can't draw a Pelican, write a solid college-level essay, or fold clothes? The ability to do a ton of book learning in training, and pull in tons of related context at once, is superhuman in some ways, but lags a lot in others. > What if it can be Einstein, but can’t draw a Pelican, write a solid college-level essay, or fold clothes? Then it’s an expert system. Stephen Hawking wasn’t very good at folding clothes. The ‘General’ part of the term ‘AGI’ seems like a trap to me, because there will always be new workflows to master. Can Astra one-shot level completion on some yet-to-be-released video game? If no, does that mean it’s not yet ‘Generally’ intelligent? You won’t get pure ‘general’ intelligence until you find Einstein’s hidden variables and load the state of the entire universe into context. Meanwhile, building a series of expert systems targeting specific valuable workflows is useful today and seems like it’ll continue to scale to cover huge swathes of economically valuable workflows. I think that’s the more interesting thing to be measuring. The surface area of useful economic workflows that can be addressed with expert systems built with today’s tech. Hitting some ‘Artificial Expert Intelligence’ coverage threshold on economically valuable workflows is what will matter for humans well before pure ‘general’ intelligence. Folding clothes is happening. https://www.youtube.com/watch?v=cRZNwgvcWUg AI in math is ongoing. https://spectrum.ieee.org/ai-in-mathematics Adding sibling comments, I think some people may be overestimating how well the median human can draw a pelican, or create an SVG of a pelican (depending if we’re comparing to an image generation model, or SVG generation). Pelicans are a solved problem at this point. An open-weight model on my own machine gave me this: https://crimson-jeri-74.tiiny.site/ And the only reason LLMs can't write essays indistinguishable from human output is because they aren't RLHF'ed to write like humans. Folding clothes isn't an LLM's job but if you were to insist, they could certainly do it, as any number of videos from robotics labs will attest. That particular future is already here but definitely not evenly-distributed. How good was Einstein at drawing pelicans on bicycles by writing SVG code? Checkmate, meatbags. Wouldn’t that mean producing novel work like relativity and QED? I would maybe argue that Einstein was the most LLM-like of great thinkers. A lot of his great discoveries were mostly that he was very knowledgeable about the bleeding edge research in a number of disparate areas, and was able to have the aha moment where he could make the connections for how to integrate them. A lot of other thinkers who created new fields from scratch are probably way harder for an LLM to crack. That is very aligned with an LLMs ability to have superhuman knowledge in wide areas. What is novel physics? I think they mean improve our understanding of physics with new theoretical results or paradigms. Like if it’s 1899, would Astra develop General and Special relativity on its own? This is as good a time as any to note that we might be closing in on a new conceptual revolution in our own time as it relates to holography and an information centric approach to spacetime. Obviously It's the furthest possible thing from a guarantee, but it has much of the enthusiasm and motivation that string theory had previously enjoyed in prior decades. So it could be a natural experiment for whether AI can contribute to novel physics. Specifically, there's a big question about weather. Something like our informational understanding of black holes where information inside it is equivalent to information on its boundary (which I'm sure I'm not saying correctly), might be generalized to regular space-time. More people should be freaking out with excitement about this and perhaps it's something to which AI can contribute. There was symbolic AI programs in the 1980’s that “discovered” Kepler’s laws and the resulting solar system model from just tycho brache’s astronomical observations. That was the the very first “new physics” ever. i wonder if we could train a modal, and omit all data prior to 1899, and see what happens? Do you mean after? People do this!! But I think it’s a bit different. It won’t be apples to apples because the data volume I think is just so much different. Maybe there are good experiments for something like this. I assume solving one of the major open problems of physics? Would this be possible without it being able to run novel real-world physics experiments autonomously? (Note: I am not suggesting we let it do this. Please don't, in fact) AI could discover candidate novel physics without autonomously operating new physical experiments, and humans or instruments can later independently validate the result. This is analogous to how Einstein developed theories whose predictions were confirmed by experiments and observations only years or decades later. Do we have any examples of an current day AI system introducing a novel concept or perspective. We've got plenty of counterexamples discovered and some theorems proven, but afaik nothing analogous to a new definition. That’s already been done. I know of at least one novel result contributed by Claude to frontier physics. I’m sure there is more. For it to be like a human it wouldn't just need to solve existing phsyics problems, it would need to push the field forward and introduce new paradigms. Solving "open problems" will push the field forward. My comment wasn't very long, yet you somehow still ignored the main part, "and introduce new paradigms". The point is whether it can do everything humans can, entirely new theoretical frameworks and ideas, such as string theory or dark matter, are not coming out of AI at the moment. If it makes you feel any better the curmudgeon who drives the ARC-AGI tests feels the same way, that's why we're on 3 and I'm sure we'll see 4. Also, we can all avoid calling it "moving the goalposts" so no one feels talked down to. I wonder if Altman's definition also includes taking on the same liability as a coworker would. Probably not. Would 100% need to be fully insured for liability, with a sizable war chest that OpenAI cannot even afford. This implies that human employees don’t have insurance. But they do. My company’s cyber insurance for example covers breaches due to employee mistakes. Most companies also have umbrella liability policies. It’s just that for now, AI “employees” need a LOT more coverage. And that's the rub, isn't it? If it can replace a worker but does too much work to be checked routinely by a human, and bears no real responsibility for its actions, well... it's really just a way to jack up the value of the settlement the company using it gets to pay out when it does something that causes a lawsuit. If OpenAI had just simply stuck to making "good enough" models that were open sourced (like they promised they would be when starting out) and could be used to augment a human doing a task - a human that could be given actual consequences for messing up - they wouldn't have burned all of this money trying to reach this nebulous definition of AGI. Hell, "good enough" is what many open-source models are, and that's what terrifies Altman. What do you mean by bearing no real responsibility for its actions? If I use a model to accomplish a task and it fails, I use something else to try to accomplish the task. If it's my responsibility to complete the task, it can't be the model's responsibility unless I've agreed to some sort of guarantee from the provider. If the provider says "the model will always be right or your money back" then the provider has got responsibility. If there's no guarantee, there's no responsibility on their part, just on the person whose job it is to try and solve a problem with the model. The "dream" that these labs are mostly selling is the ability for capital to subscribe to their AI for cheaper than it costs a human to do some task. Not to have a "human + AI hybrid where the human is responsible". It's what the whole AGI valuation is based off of, in that scenario, with no human oversight, the agent they lease has to be responsible for the task? It's not really a dream though? You can subscribe to their AI right now and complete many tasks for cheaper than it would cost to pay a human to do those tasks. Another person can do the same thing but also keep a human in the loop. You and the other person may compete in the market for whatever your product or service is, and you both might do very well or one of you might do better than the other because of a whole host of different reasons. Nothing in the process of developing and selling access to a more advanced LLM requires the customers to do away with human labor, nor does it require them to offer the LLM service with a guarantee that it will never make any mistakes. So far the LLMs have always made lots of mistakes and the companies sure keep making a lot of money. > What do you mean by bearing no real responsibility for its actions? If you give an "intelligent agent" offered by one of these model providers a task of updating the content of your website, and it updates it with inappropriate adult content, who incurs the cost of the machine's error? The model provider generally does not. It it makes a mistake and deletes your website from AWS, who is responsible? If it targets another website because it decides that it is "part" of your website and attempts to break into it, who is responsible? In all of these cases, it's you. It would be the same if you downloaded an open model, ran it locally, and it happened to make the same catastrophic mistakes. The consequence to the provider is that if they offer a product that does these things, people don't buy the product. In general, the person whose job it is to provide the company with a working, non-adult website and not hack into other websites is the one who would receive consequences for failing to meet those expectations.
dang - 7 hours ago
intenex - 6 hours ago
mvkel - 6 hours ago
giancarlostoro - 6 hours ago
chrsw - 5 hours ago
azan_ - 5 hours ago
petilon - 4 hours ago
hackerbrother - 3 hours ago
jmalicki - 2 hours ago
reasonabl_human - 16 minutes ago
AareyBaba - an hour ago
no-name-here - 37 minutes ago
CamperBob2 - an hour ago
vel0city - 42 minutes ago
block_dagger - 2 hours ago
jmalicki - 2 hours ago
throwawayq3423 - 3 hours ago
astro1234 - 3 hours ago
glenstein - 2 hours ago
adastra22 - 3 hours ago
verelo - 3 hours ago
astro1234 - 2 hours ago
senderista - 3 hours ago
type_enthusiast - 3 hours ago
petilon - 3 hours ago
auntienomen - 2 hours ago
adastra22 - an hour ago
colordrops - 3 hours ago
petilon - 3 hours ago
colordrops - 24 minutes ago
refulgentis - 4 hours ago
lenerdenator - 6 hours ago
giancarlostoro - 6 hours ago
m-s-y - an hour ago
lenerdenator - 5 hours ago
alex0015 - 5 hours ago
Avicebron - 5 hours ago
alex0015 - 4 hours ago
degamad - 5 hours ago
alex0015 - 4 hours ago