GPT-5.6 vs. Claude Fable 5 for Physical AI, which performs best?

juliahub.com

68 points by mbauman 6 hours ago


draginol - an hour ago

So Fable "won" but it cost $124.76 for marginal performance benefits over the $22.56 5.6 Sol run.

xnorswap - 4 hours ago

A really frustrating partial presentation, given an apparent lack of testing with a spread of efforts for each model.

Given that there's no reason to believe that Fable's xhigh is comparable to GPT-sol's xhigh, or Opus xhigh, for that matter, it would be far more useful to see the effort level where these tasks no longer achieved their goals.

DwarvenEngineer - 12 minutes ago

Honestly, I just hate the term "physical AI". They're robots. It's unfortunate that we had to adopt a term with the words AI in it, just to get investor's attention.

giwook - 5 hours ago

Please forgive my naivety, but are world models (once they are in a consumer-ready form) expected to outperform any currently existing LLM on these sorts of tasks (i.e. of the physical world)?

hartator - 4 hours ago

It's kind of interesting this is already out of data as it's missing Kimi 3 and Opus 5.

effnorwood - 2 hours ago

Define "best" and "performs"

grim_io - 5 hours ago

I'd expect google to do well here, since they were historically strong at multimodal and physics.

jespinel - 5 hours ago

Nice! It is missing Codex in the agent harnesses comparison IMO.

arisAlexis - 3 hours ago

Google with apptronic should have good models soon

gizmodo59 - 5 hours ago

Yet another "benchmark to promote their own harness"