Goodhart's Law Comes for Every Benchmark You Trust

cacm.acm.org

97 points by pseudolus 6 days ago


StilesCrisis - a day ago

Had to stop reading when the article devolved into Claude spam. "defensible in isolation," "honestly ranked," ugh. Please write your own blog post.

xnx - 2 hours ago

Ultimately, the only evals that matter are for the tasks you're interested in. The only way to game those is by being better.

throw10920 - 19 hours ago

The obvious solution is to have non-public benchmarks.

It is exceedingly difficult to train on a proprietary benchmark administered by someone with half a brain (i.e. don't sign up for a ChatGPT account with your benchmark@artificialanalysis.ai email) - you have to find a tiny needle in a vast haystack.

In fact, it can be difficult enough that it's simply not economically viable - that is, that it's cheaper to make the model better than it is to try to find the account running the benchmark.

In the limit case, the benchmark is indistinguishable from...normal problems that need to be solved.

astro1234 - a day ago

I agree and that’s why we need and indeed have an ever evolving landscape of benchmarks

> Private, refreshed test sets attack the mechanism itself, and in my view they are the only intervention that does. If the questions have never touched the public Web, they can’t be in the training data; if they rotate, memorizing this year’s set doesn’t help next year.

That’s what we have. A fresh public benchmark is also good, and teams do make efforts to decontaminate training data but there’s likely just no great way around leakage.

Btw, lots more issues in benchmarks than the ones discussed; for instance you can leak answers from the questions themselves or in the case of e.g. multiple choice formats in the actual answers. You just pass the MCQ choices themselves to the model and it may be able to guess way above chance. Coding agent benchmarks sometimes forget to delete .git. They mention e.g. a 6.9% error rate in one of the benchmark items, this seems pretty typical and I would actually be fine shipping that.

Benchmarks are very ugly, but if they didn’t exist we would need to invent them. All of the problems above and more do not explain the progress we see. There are probably 50,000 benchmarks in the literature and new ones get created frequently with varying levels of quality and usefulness.

cryptolobster - 3 hours ago

Benchmarks are fine, but people take them way too seriously.

A model can crush some test and still be a pain to use in real life.

zahlman - a day ago

Clearly, the solution is to judge society by how many currently-un-gamed benchmarks it has produced.

throwatdem12311 - a day ago

This reminds me of many years ago when Mozilla/Firefox (I think it was) said that they stopped focusing on mainstream benchmarks because they didn’t really translate to real world browser performance gains.

I view these AI benchmarks the same. No I do not care that GPT got 1200 on FartAGIMaX-4.0-Extreme and Claude got 1350. I care about how much it costs and how correctly it does the tasks that I give it. Unfortunately the only way to know is to use them all myself and measure it myself.

At the end of the day these things are all so damn close in how they behave in whatever harness so it realy just does boil down to whatever is actually cheapest.

This is why Deepseek is great: it’s so much cheaper it doesn’t matter if I burn way more tokens because it’s still orders of magnitude cheaper than the US SotA models. If it doesn’t get it quite right immediately I just do a few more turns and then it’s fine. Barely an inconvenience.

ddp26 - a day ago

Not forecasting though. You can't goodhart predicting real-world events

teddyh - a day ago

“When you place a tangible value on trust, trust becomes a commodity to be bought and sold.”

— <https://news.ycombinator.com/item?id=27432186>

hotsalad - 2 hours ago

Broken link?

- a day ago
[deleted]
Legend2440 - a day ago

I think Goodhart's law is just a consequence of correlation vs causation.

It is very easy to find a metric that is correlated with what you want. But once you start trying to influence a system, you quickly push it out of the range where the correlation holds.

In order to optimize for something, you need to maximize the actual causative variable. This is much harder.

qsera - 18 hours ago

Everything, every. single. thing. that can cause people to spend money on, will be ultimately manipulated.

Will be true as long as humanity exists.

cyanydeez - a day ago

obviously, the best benchmark is the one you tell no one about.

functionmouse - a day ago

Jokes on them, I don't trust benchmarks

Once something becomes a benchmark it is no longer a good benchmark.

federicoTXTS - 5 days ago

[flagged]

kzmttkc - a day ago

[flagged]

Ozzie-D - 21 hours ago

[flagged]

rdevilla - a day ago

[dead]