Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

Saturation is mostly just selection effects in play. Throw out the "90% easiest" of tasks, and what remains is a jagged ladder of high difficulty outliers.

Hard to climb, and hard to measure the climb - because you have less effective data points and the datapoints themselves are less linear, while you're still being subject to the measurement noise.

Not having the mislabeled tasks would reduce the saturation, but it wouldn't drive it to zero. Even without the "infinite difficulty tasks", bell curve would do its thing.



You almost can't give a correct answer to a task with expected wrong answer unless you know it expects wrong answer. ... maybe on a yes/no question by chance. But as you become more knowledgeable, the probability that you give wrong but expected answer reduces.

If a benchmark saturates to 100% it's very likely that answers leaked into the training data.

In college I had a funny exam. It was on C++. One question I had to answer incorrectly because there was a mistake in the question. So I gave two answers for it, one that answered the question as it was and the other that answered the question as I inferred it was intended to be. It was appreciated. I got a honorary mark above the top possible (I gave correct answers to all other questions). I wouldn't be surprised if across so many, so huge benchmarks, there were tasks with wrong questions or answers in the key, that some LLM answered in and expected manner in the same fashion.




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: