Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

We thought the same thing. Proprietary or open didn't matter, because with enough attempts at something you can eventually sus out its operations and optimize accordingly (or distill, as we've seen with LLMs).

The only solid way to reduce gaming the benchmarks is to ensure they're always, always changing. Enough pelicans on bicycles, start asking for turtles piloting gliders, or rabbits pushing skateboards. Broaden the focus, not narrow it into repetitive measures.



> Proprietary or open didn't matter, because with enough attempts at something you can eventually sus out its operations and optimize accordingly (or distill, as we've seen with LLMs).

Please explain how, if I'm OpenAI and I'm making ChatGPT 5.7, and I release it, and Artificial Analysis goes off and runs one of their proprietary benchmarks on it from a random account, how I can optimize for that benchmark.


Search for accounts @artificialanalysis.com. read their chat history. optimise for that.


Please imagine that the benchmark runners are not making mistakes of the fourth grade level - which they won't be. If you assume this level of incompetence, then literally everything is possible.


Why would I assume anything other than maximum incompetence from the AI ecosystem?


This is such a ridiculous and shallow cop-out.

And it also is completely irrelevant to my challenge to show how proprietary benchmarking can be gamed, because it presumes (absolutely insane and divorced from reality) circumstances that have nothing to do with benchmarking as a concept or process.


Does OpenAI give Artificial Analysis early access to test models? If so it's definitely not "some random account".


Artificial Analysis was an clearly meant to be an example. I was obviously talking about the ideal scenario of proprietary benchmarking, not how it might be being screwed up in practice.




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: