We thought the same thing. Proprietary or open didn't matter, because with enough attempts at something you can eventually sus out its operations and optimize accordingly (or distill, as we've seen with LLMs).
The only solid way to reduce gaming the benchmarks is to ensure they're always, always changing. Enough pelicans on bicycles, start asking for turtles piloting gliders, or rabbits pushing skateboards. Broaden the focus, not narrow it into repetitive measures.
> Proprietary or open didn't matter, because with enough attempts at something you can eventually sus out its operations and optimize accordingly (or distill, as we've seen with LLMs).
Please explain how, if I'm OpenAI and I'm making ChatGPT 5.7, and I release it, and Artificial Analysis goes off and runs one of their proprietary benchmarks on it from a random account, how I can optimize for that benchmark.
Please imagine that the benchmark runners are not making mistakes of the fourth grade level - which they won't be. If you assume this level of incompetence, then literally everything is possible.
And it also is completely irrelevant to my challenge to show how proprietary benchmarking can be gamed, because it presumes (absolutely insane and divorced from reality) circumstances that have nothing to do with benchmarking as a concept or process.
Artificial Analysis was an clearly meant to be an example. I was obviously talking about the ideal scenario of proprietary benchmarking, not how it might be being screwed up in practice.
The only solid way to reduce gaming the benchmarks is to ensure they're always, always changing. Enough pelicans on bicycles, start asking for turtles piloting gliders, or rabbits pushing skateboards. Broaden the focus, not narrow it into repetitive measures.