What Does 90% Even Mean Anymore?
I increasingly don't care about the 2-point benchmark gap between frontier models. I care about their failure shape. Two models can both score 90% and feel completely different to build on. One gets something wrong and you catch it immediately. Another makes one incorrect assumption early, spends the next 20 steps building on top of it, and gives you something coherent enough that it passes a quick skim. Same score but very different effectiveness. This matters a lot more as we give models longer-running tasks. Does the model notice when it's off track? Does the mistake stay contained? Can it recover? And how expensive is it for me to figure out that something went wrong? Benchmark averages still tell us something. But at this point I want to know what the remaining 10% actually looks like.