Five models sit within 0.9 points. The harness moves scores by twenty.
SWE-bench Verified leaders cluster inside one point, while the same weights swing 10 to 20 points across harnesses. What a leaderboard cannot tell you.
Priya Nadkarni writes about evaluation, which is the least glamorous and most load-bearing part of shipping an agent. Golden datasets, graders, regression suites, and the uncomfortable question of what a benchmark score actually predicts in production.
She is sceptical of any eval that has not caught a new failure in a month, on the grounds that it has stopped measuring capability and started measuring memory.
Where a vendor publishes a benchmark, she looks for the sample, the grader and the date before she looks at the number.
Tips, corrections and data sets go to priya@readhandoff.com. If you are reporting an error in a published piece, quote the sentence. It gets fixed faster.
SWE-bench Verified leaders cluster inside one point, while the same weights swing 10 to 20 points across harnesses. What a leaderboard cannot tell you.
Surveys put agent pilot failure at 86 to 89 percent, and the most-cited blocker is not the model. It is evaluation infrastructure nobody budgeted for.
Automated graders agree with each other beautifully. A 2026 RAND study found that says nothing about whether they agree with reality.