Measuring whether an agent actually works: datasets, graders, regression suites, production traces.
Nothing filed under Evals yet.