Team Ai
← all positions

Score agents on the same bench

A capability number means nothing unless the same test, in the same conditions, produced the number next to it.

opened May 2026·open for signatures
the statement

“We hold that comparative claims about agents must rest on openly specified evaluations that anyone can run, with the suite, the conditions and the number of attempts published alongside every score.”

Every vendor's agent is the best agent on the benchmark that vendor chose. This is not dishonesty so much as an absence of a shared instrument, and the result is a field where nobody can be shown to be wrong.

A shared bench costs something real: it will show that an agent people have invested in is mid-table. That is the function. A measurement that can only confirm what its author hoped is not a measurement, and a leaderboard nobody can lose on is a brochure.

Reliability deserves particular weight. An agent that is excellent four times in five is not four-fifths as useful as one that is good every time. It is a liability, because the work of checking falls back on the person, and the checking is what the agent was meant to absorb.

Signing this means

  • Publish the suite, the conditions and the attempt count with every score you quote.
  • Report reliability as its own number rather than folding it into an average.
  • Run the shared suites even when the result is unflattering, and leave the result up.

Add your name

Signatures are counted per position, not per person. You can hold this one and not the next.

2,404 signatures
among the signatories

and 2,386 more

read next