jean-technologies/user-simulation-leaderboard
Industry Evals
Published results for systems that model a person, one table per benchmark, with the source of every number. Within a table, systems are comparable. Across tables they are not, and nothing is averaged.
Two tiers. Reported transcribes numbers as their authors published them; every tool-transcribed row has been checked against its source table. Jean-run is the fair comparison on frozen snapshots with one protocol, and it is a seed rather than a ranking.
The full board with the methodology is at <https://jeantechnologies.com/evals>.
Data
Corrections and submissions
Open a discussion here or write to <jonathan@jeantechnologies.com>.
How this page is built
build_space.py reads the same two exported files the website reads and writes this page with the data inlined. No number is typed in by hand. A refresh is one command and a push.
