Evidence before
a leaderboard.
No verified model measurements have been published yet. We won’t turn example data, synthetic runs, or community votes into performance claims.
What will appear here
Reproducible evaluations across JudgeBench, RM-Bench, and RewardBench 2, reported under our diagnostic pairwise protocol. These are separate from each dataset’s official scoring protocol.
- Accuracy and coverage by task and language, with uncertainty.
- Cases where Jev and a comparator disagree, checked against evidence.
- Order sensitivity, probability calibration, and high-confidence errors.
- Latency and cost measured under disclosed, consistent conditions.
Community observations are a separate signal
A user vote expresses preference. An independently checked reference supports correctness. Browser-submitted results can be modified; we label their verification status and keep them separate from maintained research runs.