Evidence before
a leaderboard.

No verified model measurements have been published yet. We won’t turn example data, synthetic runs, or community votes into performance claims.

What will appear here

Reproducible evaluations across JudgeBench, RM-Bench, and RewardBench 2, reported under our diagnostic pairwise protocol. These are separate from each dataset’s official scoring protocol.

Community observations are a separate signal

A user vote expresses preference. An independently checked reference supports correctness. Browser-submitted results can be modified; we label their verification status and keep them separate from maintained research runs.

Explore starter challenges · Read the methodology