Where Do You Actually Run an Evaluation? Matching LMArena, HELM and lm-eval-harness to the Need
Seven evaluation platforms, organised by the need each one serves rather than by rank: human preference, trade-off transparency, reproducible plumbing, open-weights shortlisting, standardisation, custom task suites and hardware speed. For teams deciding where to run their next evaluation.










