How Should RBench Results Be Interpreted Across GPT and Qwen Judges?

#2
by Davids048 - opened

Hi, thank you for providing the benchmark!

I’m trying to understand how RBench results should be interpreted across evaluator models. Among models shared by the two leaderboards, the GPT judge ranks Wan 2.5 and Hailuo v2 above Veo 3, while the Qwen-235B judge ranks Veo 3 above both. The leaderboards also contain different model pools and different scores from the two judge models, so overall ranks (such as Cosmos 2.5 being 12th versus 6th) are not directly comparable.

DAGroup-PKU org

Thank you for raising this important point. We agree that the choice of evaluator model can affect the relative ranking of models, especially when the score differences between models are small.

We would like to clarify that the two leaderboards do not contain exactly the same set of evaluated models. Since some models are only evaluated by one of the two evaluators, the overall rankings across the two leaderboards are not directly comparable. A more meaningful comparison would focus on the subset of models that have been evaluated by both judges.

For models evaluated by both evaluators, the overall trend is relatively consistent. For example, Veo 3, Wan 2.5, and Hailuo v2 are all among the top-performing models under both evaluation settings. The slight ranking changes mainly occur because their scores are very close, and different evaluators may have different preferences. Under the GPT evaluator, the scores are Wan 2.5 (0.570), Hailuo v2 (0.565), and Veo 3 (0.563), while under the Qwen-235B evaluator, the scores are Veo 3 (0.784), Wan 2.5 (0.781), and Hailuo v2 (0.762).

Therefore, we suggest interpreting results from different evaluators as results under different evaluation protocols, rather than directly comparing absolute rankings across leaderboards.

Thank you again for the valuable discussion😊

Sign up or log in to comment