EasyDeepLearn

What is MT-Bench and how does it differ from Chatbot Arena?

medium

Answer

  • MT-Bench (LMSYS 2023): 80 multi-turn prompts across 8 categories; GPT-4 judges each response on a 1-10 scale.
  • Fixed prompt set → deterministic, reproducible.
  • Best for comparing models systematically on chat quality.
  • Weaknesses: (1) 80 prompts is small; (2) known prompts get gamed after publication; (3) GPT-4 judge has biases (favors GPT-4-like styles).
  • Arena is broader but slower; MT-Bench is fast and precise.
Check yourself — multiple choice
  • Same as Arena
  • 80 fixed multi-turn prompts + GPT-4 judge on 1-10 scale — reproducible but small and gameable
  • Only for math
  • Only human eval

MT-Bench: 80 prompts + GPT-4 judge; reproducible chat quality proxy.

#evaluation#benchmarks

Practise LLMs & GenAI

214 interview questions in this topic.

Related questions