What is MT-Bench and how does it differ from Chatbot Arena?
mediumAnswer
- MT-Bench (LMSYS 2023): 80 multi-turn prompts across 8 categories; GPT-4 judges each response on a 1-10 scale.
- Fixed prompt set → deterministic, reproducible.
- Best for comparing models systematically on chat quality.
- Weaknesses: (1) 80 prompts is small; (2) known prompts get gamed after publication; (3) GPT-4 judge has biases (favors GPT-4-like styles).
- Arena is broader but slower; MT-Bench is fast and precise.
Check yourself — multiple choice
- Same as Arena
- 80 fixed multi-turn prompts + GPT-4 judge on 1-10 scale — reproducible but small and gameable
- Only for math
- Only human eval
MT-Bench: 80 prompts + GPT-4 judge; reproducible chat quality proxy.
#evaluation#benchmarks
Practise LLMs & GenAI
214 interview questions in this topic.