What is Chatbot Arena and why is it important?
easyAnswer
- LMSYS-run open leaderboard where two anonymous models answer the same user prompt and humans vote which is better; Elo ratings computed from millions of pairwise comparisons.
- Foundation for a leaderboard less prone to benchmark contamination (fresh prompts every day, human judgments).
- Limitations: (1) casual users bias toward friendly / verbose outputs; (2) prompt distribution skews toward common tasks; (3) topic mix isn't stable.
- Standard reference for 'how do humans rank frontier models today'.
Check yourself — multiple choice
- Automatic benchmark
- Human pairwise votes on real prompts → Elo leaderboard; less contaminated but biased toward friendly verbose outputs (LMSYS Arena)
- MCQ only
- Same as MMLU
Chatbot Arena: LMSYS human-vote Elo; robust to contamination but biased toward verbosity.
#evaluation#benchmarks
Practise LLMs & GenAI
214 interview questions in this topic.