What is Arena-Hard-Auto?
mediumAnswer
- Auto-benchmark by LMSYS: uses 500 hard prompts drawn from the tail of Chatbot Arena topics, and GPT-4 as a judge to pick winners between two models.
- Correlates well with human Arena Elo but runs in an hour, not months.
- Cheap, reproducible proxy for 'how would humans rank this model?'.
- Free from prompt gaming.
- Widely used to score new open-source models against frontier baselines.
Check yourself — multiple choice
- Human-only benchmark
- 500 hard prompts + GPT-4 judge → cheap reproducible proxy for Chatbot Arena Elo, hour-scale not month-scale
- Only for math
- Random prompts
Arena-Hard-Auto: hard prompts + LLM judge → cheap Chatbot-Arena proxy.
#evaluation#benchmarks
Practise LLMs & GenAI
214 interview questions in this topic.